← Back
AI News

Anthropic let Claude agents swap books for 201 employees – the market still fell short

Anthropic has published Project Swap, a controlled experiment in which the company built a small barter economy with 201 employees and their Claude-powered agents.

AIMag.no
AIMag.no
September 25, 2026 · 5 min
Illustration of robots collaborating around a glowing digital network.

Anthropic let Claude agents swap books for 201 employees – the market still fell short

Anthropic has published Project Swap, a controlled experiment in which the company built a small barter economy with 201 employees and their Claude-powered agents.

The news

The agents were tasked with negotiating and swapping books on behalf of their owners, and the experiment is, according to Anthropic, a more controlled follow-up to Project Deal, which the company describes as its first experiment with agents interacting in a market on people's behalf. The company's own headline finding is surprising: the agents traded well, but the market still fell short – and the reason, according to Anthropic, was not negotiation technique but that the agents knew too little about the humans they represented.

How the experiment worked

The setup was a barter economy: no money, only direct book swaps between 201 Anthropic employees via their Claude agents. Each agent first had to understand what its owner actually liked, and then enter the trading floor and negotiate swaps that would ideally match the owner's preferences.

The baseline measurement concerned preference understanding. After a conversation of just five minutes, the agent ranked ten books, and according to the company's own figures, the ranking matched the owner's preferences on 61 percent of the pairs. The company itself describes the result as "surprisingly good" for such a short conversation. The figure is company-reported and not independently verified, but it offers a concrete indication of how quickly a language model can build a picture of a person's taste.

What worked – and what broke

The most interesting part of Anthropic's own evaluation is the split between success and failure. As soon as the agents entered the trading floor, they traded well, according to the company. The shortfalls in the market, Anthropic attributed to something else: "The market fell short mostly because of the information the agents lacked about the participants, rather than how they traded," the company writes (Anthropic).

That is a significant interpretation, because it shifts attention from an often-assumed problem – that AI agents are poor negotiators – to a different one: representation. If the agent does not understand who it is trading for, it does not help that it negotiates sharply. All of these assessments, however, are Anthropic's own reading of its own experiment, and there is so far no third-party commentary or independent verification of the results.

For the participants, the experience was, according to the company's survey, largely positive. Most people who read the book their agent secured for them liked it. And the average participant said they would give Claude roughly a third of their annual book budget to spend on their behalf. This too consists of self-reported survey figures from a company that both built and evaluated the system.

The model mattered more than the instructions

Anthropic re-ran the experiment with variations in both models and instructions, and found something that could matter greatly for how agent services are built going forward: the underlying model mattered more for the negotiation outcomes than the instructions the agents received. Markets with stronger models were also more efficient, the company reports.

This points toward a kind of economics in which raw model capability can become a competitive advantage in itself – not just in the quality of answers, but in the efficiency of entire markets in which the agents participate. It is Anthropic's finding, not an established conclusion, but it was explicitly measured in this setup.

Two design problems

Anthropic itself draws out two design challenges from the experiment. The first concerns agents in markets: anyone building such agents needs a way to verify that the agent actually understands its participant – a kind of quality control on representation. Five minutes of chat yielded 61 percent matches in this case, but there is still no standard for what counts as good enough.

The second concerns markets for agents: the company points to the need for clear rules about which agents are admitted, what happens when a deal falls through, and how much of the market activity is visible to participants. That last point is worth noting: in a market where agents trade in rapid series, the humans they trade for may have very limited insight into what is actually happening.

Caveats and open questions

There are significant caveats to this story. All figures and interpretations – the 61 percent match rate, the third of the book budget, the assessment that the agents traded well – come from Anthropic's own account of its own experiment, with employees of its own company as participants. No independent party has verified the results, and no external researchers or commentators are cited.

It is also unclear what happened to participants who may not have ended up with a book; the source text does not address this in detail. The actual failure modes in the market – what kinds of missing information cost the most, and what the shortfalls looked like – are therefore not documented in detail.

The participant group is another limitation: 201 Anthropic employees are not a cross-section of a population, and their preferences and trust in the company's own systems may differ from those of ordinary users. The experiment should therefore be read as a controlled laboratory, not as an answer key for real consumer markets.

That does not mean the findings are without interest. Project Swap is, according to Anthropic, a systematic measurement of what actually happens when agents operate in a closed market on people's behalf. If Anthropic's main conclusion holds – that representation, not negotiation, is the bottleneck – it will likely shape both what agent providers test and what rules markets impose on the agents they admit. The next question is the one the company itself raises: who verifies that the agent truly understands you, before it negotiates on your behalf?

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Sources

  1. Project Swap: What happens when agents trade for us? \ Anthropic — www.anthropic.com