When the Same Agent Task Can Cost 30 Times More or Less
Nearly nine in ten companies use AI regularly, and most are experimenting with AI agents. Yet only 37 percent can attribute any positive impact to their AI programs on operating margin (EBIT).
That is the core diagnosis of McKinsey's Technology Trends Outlook 2026, and it puts words to a problem that has followed the agent wave since it began: the effort is enormous, the returns fail to materialize. In late September 2026, McKinsey partners addressed the problem publicly, with hard numbers on both sides of the ledger — both on usage and results, and on the costs driving the gap between them.
This article walks through the numbers, the mechanisms behind the cost variations, the coding agents' mixed results, and the advice McKinsey itself gives for governing usage and spend. All the underlying data is McKinsey's own, reported afterward by CIO, Fortune, and Business Insider — the reports themselves are not available, so nothing here has been independently verified.
Usage Yes, Returns No
According to the Technology Trends Outlook 2026, 89 percent of companies use AI regularly, and a majority are at least experimenting with AI agents. But between usage and economic success lies, in McKinsey's own words, a significant gap: "We find that only 37 percent of companies attribute any positive EBIT impact to their AI programs, let alone newer agent-based rollouts," the consultancy writes in the report, as quoted by CIO.
The figure is worth reading precisely: the 37 percent applies to AI programs overall — not specifically to agents, which are the newest and most expensive part of the portfolio. The share seeing measurable positive returns from the agents alone is therefore likely lower, even though the sources do not give a separate figure.
Why the Agents Are So Expensive
The mechanism behind the cost problem is that agentic systems do not answer a single query, but carry out multistep tasks: they plan, call tools, read results, adjust course, and repeat. According to McKinsey's own calculations, a single agent workflow requires between five and 30 times more compute than a typical chatbot query (CIO).
In a recent survey, 93 percent of respondents said this led them to exceed their AI budgets, according to McKinsey via the same source. It is a striking figure, but also one that must be handled carefully: CIO reports it without naming the survey, sample size, or method, so it cannot be verified from the available source material.
To this comes a separate and perhaps even more troubling feature of agent economics: the cost is not predictable from task to task. Such tasks are often multistep processes where the same goal can be reached in many ways, and costs vary correspondingly widely. According to a McKinsey study, quoted by senior partner Lari Hämäläinen to Business Insider via Yahoo Finance, the execution cost for different agents could diverge by up to 30 times — for the same type of task.
A caveat is needed here: the CIO article states that a workflow requires 5–30 times more compute than a chatbot query, while the study Hämäläinen cited refers to up to 30-fold cost variation between runs. It is unclear from the sources whether this is two sides of the same measurement or two distinct findings — the first concerns agent workflows versus chat, the second variation between runs of the same task. Readers should therefore not merge them into a single "30x" claim.
The Coding Agents: Gains for Some, Declines for Others
The clearest documented example of uneven returns comes from software development — the field McKinsey itself believes has the greatest potential. In a McKinsey study, only a quarter of companies using AI agents for software development managed to significantly accelerate their development cycles; McKinsey counts at least a doubling of productivity for more than a quarter of teams as significant progress (CIO).
According to CIO's summary of McKinsey's data, roughly 80 percent of developers saw productivity gains of only about 3 percent, while the top 20 percent gained 55 percent — and productivity fell at 30 percent of companies after the introduction of coding agents. These splits, too, are reported without study names, samples, or dates, so the methodology cannot be verified from the source material. But the picture they paint is consistent with the cost data: the gains are concentrated among a minority, while the majority is left with marginal effects — and a significant minority actually finds things getting worse.
A Spending Picture Pointing One Way
Against this stands the spending side. In McKinsey's State of AI 2026 survey, about a third of organizations spend more than 10 percent of their technology and communications budgets on AI. Sixty percent of respondents planned to increase AI spending next year, and around one in five said AI spending was beginning to create constraints in operating costs (Business Insider via Yahoo Finance).
Taken together, this paints a paradox that McKinsey senior partners Tanguy Catlin and Lari Hämäläinen raised in a McKinsey Live seminar titled "Improving the Economics of Agentic AI," reported by Fortune (via MSN): the models are getting cheaper per token, but the bills are getting bigger — because the agents use so many more tokens per unit of work.
McKinsey's Countermeasures: Measure the Task, Not the Token
It is worth noting that McKinsey does not conclude that the agents are worthless — the consultancy is a major stakeholder in the market and argues that the problem is primarily one of governance. The advice falls into two categories: evaluation and cost control.
Evaluate at the task level. Hämäläinen argued in the seminar that companies should assess agents at the task level rather than per token, with a rule of thumb: an agent can make sense when the time it takes to verify the output is a small fraction of the time a human would spend doing the task from scratch. If a task takes a human one hour, but the agent's work can be verified in six minutes, then an agent with a success rate above 10 percent can already start creating value, he said (Fortune via MSN). The logic is that verification — not generation — is the human cost that determines the net gain, and that an agent with a high error rate can still pay off if the errors are caught quickly and cheaply.
Three levers on the cost side. Catlin pointed to three governance levers. The first is visibility: companies need insight into which use cases, business units, agents, models, and users are driving consumption, he said according to Fortune. The second is workflow optimization — model routing (cheaper models for simpler steps), caching, and limiting tool calls and agent loops, meaning the number of times an agent is allowed to execute steps in the same task. The third is sourcing discipline: removing unused licenses, managing quotas, and negotiating terms.
This advice is sensible in itself, but it also points to an organizational reality behind the numbers: most companies apparently lack task-level evaluation and spend governance today. That 93 percent exceed their AI budgets is not a technology problem alone — it is a governance problem.
The Capability Debate: When Numbers Are Misread
A side story illustrates how easily agent metrics are misread. McKinsey cites METR, an independent benchmarking institute, in its reporting on what the agents actually accomplish. For Claude Opus 4.6, METR estimates that the "50 percent time horizon" for software development tasks lies at roughly twelve hours — that is, for tasks of this size, successful autonomous completion is expected in half of the cases (CIO).
But METR researcher Sydney Von Arx has, according to CIO, which relays MIT Technology Review, pointed out that McKinsey has misinterpreted the metric: the twelve-hour figure refers to the time a human would need for the task, not how long the AI can work autonomously. The metric thus measures human task time as a reference point, not the AI's capacity in hours.
The CIO article presents the METR critique as a settled verdict of misinterpretation, but it is itself reported secondhand, and no known response from McKinsey exists in the source material. This is worth noting for two reasons. First, it shows how fragile the basis is for the most aggressive readings of agent capability — the twelve-hour figure has been circulated as evidence that agents can handle a full workday on their own, but that is not what the metric says. Second, it is a confirmation of McKinsey's own point: there is little in the way of a common, well-understood measurement methodology for what the agents actually deliver per task and per dollar — which is precisely the gap the companies themselves must close.
What McKinsey Says About the Future — and What Remains
Despite the numbers, McKinsey remains optimistic. According to CIO, the company estimates that AI agents for software development, if successfully deployed, could unlock nearly a trillion dollars in added value for companies. This is thus McKinsey's own forward-looking estimate, not a verified result. McKinsey also argues that the tools are getting better — with reference to agent tools from Anthropic, OpenAI, and frameworks like LangGraph.
There are, then, two stories in this material, and they are not equally well supported. The first is the data story: a clear and consistent description of an economic gap between high effort and low returns, with concrete mechanisms — multistep workflows that multiply consumption, variation between runs that destroys cost predictability, and evaluation happening at the wrong level. This story rests on McKinsey's own surveys and studies and therefore cannot be independently verified from the available source material, but it is internally consistent across three different publications.
The second story is the future story: that the payoff will come as the tools get better, the workflows are correctly designed, and the governance tooling matures. It is a reasonable hypothesis, but it is McKinsey's own — from a company that profits when companies continue to invest. The questions that remain open are essentially three: whether the 5–30x in compute and the 30x in cost variation concern the same or different measurements, whether the splits in the coding-agent data reflect samples or technology, and whether the payoff actually materializes as governance tooling matures — or whether agent economics ultimately needs cheaper models faster than companies can manage their consumption.
The latter is not just McKinsey's problem. If 60 percent of companies increase their AI budgets next year, while roughly six in ten still cannot attribute positive EBIT impact to their AI programs, the gap between expectations and returns will grow before it shrinks — and then task-level evaluation and spend governance will no longer be advice from consultants, but a prerequisite for keeping the books in balance.

