300 Million Work Events Show: AI Agents Write More Code, Not More Software
New research built on data from over 700,000 employees at more than 700 software companies finds 30 percent more lines of code after AI agents were introduced – but no statistically significant increase in the number of resolved software features.
When companies adopt AI coding agents, their employees measurably write more code – 30 percent more lines of code, 20 percent more commits, and 23 percent more pull requests on average. Yet fewer of the tasks that actually matter get resolved: the resolution rate for Issues and Epics tracked in tools like Jira does not change in a statistically significant way after the AI tools are introduced.
That is the core finding of new research from Harvard University, presented by researchers Fiona Chen and James Stratton and reported by Ars Technica in October 2026. The study is among the first to measure the effect of AI agents on actual software delivery at the enterprise level, not just on code volume – and it points toward an uncomfortable possibility: that the gains from automated code generation are being eaten up elsewhere in the development process.
What the Study Is Built On
Chen and Stratton used aggregated analytics data from Jellyfish, a company that measures the detailed output of engineering teams. The data covers 300 million individual "work events" – such as commits and pull requests – and data from issue-tracking software, spanning more than 700,000 employees at over 700 relevant software companies, from 2021 through March 2026.
What makes the dataset interesting is that it links two sources that are rarely seen together: detailed activity in version control and development tools on one side, and task tracking in systems like Jira on the other. That makes it possible to ask the question most prior measurements have skipped: Does more code actually mean more software delivered?
The answer in this data is no – at least not when software is measured as resolved Issues and Epics, that is, coherent software features. The researchers also found no "compositional shift": the size or complexity of tasks did not change across the AI adoption. This means the pattern cannot easily be explained by teams simply taking on bigger chunks at a time, or by resolving many more small tasks and fewer large ones.
How the Numbers Were Calculated
An obvious objection to this type of study is selection bias: companies that adopt AI agents may be a different kind of company than those that don't, or they may be in a particular growth phase. Chen and Stratton address this with a difference-in-differences approach, according to Ars Technica's walkthrough of the methodology.
The researchers first identified when the various organizations adopted AI assistants versus AI agents. They did not do this using the companies' own statements, but through direct measurement of AI usage and analysis of GitHub activity. They then ran regression analyses of key variables before and after adoption, at different times in different organizations – and compared the change in companies that adopted the tools with those that had not yet.
The method does not remove all uncertainty, but it is a standard way of drawing causal inferences from observational data, and it is far more robust than simple before-and-after comparisons of output.
Why More Code Isn't More Software
The study documents a gap: up in code production volume, flat in delivered functionality. Why the gap arises is a matter of interpretation.
The Japanese AI Times newsletter (October 10, 2026, written by 江守義樹 at ALL WEB CONSULTING) offers one interpretation of the finding: the gains are absorbed by human review capacity. If code generation becomes cheaper and faster, but every pull request still has to pass through human review, it is the review – not the generation – that sets the ceiling on how much actually gets finished. "Even if costs go down [in the earlier stages], throughput will not increase unless review capacity also increases," AI Times writes in its analysis of the Ars Technica-reported research.
It is worth emphasizing the limitation: "review as bottleneck" is the commentator's interpretation, not a conclusion Chen and Stratton themselves draw in the available source material. Ars Technica's article notes that significant review effort is required, but the specific mechanism – that the gains are "absorbed" – comes from AI Times. The study establishes that the gap exists; the mechanism behind it remains open.
Still, the interpretation is plausible enough that it has practical consequences if it holds. It would follow that companies measuring the effectiveness of AI adoption through generation volume or API costs are measuring the wrong things. AI Times instead suggests metrics such as review wait time (how long a pull request waits before being reviewed) and lead time to merge (the time from code submission to when it is merged in). These are metrics that capture where in the process work actually stalls – and that would reveal a review bottleneck directly, rather than implying it indirectly through missed feature delivery.
What a Company Can Take Away
For organizations that have already adopted AI agents, or are considering it, the study points to three concrete things:
First: code volume is a weak indicator of value. A 30 percent increase in lines of code without a corresponding increase in resolved Epics means that volume metrics can create a false sense of progress – or, worse, be used to justify investments that don't pay off.
Second: if review is truly where the gains disappear, then review capacity is an investment decision on par with the AI licenses themselves. More or faster review rounds, better automated testing, and clearer merge criteria may be the precondition for agent gains to materialize as delivered software.
Third: the data covers the period through March 2026 – an early phase of agent adoption. How the tools and work processes evolve from here – for example, whether AI takes on a larger role in review as well – could change the picture considerably.
Open Questions
The study's numbers deserve a degree of academic humility. The underlying paper by Chen and Stratton is not available in this source material, and whether it is peer-reviewed, or where it may have been published, is unknown. All figures come via Ars Technica's reporting on the research, and AI Times itself states that its own figures are unverified source information – the newsletter could not access the primary sources directly.
Questions the study does not answer also remain: How much of the extra code is usable? Does it sit as unresolved pull requests, or is it merged in as code that moves no Jira task? And what does the picture look like over time – does review capacity recover as organizations adapt, or is the gap structural?
But as one of the first large-scale, enterprise-level measurements of AI agents' effect on actual software delivery, the main message is nonetheless clear – with the caveat that the interpretation of the mechanism is not the researchers' own: if code generation is no longer the bottleneck, companies must figure out what is – review, requirements clarity, or something else entirely – and measure their way to the answer.

