Companies deploying AI coding agents are generating more code without producing more software, according to Harvard University research that examined aggregated analytics data from engineering intelligence platform Jellyfish. The study, conducted by researchers Fiona Chen and James Stratton, found "little evidence that firms increase software output or reduce employment" following the introduction of AI coding tools.
The bottleneck appears in code review. After AI agents were introduced, the review process took significantly longer on average, pull requests became more likely to require revisions, and reviewers spent more time evaluating each submission. The productivity gains from faster code generation were, in the researchers' words, "absorbed by downstream constraints in the production process."
The finding challenges the central promise of the AI coding tool industry, which has attracted billions in venture funding on the premise that tools like GitHub Copilot, Cursor, and Devin would let smaller teams ship software faster. Microsoft, Google, and Amazon have all integrated AI code assistants into their development workflows, and enterprise adoption has climbed steadily since 2023.
The Harvard study's methodology relied on Jellyfish's aggregated analytics, which tracks engineering metrics across hundreds of companies. By comparing code output, review times, and deployment frequency before and after AI tool adoption, Chen and Stratton isolated the effect of the tools from other variables affecting engineering productivity.
Their conclusion suggests that the constraint on software production has shifted rather than disappeared. Writing code was never the only step between an idea and a shipped product — it sits alongside design, review, testing, and deployment. When AI accelerates one stage, the stages downstream must absorb the increased volume or the overall system stalls.
Code review in particular resists automation because it requires judgment about intent, architecture, and maintainability — qualities that AI-generated code often obscures. Reviewers evaluating AI-written submissions report difficulty assessing whether the code solves the right problem, not just whether it runs.
The employment finding is equally notable. Despite fears that AI coding tools would displace software engineers, the Harvard data showed no measurable reduction in engineering headcount at firms that adopted them. Companies appear to be using AI to handle more code volume with existing staff rather than cutting teams.
The research adds to a growing body of evidence that AI's productivity effects vary sharply by task and context. Studies of AI in customer service and writing have found measurable gains, while research on complex knowledge work has shown more modest results. Software development, with its long chain of interdependent steps, may be particularly resistant to end-to-end acceleration from a single tool.
Chen and Stratton's findings have not yet been peer-reviewed, and Jellyfish's customer base — primarily mid-size to large technology companies — may not represent the broader software industry. But the direction of the result is consistent with earlier reports from engineering leaders who noted that AI-generated code increases review burden without proportionally increasing shipped features.