00 Walkthrough

Attention Is All You Need, Really Need

We hold less and less of what we ship. AI does blast out code, and that part is real: the volume numbers are enormous and they are not in dispute. The open question is whether we are using it well, and whether understanding still matters.

Twenty four numbers. Every figure traces to a primary source and carries the date of that research. Fourteen passes, each finding attacked by a verifier told to refute it rather than confirm it. Space to advance.

01 Understanding is the job 2017 · pre-AI baseline
58%understanding it
24%finding it
5%typing it
13%everything else

This is 2017. Before any of this existed, typing was already the smallest part of the job. That is the part we automated.

Xia et al, IEEE TSE 44(10), DOI 10.1109/TSE.2017.2734091. 78 professional developers, 3,148 hours logged across every application on the machine, not just the editor. Sums to 100. "Typing" is literally changing characters: reading the code you are about to change counts as understanding, finding it counts as navigation. Say the year on camera. It has to be pre-AI to be a baseline, and nobody has replicated it since.

02 Understanding is the job January 2026
67%Anthropic study: wrote it by hand
50%Anthropic study: used AI

Quizzed on code they wrote minutes earlier. One group scraped a pass. The other failed. And that was a chat window, not an agent.

Shen and Tamkin, Anthropic, published at ICSE-SEIS '26, January 2026. Anthropic researchers, and note the tool they tested was OpenAI's GPT-4o, not their own model. 52 developers, 29 with 7+ years, 35 minutes on a Python async library none had used. Anthropic's own write up calls the gap "nearly two letter grades". The tool was a GPT-4o sidebar chat. The one study that used a file-editing agent instead (arXiv 2607.26375, July 2026) found a bigger gap still: asked to identify their own code, the agent group scored 0.59 against 0.96. Two separate studies, different tasks, so say it as a direction and not a ladder. Nobody has measured a frontier agent, on real code, with professionals.

03 The expectation 9 March 2026

200%

is what Anthropic advertises. Code output per engineer, up two hundred percent.

@bcherny, X, 2026-03-09. Posted on launch day, quote-posting the product the number justifies. "Code output" is never defined anywhere. Up 200 percent means three times, not two.

04 The expectation Your own room Mine. Firsthand, not a study.

15

pull requests per developer per day. I have sat in the room where that was the target.

No survey has ever asked leaders how much more they expect. Nobody has ever put a number on it. So this one is yours, and it is worth more than a statistic because you were there. Say it as yours.

05 And nobody checked the cost August 2026
45%leaders working more hours
41%say their team is less motivated
1 in 3managers considering leaving

We are asking for a lot more, from people nobody is checking on.

600 engineering leaders, August 2026. Never say AI causes burnout. Say demands moved and wellbeing was never measured. Both are true and only one is defensible.

06 And nobody checked the cost Search covers 2023 to 2026

0

studies have measured burnout in engineers with AI as the exposure. Since 2023.

Systematic search, validated instruments only. Not a small number. Zero. So when someone tells you the data shows AI is not hurting engineers, the data does not say anything, because nobody collected any.

07 The reality 2025 to 2026
200%Anthropic, the launch post
51%Meta, diffs per developer
50%Anthropic, own research team
24%Microsoft, agentic rollout
6%Google, internally

Four of the biggest engineering organisations on earth measured themselves. Every measured number lands between six and fifty one. The number in the room is two hundred.

Anthropic, Meta, Microsoft and Google, each measuring itself, which is what makes them hard to dismiss. Meta: diffs per developer per month up 51 percent, with more than 80 percent of that increase attributed to agentic AI; RADAR telemetry covers 535,000+ reviewed diffs at 25,000 a day, though the developer count behind the 51 percent is never given. Microsoft, arXiv 2607.01418, the cleanest purely-2026 measurement that exists: Claude Code and Copilot CLI rolled to tens of thousands of engineers on 5 January 2026, +24.0 percent merged pull requests per engineer per day through 29 April, 95 percent interval +14.5 to +33.7, p<0.001, against a synthetic counterfactual. It does not fade: +29.4 percent in February, +20.0 percent March to April. Dose response within the same person: +15.0 percent at three days a week, +50.1 percent at five or more. Caveats to say aloud: Microsoft measuring Microsoft, merged PRs reward small frequent PRs, and if you quote the tool head-to-head (Copilot CLI +24.9 against Claude Code +11.4) you must say Microsoft owns GitHub. Anthropic: Societal Impacts December 2025, n=132 self reported, against 200 percent in their own March launch post. Google: about 6 percent internally, and 12.5 percent is the default in their own buyer-facing ROI calculator.

08 The reality Data to April 2026 · agentic
+44%more work shipped, new repositories
nonein the legacy codebase. No measurable gain

The gain is real and it concentrates in new code. In the legacy monolith it vanished.

arXiv 2607.01904, CMU and Stanford, 2 July 2026. 802 developers, 196,212 non-bot pull requests, 364 repos, January 2024 to April 2026, so it spans the agentic shift. Throughput in post-2022 repositories +0.364, p<0.001. Pre-2022 legacy +0.117 with no significance marker: say "no measurable gain", never "a 12 percent gain". The difference between the two is itself significant at p<0.01, and that test is what makes it a finding. The Stanford researcher behind the original pre-agentic version of this split is a co-author here, and Google's own DORA guidance from February 2026 still quotes that older number. Caveat aloud: one company, one legacy monolith, and repository age is a proxy for complexity, not a measurement of it.

09 The reality METR · 24 February 2026
19% slowerwith AI than without · 2025
18% fasterwith AI than without · re-run to 2026

Take the same task. Let one developer use AI and the other not. In 2025 the AI one finished later. In the re-run they finished sooner. Neither gap is big enough to be sure it is real.

METR, "We are Changing our Developer Productivity Experiment Design", 24 February 2026. Real software engineers, median 10 years, working real issues on their own repositories. Each issue is randomly assigned: AI allowed, or AI not allowed. Then METR compares how long the two groups took. 2025 run: the AI group took 19 percent longer. Re-run, August 2025 to February 2026: returning developers 18 percent quicker, new recruits 4 percent quicker. Why neither counts: the re-run's real answer could sit anywhere from 38 percent quicker to 9 percent slower. A range that wide includes "no difference at all", so the experiment cannot rule that out. Same problem with the original. METR has since put a warning banner on the 2025 study. Never quote the 19 percent alone.

10 The reality METR · 24 February 2026

30-50%

of developers stopped putting their real work into the study, rather than risk being told to do a task without AI.

METR, "We are Changing our Developer Productivity Experiment Design", 24 February 2026. How the study works, because the number only makes sense once you know: a developer submits real issues from their own repository, and each issue is randomly assigned to "you may use AI" or "you may not". You do not know which bucket yours will land in. So developers began quietly not submitting work at all, rather than gamble on having to do it by hand. 57 experienced developers, median 10 years, 143 repos, 1,134 randomised real issues, August 2025 to February 2026, a third of them randomised inside 2026. One developer completed none of the AI-disallowed tasks at all. METR's words: developers "would not want to do 50% of their work without AI, even though our study pays them $50/hour", and their own verdict on the result is that it is "only very weak evidence". When people will not give up the tool long enough for you to measure it, you have stopped measuring productivity and started measuring dependence.

11 What actually grew GitHub · data to Q1 2026
+7%public code pushes, growth in 2024
+78%public code pushes, growth in 2026

Every push is somebody sending code somewhere. The base doubled in two years, and the growth rate went up twelvefold.

GitHub Innovation Graph, git_pushes.csv, Q1 2026 released 7 July 2026, CC0. 184,167,232 pushes in Q1 2024 against 380,012,766 in Q1 2026: 2.06x, or 4.22 million a day. Year on year by quarter: 6.6, then 15.7, 30.9, 40.0, 46.1, and 78.4 percent. A balanced panel of the 150 economies present in all 25 quarters gives 2.062x, so this is not roster growth. Say "pushes", not "pull requests". For a PR-specific figure: Octoverse 2025 has 47.5 million created per month against 39.5 million.

12 What actually grew Snapshot · 24 August 2026

10,028,364

agent pull requests opened, across six tools, from a standing start in 2025.

prarena.ai, snapshot 2026-08-24 01:51 UTC. Six coding agents that did not exist as pull request authors before 2025. Weekly run rate roughly doubled: 106,317 a week across the first eight full weeks against 213,503 across the last eight, so 2.01x in fourteen months. July 2025 against July 2026 is 2.21x. Volume only. Do not quote a merge rate off this: the six use different draft workflows, Codex alone is 64 percent of the total, and Cursor's own success rate prints as 103.8 percent. Public repos only, detected by branch prefix, so agents are about 2 percent of public PR creation. The repo carries no licence, so do not claim one on screen.

13 What actually grew DX · Q1 to Q2 2026
34%of code AI generated, Q1 2026
52%Q2 2026. One quarter later

More than half the code is now written by AI. It crossed the halfway line in a single quarter.

DX, State of AI Impact in Engineering, Q2 2026. 500+ organisations. Over the same four quarters, throughput per engineer rose 37 percent, from 1.42 to 1.94 pull requests a week, while DX's own developer experience index fell from 67 to 65. And the piece a human has to look at got bigger with it: median pull request size went from 44 lines to 72 between July 2025 and June 2026 across 400+ companies. More code, bigger changes, fewer people enjoying it.

14 What actually grew NBER · agents, Feb to Dec 2025
17.3xmore code written
1.3xmore software released

Seventeen times the code. One point three times the software. Everything in between is where it went.

NBER Working Paper 35275, "Writing Code vs. Shipping Code", 100,000+ GitHub developers with usage telemetry. Cumulative multipliers for adopting every generation up to async agents: lines of code 17.3x, files touched 3.9x, commits 2.8x, pull requests 2.5x, repositories touched 1.5x, releases 1.3x. Each rung down the pipeline the gain shrinks. Synchronous agents alone raise lines changed 741.3 percent and releases 20.3 percent. Broken out by tool it gets worse: Claude Code +29.2 percent on releases, GitHub Sync Agent +1.3 percent, OpenAI Codex +0.2 percent. Two of the three are statistically indistinguishable from zero. Window: agent outcomes measured through 2 February 2026, with a 2022 autocomplete component stacked underneath, so the 17.3x mixes two eras. Say the tool-level release numbers if anyone pushes: they are cleaner. And "releases" is a GitHub release counter, not shipped value. This is the whole argument in one number pair, from an economics working paper with no product to sell.

15 And we call review the bottleneck Faros AI · April 2026
+157%time to first review
+200%time spent in review
+441%median time in review

Three measures of the same squeeze, and every one of them points the same way.

Faros AI, "The Acceleration Whiplash", published April 2026. 22,000 developers, 4,000+ teams, two years of telemetry, analysis as of March 2026. Exact figures 156.6, 199.6 and 441.5 percent. Never say "since 2024". Faros compared each team's two lowest AI adoption quarters against its two highest: the charts are labelled "% change from low to high AI adoption". The words 2024 and 2023 appear zero times in the report. Read their next sentence aloud: "The engineers with the deepest knowledge of the system are spending their most valuable hours unraveling plausible-looking code that should never have reached them in the state it did."

16 Review was never deep July 2026

49%

accuracy judging AI written assertions that were wrong. Same confidence either way.

arXiv 2607.08885. 86 programmers, two thirds professionals. 74 percent accurate on the correct ones. Add a vague explanation and people got less accurate and more confident at the same time.

17 Review was never deep Faros AI · April 2026

+31%

more pull requests merged with no review at all. Human or agentic.

Faros AI, "The Acceleration Whiplash", published April 2026. Exactly 31.3 percent, comparing each team's lowest AI adoption quarters against its highest. And say their explanation, not yours, because it stops the accusation that engineers got lazy: "We do not believe this reflects a deliberate decision to bypass oversight. The more likely explanation is that reviewers cannot keep pace with the volume of AI-generated code arriving for their attention." And Meta says the same thing about itself: the share of diffs reviewed within 24 hours is dropping, and they have never published that number. There is a measured downstream cost too: across May 2025 to May 2026, agentic contributions needed significantly more corrective maintenance and introduced more security weaknesses, with each 10 point rise in a project's no-review rate associated with roughly 6 percent more agentic maintenance burden (arXiv 2607.09902).

18 Review was never deep MSR 2026

1 min

median time from opening an agent pull request to merging it, in small repositories.

Jellyfish, roughly 1 million developer-weeks, published 13 August 2026: agent pull requests merge at about 61 percent against about 79 percent for humans, which is the clean comparison PR Arena could not give you. Yoshioka et al, MSR 2026. DOI 10.1145/3793302.3793567. 72.6 percent merge inside five minutes. Nobody read that.

19 The place people call brain rot June 2025 · still a preprint

0 of 18

could quote their own essay, minutes after writing it.

MIT Media Lab, arXiv 2506.08872. 54 students, SAT essays, not code, still a preprint. The authors published a page asking people to stop saying "brain rot," "harm" and "damage" about it. Correct it, then use their real number: nothing rotted. It never went in.

20 So people blast it out Jellyfish · April to May 2026

84%

of developers top out at one or two agents at a time. The ones running four or more still spend over eighty percent of the session watching a single one.

Jellyfish, State of AI in Software Engineering H1 2026. Sub-study windows April and May 2026. Agent-opened PR share at the most advanced companies went 10 percent in January 2026, 14.5 percent by end of February, 35 percent by April, against roughly 2 percent a year earlier. The median company stayed near 2 percent. This is the title of the video, measured: attention does not parallelize.

21 So people blast it out 13 August 2026

388 → 180

pull requests opened by automated routines. Then merged.

@bcherny, X, 2026-08-13, archived. Anthropic's own maintenance routines, org-level, six platforms, after AI review plus human review. Do not turn this into a percentage. He never wrote one.

22 So people blast it out 13 August 2026

208

never accounted for. He published a numerator and half a denominator, and stopped.

Same post. And across 11,048 closed agentic pull requests, 33 percent of rejections leave no recorded reason at all. Nobody is keeping the ledger. That is the whole point.

23 Because you sign it Live · August 2026

0 of 1,935

AI systems have ever been sanctioned. All 149 sanctions landed on a person.

AI Hallucination Cases database. 144 lawyers, 3 judges, 2 self-represented. The database has a column for the AI tool. No AI has ever appeared in the party column.

24 Because you sign it Live · August 2026

63x

more cases in three years. The accountability moved zero.

Same database. 16 cases in 2023. Over a thousand this year with a third of the year still to run. You cannot make an agent liable. There is nothing to put there but you.

25 The ask

You still have to
know what is
going on.

Speed up the parts that cannot hurt you. Slow down the parts that can. Decide which is which before you start, not at two in the morning.

Understanding is not a bottleneck. It is the job. And it is the only part of this you cannot hand to something that cannot be held responsible for it.