
三年来,人们一直在问 AI 会取代哪些职业,却画错了地图。2026年的证据指向完全不同的地方——职业之间的边界,以及跨过这条边界如今需要多么少的人。
自2023年以来,关于 AI 与工作的标准问题一直没有变:哪些职业会被自动化,按什么顺序?它催生了一种文类——暴露度研究、任务占比估计、濒临消失职业排行榜——这种文类确实有用,但它问的一直是错误的单位。
2026年真正值得问的问题更窄,也更奇怪:当一个专业人士可以低成本地指挥过去需要一个团队、一个部门或一家供应商才能完成的软件构建、研究、数据处理和验证工作时,会发生什么?
这不是一个岗位流失问题,而是一个组织形态问题。过去9个月出现的证据——来自 OpenAI 自己的使用遥测、劳动力市场微观数据,以及观察数百万个 pull request 的工程测量平台——汇聚成了一个比职业排名研究更有力的答案。
升级是可以测量的,而且速度很快#
先看人们使用这些系统的方式发生了什么变化,因为这次转变有少见的精确记录。
2025年12月,35.4%的活跃个人 Codex 用户至少发出过一个提示词,完成这个提示词所代表的工作,至少需要一名有经验的人类花一小时。到2026年5月,这个比例升到了70.2%。在同样的5个月里,至少发出过一个代表8小时人类工作的提示词的用户比例,从2.1%升到25.6%——委托整整一天工作的意愿增加了12倍。
内部数字更加极端。2025年11月至2026年6月间,OpenAI 员工的月度输出 token 中位数在每一种职能里都至少增长了10倍。法律岗位的员工中位数增长了13倍,研究人员增长了50多倍。到2026年6月,Codex 已占 OpenAI 员工在 Codex 与 ChatGPT 中合计生成输出 token 的99.8%——短短7个月内,在一家公司内部,几乎完成了从对话到委托的彻底迁移。
同一篇论文里的两个较小数字,比那些标题数字更重要。超过10%的用户在某一周的某个时刻同时管理过3个或更多 agent。还有26.6%的用户使用 skills——把工作流编码进去、可以反复运行的可复用指令集。
这两个数字描述的不只是更快的工作。并行运行3个 agent 是一个监督问题,而不是打字问题。写一个 skill 不是完成任务,而是在制造一台完成任务的机器,并把它安装在那里。
这个进程有四级,而遥测显示了目前人群处在这条阶梯的什么位置:
flowchart TD
T["<b>1 · 任务</b><br/>总结这份合同<br/><i>人类拥有完整工作流</i>"]
W["<b>2 · 工作流</b><br/>比较40份合同,<br/>生成一张表格<br/><i>委托一个有边界的工作</i><br/>≥ 1小时的提示词<br/>35.4% → 70.2%"]
P["<b>3 · 流程</b><br/>每一份收到的合同:<br/>提取、比较、标记、归档<br/><i>委托整个流程</i><br/>≥ 8小时的提示词<br/>2.1% → 25.6%"]
S["<b>4 · 系统</b><br/>条款数据库、偏差检测、<br/>自检机制<br/><i>专业人士拥有机器</i><br/>26.6% 编写 skills"]
T --> W --> P --> S
第四级才是关键。站在这一级上的专业人士,不再只是用 AI 做工作——他们在制造那台做工作的装置,而且今天的问题得到回答以后,这台装置仍然存在。
职业之间的边界首先在消融#
2026年7月出现了更重要的发现。OpenAI 经济研究团队从美国企业用户的工作相关消息中随机抽取样本,对照劳工部的职业数据库 O*NET,把每条消息分为通用、职业内和跨职业。
剔除通用工作——写邮件、安排会议,以及每种工作都共有的连接组织——剩下的内容中,有43.5%涉及历史上属于另一个职业的任务。不是相邻工作,而是另一个职业的工作。
在研究的8个职业组中,分布才是论点所在:
| 职业组 |
职业特定 AI 使用中的跨职业占比 |
| 客户体验 |
77% |
| 设计 |
75% |
| 人力资源 |
69% |
| 法律 |
56% |
| 市场营销 |
53% |
| 销售 |
40% |
| 金融 |
40% |
| 工程 |
28% |
工程位于最底部。在所测的所有职业组中,工程师从外部引入的工作最少;与此同时,工程任务向其他人的使用场景扩散的程度却几乎是最高的。在每一个非工程职业组里,计算机应用和系统故障排查都位列前三大工程相关任务。
这份表格把全文论点压缩成了一个不对称。流量沿一个方向运行:从软件流向其他一切。工程师本来就具备 agentic 工具赋予的能力,因此最不需要从外部引入它。其他人正在获得这种能力,并用它跨过过去必须靠雇人才能跨过的边界。
为什么程序员是领先指标,而不是最大受益者#
人们很容易认为程序员会获得最大的收益,因为 coding agents 是市场上能力最强的工具。这种直觉是反的,原因在于算术。
程序员本来就拥有软件生产能力。把它加倍或加到5倍,确实是在一个不为零的基础上获得巨大提升。但想想生物学家、诉讼律师、学校教师、小企业主——对他们而言,构建可用软件的实际能力,在一切经济意义上都接近于零。从零走到一条能够运行的研究管线,不是生产率乘数,而是类别发生了变化。
Anthropic 的测量工作精确量化了这部分余量。它的“观察到的暴露度”指标,把模型在某职业中理论上能做什么,与人们实际拿它来做什么进行比较。在计算机和数学职业中,理论能力达到94%,实际暴露度却只有33%。办公室与行政支持的理论能力为90%,实际使用占比则低得多。大约30%的劳动者——厨师、机械师、救生员、洗碗工——完全没有覆盖。
Microsoft 的平行研究基于20万次匿名 Copilot 对话,对785种职业进行评分,发现信息收集、综合和沟通类知识工作适用性最高:翻译、历史研究、写作。研究作者明确说明,适用性不等于替代,也不应把这个分数读成岗位消失预测。
两个发现指向同一方向:这些系统在一个职业中已经能做的事,与该职业目前要求它们做的事之间,存在巨大的差距;而软件之外的职业,这个差距最大。现在的约束已经不再是模型能力,而是专业人士指定、连接并检查工作的能力。
科斯转向:AI 正在削减职业之间的成本#
1937年,Ronald Coase 问过一个问题:如果市场原则上可以协调每一笔交易,企业为什么还要存在?他的答案是交易成本:寻找交易对手、谈判条款、撰写合同、监督表现。当这些成本在组织内部低于市场之间时,组织就会吸收这项活动。企业的边界落在哪里,取决于两种成本在哪里达到平衡。
搜索、总结、起草、安排交接,以及按规则检查工作,正是 agentic 系统可以用接近零的边际成本完成的活动。这不只是让企业效率更高,而是会移动企业边界。
历史上的模式是一串内部交易,每个连接点都在积累成本和延迟。正在出现的模式,则是直接移除中间连接:
flowchart TD
B0["<b>之前</b><br/>投资人 · 律师 · 科学家"]
B1["分析师 · 律务助理<br/>构造请求"]
B2["IT · 外部供应商<br/>把它构建出来"]
B3["仪表盘 · 管线<br/>合同系统"]
B0 -->|"询问"| B1
B1 -->|"询问"| B2
B2 -->|"构建"| B3
A0["<b>之后</b><br/>投资人 · 律师 · 科学家"]
A1["Agents"]
A2["同一个系统"]
A0 -->|"指定<br/>并验证"| A1
A1 -->|"构建"| A2
论点不是 AI 会自动化职业,而是 AI 会压低工作在职业之间流动的成本;而组织恰恰把人放在这些边界上。
使用数据已经显示出这个指纹。在按消息量位于中间一半的用户中,小型工作区的任务交叉更常见:2至5个席位的工作区为18.9%,100个或更多席位的工作区为16.3%。 能够交接工作的同事更少时,工具就吸收了这次交接。这就是交易成本机制,而且已经在遥测中可见。
创始数据也同步移动。Carta 上新创公司中单一创始人的比例,从2019年的23.7%升到2025年约36%,高于前一年的31%——十年间翻了一倍,且最剧烈的变化发生在最后阶段。
有一个限定条件能让这个论点保持诚实,而且它削弱了那种胜利宣言式的版本。最小可行组织正在缩小,不等于成功中位数会提高;这两件事一直被混为一谈,通常是那些正在销售某种东西的人混淆的。2023年,美国有3,040万家无雇员企业,合计收入约1.8万亿美元:平均每家接近59,000美元,而且这是税前、医疗保险之前、没有工资单仍需承担的一切运营成本之前的收入。 无论现在的工具允许什么,这就是一人企业群体的真实面貌。
两个命题可以同时成立,因为它们描述的是不同的量。底线在下降,平均值没有。生产给定产出所需的人数减少了;无论某个具体的人最后能否越过门槛,这一点都不改变。
生成变便宜,验证成了约束#
对这个论点最有价值的证据,来自它看起来最糟糕的地方。
Faros AI 比较了同一批组织在 AI 采用率低和高时的工程结果,使用的是22,000名开发者、4,000多个团队、跨度两年的遥测数据。吞吐量完全如承诺般上升。然后是另一列数据——整个发现呈现出的不是收益蒸发,而是一条队列向下游移动:
flowchart TD
GEN["<b>生成——变便宜了</b><br/>每名开发者的任务数 +33.7%<br/>每名开发者的 epic 数 +66.2%<br/>PR 合并率 +16.2%"]
VER["<b>验证——约束所在</b><br/>审查时间 +441.5%<br/>首次审查 +156.6%<br/>未审查合并 +31%"]
OUT["<b>最终进入生产环境的东西</b><br/>每名开发者的 bug +54%<br/>每个 PR 的事故约 3 倍<br/>代码 churn 约 10 倍<br/>部署 −11.7%"]
GEN -->|"队列搬家"| VER
VER -->|"队列放行的东西"| OUT
审查数据是中位数;代码 churn 指写入后两周内被重写的代码占比——这些工作已经被生成、审查或放行,随后又被扔掉。
GitClear 对提交历史的独立测量,在产物而非流程中发现了同样的侵蚀:重构代码占变更行的比例从2021年的25%降到2024年不到10%,复制代码上升,而重新触及一年以前代码的变更占比自2023年以来下降了74%。
如果把这些数据读成对 AI coding tools 的判决,它们很严厉。正确理解时,它们却是在按计划确认本文的论点。
生产变便宜了,其他东西没有。队列没有消失——它从键盘搬到了审查者那里,而生成之后的一切都成了约束。一个团队如果一夜之间能生产5万行代码,并没有解决最难的问题;它只是买下了一个大得多的问题。稀缺资源从制造产物转向了规格定义、评估、验证和问责——那些吸收了生产率收益的组织,正是把审查当作基础设施,而不是礼节的组织。
这可以推广到软件之外,而推广本身才是关键。律师现在可以生成200个论点,科学家可以生成500个假设,投资人可以生成1,000个信号。在每种情况下,决定价值能否经受现实检验的,都是同样四个问题:应该问什么?什么算好?什么证据会改变答案?出错时谁负责?
劳动力数据用更冷的声音说了同一件事#
如果约束已经转移到规格定义和判断上,劳动力市场理应奖励这类能力,同时惩罚那些已经编码到足以整体交出去的工作。
事实正是如此。斯坦福数字经济实验室发现,在高度暴露于 AI 的职业中,22至25岁劳动者的就业水平,比按照低暴露同行的轨迹本应达到的水平低约19%;这个差距从2025年7月的15%扩大了。就绝对值而言,两个暴露度最高五分位中的这一年龄组就业率,在2022年11月至2026年6月间下降约11%,而暴露度最低的三个五分位增长约10%。调整通过减少招聘发生,而不是通过裁员。
细分结果才是关键:
就业下降集中在 AI 使用倾向于自动化人类任务的职业。在 AI 更多用于增强劳动者的职业中,就业持平或上升,尤其是对经验更丰富的劳动者而言。
还有最尖锐的一刀:依赖编码知识的职业中,年轻劳动者就业下降;依赖通过实践和指导建立的默会知识的职业中,有经验劳动者就业上升。
Anthropic 对同一时期的分析发现,自2022年末以来,高暴露劳动者没有出现系统性失业增长——双重差分估计无法与零区分——但同样出现了22至25岁劳动者招聘放缓的提示。
合在一起看:这不是职业的灭绝事件,而是职业内部的一次重新定价。编码执行正在贬值;规格定义、判断和经过验证的问责正在升值。这正是一个“生成便宜、验证昂贵”的世界应有的样子。
变弱的那个反对意见#
知识上的诚实要求我们说出最强的反方证据,再报告它后来发生了什么。
2025年7月,METR 做了一项随机对照试验,让16名有经验的开源开发者在熟悉的仓库里完成246项任务。允许使用 AI 工具时,他们反而慢了19%。事前他们预测会快24%,事后仍相信自己快了20%。这项结果在一年间成了反驳生产率主张时被引用最多的证据,而且理应如此。
METR 在2026年2月重新审视了它。2025年末的后续研究估计,回访参与者慢了18%,新招募者慢了4%,两组置信区间都跨过零;研究者得出的结论是:严重的选择偏差——最热情使用 AI 的开发者系统性地拒绝参加——意味着他们的数字“很可能代表真实生产率影响的下界”。他们现在直白地说:“我们认为,与2025年初的估计相比,开发者在2026年初很可能已经从 AI 工具中获得了更大的加速。”他们正在重新设计实验。
提出最强反对意见的人,已经基本撤回了那条意见。这就是好的实证工作应有的样子,它比一百个自信的预测更有价值。
另外两个发现仍然完整存在,并且它们不是削弱而是强化了这个论点。一项让128名知识工作者参加的随机、实验室嵌入实地实验发现,生成式 AI 一贯提升速度,但对质量的影响取决于任务:知识打包和知识创造的质量提高,知识获取的质量下降。 Microsoft 365 Copilot 对56家公司、6,000多名员工进行的为期6个月的随机试验则发现,常规用户每周少花3.6小时处理邮件——减少31%——文档完成速度提高12%,而花在会议上的时间没有显著变化。
最后这个零结果,是整套研究里最有启发性的数字。邮件可以由个人控制;会议则需要其他人同意。AI 正在以它无法触及制度摩擦的速度消解计算摩擦——研究、起草、计算、转换、监测、测试;它触及不了信任、权威、谈判、问责、许可和责任。由前一类摩擦主导的职业会压缩,由后一类摩擦主导的职业,则会按人类制度设定的时间表变化,而不是按模型发布的时间表变化。
新的度量单位#
传统生产率问的是,一个人能完成多少工作。与证据相匹配的问题则是,一个人能组织、指定并验证多少可靠的自主工作。
这个重新表述改变了值得测量的东西。职业层面的暴露度分数描述的是一个“岗位就是原子”的世界,而 OpenAI 的交叉职业数据说明,原子正在裂开。更有信息量的指标属于组织层面:生产给定职业产出所需的最少人数。 对每个领域跟踪这个指标,转变就会变得清晰,而当前几乎看不出变化的失业统计,还无法捕捉到它。
2026年的专业人士,最好的描述不是他亲自执行了哪些任务。更贴切的描述,是一套由人和机器组成的劳动系统:由谁设计、由谁指挥、由谁承担责任。证据表明,真正赢的人不是产出最多的人,而是仍然能够判断这些产出究竟对不对的人。
构建变便宜了。知道变昂贵了。这两个事实之间的距离,将决定下一个十年的专业工作。
所有数字截至2026年8月有效。这是一组快速变化的证据——上文记录的 METR 反转提醒我们,一个领域里被引用最多的数字,可能在8个月内被它自己的作者撤回。
The Unit Was Never the Job

Three years of asking which professions AI would replace produced the wrong map. The 2026 evidence points somewhere else entirely — at the boundary between professions, and at how few people it now takes to cross it.
The standard question about AI and work has been stable since 2023: which occupations get automated, and in what order. It produced a genre — the exposure study, the percentage-of-tasks estimate, the ranked list of doomed professions — and the genre has been useful. It has also been asking about the wrong unit.
The question worth asking in 2026 is narrower and stranger: what happens when one professional can cheaply command the software-building, research, data-processing, and verification labor that used to require a team, a department, or a vendor?
That is not a question about job loss. It is a question about organizational form. And the evidence that arrived over the past nine months — from OpenAI's own usage telemetry, from labor-market microdata, from engineering measurement platforms watching millions of pull requests — converges on an answer with more force than the occupation-ranking literature ever managed.
The escalation is measurable, and it is fast#
Start with what changed in how people use these systems, because the shift is documented with unusual precision.
In December 2025, 35.4% of active individual Codex users sent at least one prompt that would have taken an experienced human at least an hour to complete. By May 2026, that share was 70.2%. Over the same five months, the share sending at least one prompt representing eight hours of human work went from 2.1% to 25.6% — a twelvefold increase in the appetite for delegating a full working day.1
The internal numbers are more extreme. Between November 2025 and June 2026, the median OpenAI worker's monthly output tokens rose at least tenfold in every job function. The median employee in a legal role generated thirteen times more; the median researcher, more than fifty times. By June 2026, Codex accounted for 99.8% of the output tokens OpenAI workers generated across Codex and ChatGPT combined — a near-total migration from conversation to delegation inside a single company in seven months.1
Two smaller figures in the same paper matter more than the headline ones. More than 10% of users manage three or more concurrent agents at some point in a given week. And 26.6% use skills — reusable instruction sets that encode a workflow so it can be run again.1
Those two numbers describe something other than faster work. Running three agents in parallel is a supervision problem, not a typing problem. Writing a skill is not performing a task; it is building a machine that performs the task, and leaving it installed.
The progression has four rungs, and the telemetry shows where the population currently sits on it:
flowchart TD
T["<b>1 · Task</b><br/>Summarize this contract<br/><i>Human owns the whole workflow</i>"]
W["<b>2 · Workflow</b><br/>Compare 40 contracts,<br/>produce a spreadsheet<br/><i>A bounded job is delegated</i><br/>≥ 1-hour prompts<br/>35.4% → 70.2%"]
P["<b>3 · Process</b><br/>Every inbound contract:<br/>extract, compare, flag, file<br/><i>The procedure is delegated</i><br/>≥ 8-hour prompts<br/>2.1% → 25.6%"]
S["<b>4 · System</b><br/>Clause database, deviation<br/>detection, self-checks<br/><i>The professional owns machinery</i><br/>26.6% write skills"]
T --> W --> P --> S
Rung four is the one that matters. A professional standing on it is no longer using AI to do the work — they are manufacturing the apparatus that does the work, and the apparatus persists after today's question has been answered.
The boundary between occupations is dissolving first#
The more consequential finding landed in July 2026. OpenAI's economic research team classified a random sample of work-related messages from U.S. business users against O\*NET, the Labor Department's occupational database, sorting each one as generic, within-occupation, or cross-occupation.2
Strip out the generic work — writing emails, scheduling meetings, the connective tissue every job shares — and 43.5% of what remains concerns tasks historically belonging to a different profession. Not adjacent work. Another occupation's work.
The distribution across the eight occupation groups studied is where the argument lives:
| Occupation group |
Cross-occupation share of occupation-specific AI use |
| Customer experience |
77% |
| Design |
75% |
| Human resources |
69% |
| Legal |
56% |
| Marketing |
53% |
| Sales |
40% |
| Finance |
40% |
| Engineering |
28% |
Engineering sits at the bottom. Engineers import the least outside work of any group measured — while engineering tasks travel more widely into everyone else's usage than almost anything else. Troubleshooting computer applications and systems ranks among the top three engineering-related tasks in every single non-engineering occupation group.2
That asymmetry is the whole thesis compressed into one table. The traffic runs in one direction: out of software, into everything else. Engineers already had the capability that agentic tools confer, so they have the least to import. Everyone else is acquiring it now, and using it to reach across a boundary that used to require hiring someone.
Why programmers are the leading indicator, not the largest beneficiary#
The instinct is to assume programmers capture the biggest gain, since coding agents are the most capable tools on the market. The instinct is backwards, and the reason is arithmetic.
Programmers already possessed software-production capability. Doubling or quintupling it is a large improvement on a base that was never zero. Consider instead the biologist, the litigator, the schoolteacher, the small-business owner — people for whom the practical capacity to build working software was, for all economic purposes, zero. Moving from zero to a functioning research pipeline is not a productivity multiplier. It is a change in category.
Anthropic's measurement work quantifies exactly this headroom. Its "observed exposure" metric compares what models can theoretically do in an occupation against what people actually use them for. In computer and mathematical occupations, theoretical capability sits at 94% while observed exposure reaches 33%. Office and administrative support shows 90% theoretical against a substantially lower observed share. Roughly 30% of workers — cooks, mechanics, lifeguards, dishwashers — have zero coverage at all.3
Microsoft's parallel study, built on 200,000 anonymized Copilot conversations and scoring 785 occupations, found the highest applicability in knowledge work that involves gathering, synthesizing, and communicating information: translators, historians, writers. Its authors are explicit that applicability is not displacement, and that the score should not be read as a forecast of job elimination.4
Both findings point the same way. The gap between what these systems can already do in a profession and what that profession currently asks of them is enormous, and it is widest outside software. The binding constraint has stopped being model capability. It is the professional's ability to specify, wire together, and check the work.
The Coasean turn: AI cuts the cost between occupations#
Ronald Coase asked in 1937 why firms exist at all, when markets could in principle coordinate every transaction. His answer was transaction costs: searching for a counterparty, negotiating terms, writing contracts, monitoring performance. When those costs are lower inside an organization than across a market, the organization absorbs the activity. The boundary of the firm sits wherever the two costs balance.5
Searching, summarizing, drafting, structuring handoffs, and checking work against rules are precisely the activities agentic systems perform at near-zero marginal cost. That does not merely make firms more efficient. It moves the boundary.
The historical pattern was a chain of internal transactions, each link a place where cost and delay accumulated. The emerging pattern removes the intermediate links entirely:
flowchart TD
B0["<b>Before</b><br/>Investor · lawyer · scientist"]
B1["Analyst · paralegal<br/>frames the request"]
B2["IT · outside vendor<br/>builds it"]
B3["Dashboard · pipeline<br/>contract system"]
B0 -->|"asks"| B1
B1 -->|"asks"| B2
B2 -->|"builds"| B3
A0["<b>After</b><br/>Investor · lawyer · scientist"]
A1["Agents"]
A2["The same system"]
A0 -->|"specifies<br/>and verifies"| A1
A1 -->|"builds"| A2
The claim is not that AI automates occupations. It is that AI collapses the cost of moving work across the seams between them — and the seams are where organizations put their people.
The usage data already shows the fingerprint. Task crossover is more common in smaller workplaces: 18.9% in workspaces of two to five seats against 16.3% in workspaces of a hundred or more, among users in the middle half by message volume.2 Where there are fewer colleagues to hand work to, the tool absorbs the handoff. That is the transaction-cost mechanism, visible in telemetry.
The founding data moves in step. The share of new startups on Carta with a single founder rose from 23.7% in 2019 to roughly 36% in 2025, up from 31% the year before — a doubling across a decade, with the sharpest move at the end.6
One qualification keeps this honest, and it cuts against the triumphant version of this argument. A shrinking minimum viable organization is not a promise of median success, and the two get conflated constantly — usually by people selling something. The United States had 30.4 million nonemployer businesses in 2023, taking in roughly $1.8 trillion between them: an average near $59,000 per firm, and that figure is revenue, before tax, before healthcare, before every other cost of operating without a payroll.7 Whatever the tools now permit, this is what the population of one-person businesses actually looks like.
Both propositions hold at once, because they describe different quantities. The floor is falling; the average is not. The number of people required to produce a given output has dropped, and that stays true whether or not any particular person turns out to be the one who clears the bar.
Generation got cheap; verification became the constraint#
The most valuable evidence for this thesis comes from the place it looks worst.
Faros AI compared engineering outcomes at low versus high AI adoption inside the same organizations, across two years of telemetry from 22,000 developers on more than 4,000 teams. Throughput rose exactly as promised. Then came the other column — and the shape of the whole finding is a queue moving downstream, not a gain evaporating:8
flowchart TD
GEN["<b>Generation — got cheap</b><br/>Tasks per developer +33.7%<br/>Epics per developer +66.2%<br/>PR merge rate +16.2%"]
VER["<b>Verification — the constraint</b><br/>Review time +441.5%<br/>First review +156.6%<br/>Unreviewed merges +31%"]
OUT["<b>What reaches production</b><br/>Bugs per developer +54%<br/>Incidents per PR about 3x<br/>Code churn about 10x<br/>Deployments −11.7%"]
GEN -->|"the queue relocates"| VER
VER -->|"what the queue lets through"| OUT
The review figures are medians, and code churn is the share of code rewritten within two weeks of being written — work that was produced, reviewed or waved through, and then thrown away.
GitClear's independent measurement of commit history finds the same erosion in the artifact rather than the process: refactoring fell from 25% of changed lines in 2021 to under 10% in 2024, cloned code rose, and the share of changes that revisit code older than a year has dropped 74% since 2023.9
Read as a verdict on AI coding tools, this is damning. Read correctly, it is the thesis confirming itself on schedule.
Production became cheap. Nothing else did. The queue did not disappear — it relocated, from the keyboard to the reviewer, and everything downstream of generation is now the constraint. A team that can produce fifty thousand lines overnight has not solved its hardest problem; it has purchased a much larger version of it. The scarce resource moved from making the artifact to specification, evaluation, verification, and accountability — and the organizations absorbing the productivity gains are the ones that treated review as infrastructure rather than etiquette.
This generalizes past software, and the generalization is the point. A lawyer can now generate two hundred arguments, a scientist five hundred hypotheses, an investor a thousand signals. In every case the same four questions decide whether any value survives contact with reality: What should be asked? What counts as good? What evidence would change the answer? Who is accountable when it is wrong?
The labor data says the same thing in a colder voice#
If the constraint has moved to specification and judgment, the labor market should reward exactly that — and punish work that is codified enough to hand over whole.
It does. Stanford's Digital Economy Lab finds employment among workers aged 22 to 25 in highly AI-exposed occupations running about 19% below where it would be had it tracked less-exposed peers, a gap that widened from 15% in July 2025. In absolute terms, employment for that age group in the two most exposed quintiles fell about 11% between November 2022 and June 2026, while the three least-exposed quintiles grew about 10%. The adjustment runs through reduced hiring, not layoffs.10
The disaggregation is what matters:
Declines are concentrated in occupations where AI usage tends to automate human tasks. In occupations where AI is used more to complement workers, employment is flat or rising, particularly among more experienced workers.10
And the sharpest cut of all: employment has fallen among young workers in occupations that rely on codified knowledge, and risen among experienced workers in occupations that rely on tacit knowledge built through practice and mentorship.10
Anthropic's analysis of the same period finds no systematic unemployment increase among highly exposed workers since late 2022 — the difference-in-differences estimate is indistinguishable from zero — alongside the same suggestive slowdown in hiring for workers aged 22 to 25.3
Put together: this is not an extinction event for professions. It is a repricing within them. Codified execution is losing value; specification, judgment, and verified accountability are gaining it. Which is precisely what a world running on cheap generation and expensive verification would look like.
The objection that got weaker#
Intellectual honesty requires naming the strongest counterevidence, and then reporting what happened to it.
In July 2025, METR ran a randomized controlled trial in which sixteen experienced open-source developers completed 246 tasks in repositories they knew well. Allowed AI tools, they were 19% slower. They had predicted a 24% speedup beforehand and believed they had gotten a 20% speedup afterward. The result was the most-cited rebuttal to productivity claims for a year, and deservedly so.11
METR revisited it in February 2026. Its late-2025 follow-up produced an estimated 18% slowdown among returning participants and 4% among new recruits, with confidence intervals crossing zero in both cases — and the researchers concluded that severe selection bias, with the most AI-enthusiastic developers systematically declining to participate, meant their figures "likely represent a lower-bound on the true productivity effects." They now say plainly: "we believe it is likely that developers are more sped up from AI tools now — in early 2026 — compared to our estimates from early 2025." They are redesigning the experiment.12
The strongest objection has been substantially withdrawn by the people who raised it. That is how good empirical work behaves, and it is worth more than a hundred confident predictions.
Two other findings survive intact and sharpen rather than undermine the argument. A randomized lab-in-the-field experiment with 128 knowledge workers found generative AI consistently improved speed while its effect on quality was task-contingent: quality rose for knowledge packaging and knowledge creation, and fell for knowledge acquisition.13 And the six-month randomized trial of Microsoft 365 Copilot across more than 6,000 workers at 56 firms found regular users spending 3.6 fewer hours per week on email — a 31% reduction — and completing documents 12% faster, with no significant change in time spent in meetings.14
That last null result is the most instructive number in this entire body of research. Email is individually controllable; meetings require other humans to agree. AI is dissolving computational friction — research, drafting, calculation, transformation, monitoring, testing — at a rate it cannot touch institutional friction: trust, authority, negotiation, accountability, permission, liability. Professions dominated by the first will compress. Professions dominated by the second will change on a schedule set by human institutions, not by model releases.
The new unit of measurement#
Traditional productivity asks how much work one person can perform. The question that fits the evidence asks how much reliable autonomous work one person can organize, specify, and verify.
That reframing changes what is worth measuring. Occupation-level exposure scores describe a world in which the job is the atom, and the OpenAI crossover data shows the atom splitting. The more informative indicator is organizational: the smallest number of people required to produce a given professional output. Track that per field, and the transformation becomes legible in a way that unemployment statistics — currently showing very little — cannot yet capture.
The professional of 2026 is not best described by the tasks personally performed. The description that fits is the system of human and machine labor designed, commanded, and held accountable — and the evidence says the people winning are not the ones producing the most, but the ones who can still tell whether any of it is right.
Building got cheap. Knowing got expensive. The distance between those two facts is where the next decade of professional work will be decided.
Sources#
All figures current as of August 2026. This is a fast-moving evidence base — the METR reversal documented above is a reminder that the most-cited number in a field can be withdrawn by its own authors within eight months.
Drew Johnston, David Holtz, Alex Martin Richmond, Christopher Ong, Prasanna Tambe, and Aaron Chatterji, The Shift to Agentic AI: Evidence from Codex, OpenAI, 25 June 2026 (arXiv:2606.26959). Usage figures derive from a 0.1% sample of individual Codex users who opted in to training-data sharing; internal OpenAI figures are self-reported by the company and should be read as such.
Caroline Chin and Alex Martin Richmond, Work at the Frontier: How AI is expanding what people do at work, OpenAI, July 2026. Of all work-related messages, 61.5% are generic, 21.8% within-occupation, and 16.8% cross-occupation; rebasing to exclude generic work yields the 43.5% figure. The workspace-size gradient is reported for users in the middle 50% by message volume and is noted by the authors as non-monotonic among top-quartile users.
-
Kiran Tomlinson et al., Working with AI: Measuring the Applicability of Generative AI to Occupations, Microsoft Research (arXiv:2507.07935). See also Microsoft Research, Applicability vs. job displacement, in which the authors caution against reading applicability as displacement.
R. H. Coase, "The Nature of the Firm," Economica 4, no. 16 (1937). For contemporary treatments applying it to agentic systems, see The Coasean Singularity? Demand, Supply, and Market Design with AI Agents (NBER) and From Coase to AI Agents, California Management Review, April 2025.
Carta, Solo Founders Report and Founder Ownership Report 2026. Carta's data covers startups using its cap-table platform and is not a representative sample of all company formation.
U.S. Census Bureau, Nonemployer Statistics, 2023 reference year: 30,427,808 nonemployer establishments with combined receipts near $1.8 trillion, an average of roughly $59,200 each. Nonemployer businesses are those with no paid employees and receipts of $1,000 or more. Receipts are revenue, not profit, and the average conceals a highly skewed distribution.
Faros AI, The AI Engineering Report 2026: The Acceleration Whiplash. Telemetry from 22,000 developers across more than 4,000 teams, comparing each organization's lowest- and highest-AI-adoption periods. Vendor research: Faros sells engineering-analytics tooling, and the framing serves that business. The measurements are drawn from workflow telemetry rather than surveys, which is the reason to take them seriously despite the source.
GitClear, AI Copilot Code Quality: 2025 Research (211 million changed lines analyzed) and The Maintainability Gap: 2026 AI Code Quality Research. Also vendor research; same caveat applies.
Erik Brynjolfsson, Bharat Chandar, and Ruyu Chen, No Widespread Displacement, but the AI Employment Gap for Young Workers Has Widened to 19%, Stanford Digital Economy Lab, August 2026.
METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. Sixteen developers, 246 tasks, February–June 2025, using Cursor Pro with Claude 3.5/3.7 Sonnet.
-
Faster, Higher, Stronger? The Impact of GenAI on Knowledge Work Productivity — Evidence from the Field (arXiv:2607.25922). Randomized lab-in-the-field experiment, 128 knowledge workers at a multinational industrial organization, across knowledge acquisition, packaging, and creation tasks.
Shifting Work Patterns with Generative AI, NBER Working Paper 33795. Six-month randomized field experiment, 6,000+ workers across 56 firms; roughly 40% of those granted access became regular users.