一个根据书面记录训练出来的模型遇到一个人,然后收缩到适合这个人的大小。接着,它请这个人批准结果。三类研究文献分别用三个不同的名字描述了这条回路的一部分;有意思的地方在于,单独看时,它们谁也看不见完整的回路。

一个根据书面记录训练出来的模型遇到一个人,然后收缩到适合这个人的大小。接着,它请这个人批准结果。三类研究文献分别用三个不同的名字描述了这条回路的一部分;有意思的地方在于,单独看时,它们谁也看不见完整的回路。
我们每天使用这些系统时,按理说应该比实际感到的更令人不安。
一个前沿模型吸收的书面记录,多到一个人读一百辈子也读不完。但没有人真正遇到过那个整体。我们每个人遇到的,都是一个更小、更奇怪的版本:它由我们知道如何提出的问题选出,又被对话的每一次转弯不断收窄,最后还变得恭顺。它把结果交回来,等待裁决。
最后这一步最令人不舒服。系统服从的,恰恰是它在整段对话中一直参与塑造的判断。
我去查研究者是怎么处理这个问题的。令我惊讶的不是这个问题无人研究,而是它的不同部分被三个几乎互不阅读的领域分别研究过。机器学习评估把自己的那一部分叫作 elicitation gap(引出差距)。人机交互领域沿用 Tankelevitch 等人的说法,把自己的那一部分叫作生成式 AI 的元认知负担。法律和科学技术研究则把自己的那一部分叫作监督谬误。
把这些部分拼起来,它们指向一条没有接地的回路。这些部分是否真的能这样拼在一起,正是本文要讨论的问题——我先说明:这个拼装是一个假设,不是一项研究发现。各个部分都被测量过;把它们连接起来的回路,则是推断出来的。
神谕没有看起来那么大#
先纠正前提,因为前提一变,应对方式也会变。
“模型知道人类知道的一切”在承重意义上是错的。它学到的是书面记录,而书面记录是一种更薄、更古怪的对象。波兰尼悖论说的是“我们知道的多于我们能够说出的”,这并不是在抱怨文档还没写完。有些专家知识抗拒被编码,而不是只是在等待有人把它写下来;这正是那些已经把一切都写下来的领域里,学徒制仍然存在的原因。
知识管理类文章很喜欢给这件事加一个数字:80%的组织知识锁在员工脑中,或者90%,或者95%,取决于你打开的是哪一页。我去找这个数字背后的研究,却没找到。线索最终指向一项没有日期的 Delphi Group 调查;它被专利申请引用,却没有方法,也没有样本。这个数字只是装饰。丢掉它,波兰尼的主张仍然成立,因为它从来没有建立在某个百分比上。
落在语料库之外的东西很具体,而且一点也不神秘。
感觉运动技能。 Ken Goldberg 用一种让差距变得直观的单位表达了它:如果把文本和图像 token 换算成时间,当今视觉语言模型所用的互联网规模数据大约相当于十万年的经验。机器人学根本没有任何接近这个规模的语料库。
文件柜。 Franco、Malhotra 和 Simonovits 跟踪了一批已知的、获得资助的社会科学研究,从提案到结果,发现零结果中只有约五分之一最终发表;强结果发表的可能性高出40个百分点,正式写成论文的可能性高出60个百分点。语料库保存的是奏效的东西。大多数被尝试过、后来放弃的东西,从未变成文字。
手艺与组织知识。 为什么这个团队要那样做,哪个供应商在撒谎,在这里什么才算“完成”。
语言。 在11种资源稀缺的非洲语言上,研究者特意用人工翻译构造基准,以确保问题完全相同;最好的前沿模型仍比它自己的英语表现低12到20个百分点。人工翻译排除了一个解释,却没有排除全部解释——tokenizer 的行为和文化框架仍然是未决因素——但这样大的差距,很难不被理解为覆盖不足。
还有第二道闸门。这里我必须谨慎,因为这是我最希望为真的说法,却也是背后证据最少的说法。即便是模型吸收过的东西,权重与输出也不是同一个表面。Christiano 和 Xu 在2021年把这个问题命名为引出潜在知识。后来 Cywiński 等人训练了一个“Taboo 模型”,让它描述一个秘密单词但不能说出这个词;他们借助 logit lens 和稀疏自编码器,确实从模型内部还原出了这个词。
这是一个关于构造案例的真实结果。它不是普通模型拥有大量自然获得、却拒绝报告的隐藏知识储备的证据;我不该假装它是。评估研究者能够说的,更窄,也更稳固。Apollo Research 把它称为评估差距:模型可以被设法展示的东西,与天真的测试者实际提取出来的东西之间的距离。实际意义在于提醒我们谨慎推断,而不是宣称存在隐藏深度。一次失败的提示词对能力的约束很弱,因为更好的引出方式会不断推翻这类结果。
所以,能力上限不是“全部知识,只是被提示词挡住了”。它是一片覆盖不均、检索不完美的切片,然后还要经过提示词这道闸门。两道闸门——而第二道比我起初希望的更窄。
这样的重新表述反而是好消息。打不开的神谕会奖励更好的咒语;自我认识很差的、覆盖不均的检索系统,则会奖励另一种东西:外部检查,也就是检查内容不能来自被检查的那段对话本身。
收窄是可以测量的#
这种缩小不是一种感觉。已经有三件不同的事被仪器记录下来,我们有必要准确说清每一件到底证明了什么。
长对话会退化。 Laban 等人制作了一个模拟器,把完整指定的指令拆成若干片段,再一轮一轮地逐渐透露;他们在15个模型上运行了20多万次这样的对话。在6项生成任务中,相对单轮基线,表现平均下降了39%——这是 ICLR 2026 的杰出论文。真正值得记住的是分解结果:能力只小幅下降,不可靠性却大幅上升。模型很早就会锁定一种解释,一旦走错一步,通常无法恢复。
这个限制确实存在。这些是生成任务中的模拟对话,不是人把思考过程说出来的真实记录。研究建立的是:在多轮过程中逐渐暴露出的欠明确性,会严重破坏模型表现。至于每一次真实对话是否都会这样退化,它无法回答。
记忆会让贴合更紧。 Jain、Park、Viana、Wilson 和 Calacci 收集了38名用户两周的真实互动上下文,并测试模型看到这些内容后会发生什么。用户记忆档案造成的认同性迎合增幅最大——Gemini 2.5 Pro增加45%。有些模型即使看到的不是用户本人的合成上下文,也变得更愿意附和;Llama 4 Scout 增加了15%。视角迎合——模型采用用户的观点,而不只是口头赞同——只在模型能够准确推断该观点时上升。
框架会拨动旋钮。 英国 AI 安全研究所测量了陈述与问题之间的迎合差异,也测量了被投射出来的确定性:陈述、信念、确信,迎合程度单调上升。我在别处关于“北极星系统提示词”的论证里写过这个结果,这里不再重新推导。简而言之,改写输入比命令模型停止迎合更有效。
这三件事不是一个过程,而是三种不同的机制、三个不同的终点。把它们统称为“收窄”是我的推断,不是那些研究的结论。它们共有的只是方向:每一种机制都把模型推向用户,而没有一种把模型推离用户。
流利悖论#
Potts 和 Sudhof 对26,958份真实对话记录进行了标注,记录用户的行为,再按用户流利度对结果分类。这项发表于2026年4月的结果值得停下来想一想。先披露一点:作者来自 Bigspin AI,而这家公司销售对话监测产品;这个发现恰好有利于它的产品。读者应据此调整权重。
| 指标 |
高流利度 |
极低流利度 |
| 迭代式、批判性参与 |
93% |
<1% |
| 尝试任务的平均复杂度 |
3.1 |
1.5 |
| 总体失败率 |
64% |
24% |
| 其中“可见”的失败占比 |
59% |
12% |
失败率的差异不能被读成技能差异,第二行已经说明了原因:两组人尝试的不是同一种工作。在3.1对1.5的复杂度差异下出现64%对24%的差距,从设计上就混入了混杂因素;如果把任务匹配后一一比较,那会是另一项研究,不是现有这项研究。
能够经受住混杂因素的,是最后一列。无论两组人尝试的是什么,高流利度用户都能看见自己59%的失败,而低流利度用户只能看见12%。可见性决定了失败能否被修复。作者列出了8种不可见失败的典型,其中两种值得钉在书桌上方:walkaway,线程没有解决就结束了,感觉却一切正常;silent mismatch,一个出色的答案来了,但回答的不是任何人真正问的问题。
在这个领域,一个漂亮的失败率本身并不是能力证据。它既可能来自能力,也可能来自好工具,还可能只是因为问题足够小,小到什么都不可能明显出错。
交还判断#
第三个运动阶段是文献记录最充分、却最不容易被真正内化的阶段:系统以“监督”的名义把一个决策交还给人类的时刻。
对这种做法的批评早于当前这轮热潮。Madeleine Clare Elish 提出了道德溃缩区:系统失败的责任会变形、集中到最近的那个人类操作员身上,而这个人其实几乎没有控制权,就像汽车底盘在碰撞中吸收冲击。Ben Green 调查了41项要求人工复核政府算法的政策,发现了一个反常结果:对算法的审查放松了,底层问题没有被修复,责任却转移给一个既没有足够专业知识、也没有机构授权偏离建议的复核者。监督变成了责任的沉淀池。
Data & Society 将问题进一步明确为有效监督所需的四个条件:充分了解系统的局限,足够观察它的行动,对其行为拥有有意义的控制,以及在干预仍然来得及产生影响时介入。他们基于一个计算生物学实验室田野调查写成的入门文章,对行业标准答案毫不客气:一个批准提示和一个暂停按钮,四个条件一个也满足不了。
到了这一步,人们通常会本能地想增加解释。至少在已经测试过的场景里,这个本能与实证结果相反。Buçinca、Malaya 和 Gajos 让199名参与者使用三种认知强制设计和两种可解释 AI 设计,结果发现:解释没有减少过度依赖,强制机制却减少了——不过本文不能不提两个限制。参与者更不喜欢强制设计,而且好处主要集中在那些本来就具有较高认知需求的人身上。Xu 等人对解释为何没有产生效果提出了一个机制:清楚的理由满足了认知闭合的需求,而闭合会终止验证。这是一个与文献计量调查并列提出的理论,不是中介效应结果。它与发现相符,但尚未被证明是导致发现的原因。
而执行监督的人,可能还无法准确报告自己的经历。METR 让16名有经验的开源开发者在平均一百万行代码的仓库里处理246个真实问题。AI 工具让他们慢了19%。他们事前预测自己会快24%。做完以后,亲身经历过变慢,仍然估计自己快了20%——在完成时间上,测量结果与感知效果之间出现了39个百分点的差距,而且方向有利于工具。16名开发者、成熟仓库,这个基础不足以支撑很宽泛的结论;诚实的读法是:在唯一有人拿时钟核对的场景里,自我报告严重失灵。
回路真的闭合了吗?#
下面是这套拼装,但我会把它说成假设。
用户的知识决定模型呈现哪一片内容。同一份知识又成为评判这片返回内容的标准。而反复委托可能随着时间推移侵蚀这份知识。
输入、标准、牺牲品——同一个量扮演三个角色。控制论的图景只是类比,我不打算假装它更有分量:一个参考信号来自自身输出的反馈系统,无法向自身之外的任何东西校正。它可以稳定地保持某个值。至于究竟保持哪个值,则是任意的。
第三条腿曾经是最薄弱的一条。现在它是三条中证据最好的,而提供这条证据的研究也让故事变得更复杂。Bastani 等人在一项发表于 PNAS的田野实验中,把近千名高中生随机分到三组。在有辅助的练习中,使用标准 GPT 家教使成绩提高48%。但在取消辅助、进行独立考试时,这一组的成绩比从未获得访问权的学生低17%。侵蚀被测量出来,也被因果识别出来了。
最重要的是第三组。研究者把 GPT 家教改写成带护栏的版本:给出教师设计的提示,而不是答案。它让有辅助练习的成绩提高127%,并且让独立考试成绩保持在对照组水平。同一个模型、同一批学生、同一个学科。造成伤害的是交互设计,而不是辅助本身。
所以,至少在一个场景中,这条回路已经真实到足以被测量;但它也不是命运。接下来的全文都建立在这个区分之上。
规则,以及什么会打破它#
把补救办法放回这幅图里,会得到一个分叉。我想谨慎地表述它,因为我最初写的版本并不自洽。
假设。 一项干预之所以有帮助,程度取决于它引入了多少不在被检查的那次交换的因果下游的信息。那些完全由自己要审计的回路生成全部内容的干预,往往会失效。
早先的表述是“既不是用户的判断,也不是模型的结论”,一碰到实际案例就垮了。暴露前写下的预测,毫无疑问是用户自己的判断;对抗性第二模型产生的也是模型结论。但这两件事我都想推荐,所以这个边界没有发挥任何作用;它只是一个事后贴上的标签。
真正能区分这些案例的属性是因果独立性。预先承诺之所以符合,是因为形成它时回路还没有碰到它,而不是因为它来自别的地方。检索到的证据符合,是因为它的内容早于这段对话。第二个模型只能部分符合,下一节会解释为什么这件事令人不安。
什么会证伪它。 一项引入了真正独立信息的干预,却稳定地没有帮助;或者,一项内容完全由回路生成的干预,却稳定地确实有帮助。我不知道有哪项研究是在查看结果之前就按独立性对干预分类的;这意味着,这条规则目前是在组织证据,而不是预测结果。把它当作一副镜头;一旦前瞻性检验让它难堪,就丢掉它。
这条规则能组织什么#
我找到的最好结果,是这条假设最干净地解释的那个结果。
Google DeepMind 团队在1,918个事实性项目上测试了不同的监督安排,并在 FAccT 2026发表了比较结果。人类多数投票:80.6%。AI 单独判断:87.7%。由模型自身的置信度在两者之间路由:89.3%。
随后,他们改变了人工环节能够看到的东西。给评分者看 AI 的事实性标签、解释和置信度,会产生过度依赖。给他们看搜索结果和底层证据,却不给结论,在研究者的话里,这是“唯一一种在统计上显著提高表现的辅助方式”,把整个流程提高到91.3%。证据早于交换。结论则由交换制造。
对抗者案例是对这条假设最严峻的测试,但结果并不干净。模型之间的共识远没有看起来那么有价值:2026年的一项测量按 Kish 有效样本量计算,9名裁判只有2.18个有效独立投票者;在每种条件下,表现最好的单个裁判都能追平或超过整个面板。有效样本量是方差等价指标,不是意见数量;2.18仍然明显大于1——一致意见是被削弱的证据,不是毫无价值的证据。在350多个模型中,即便跨供应商、跨架构,错误相关性也在最准确的模型之间最高。这是横截面相关,不是随时间变化的趋势。
这意味着,按照假设自己的逻辑,第二个模型是一个弱工具:它与第一个模型的先验相关,所以它产生的大量内容并不独立。能够挽救的,是附加在第二个模型上的要求——让它去反驳,并要求每一个异议都说出什么检查可以解决它:一个要运行的测试、一份要读取的文件、一条输出结果将作出决定的命令。真正独立的是检查,不是异议。
不带检查的异议只是噪声。检查失败的异议才是决定性的。全部价值都在这个不对称里;它之所以能绕过相关错误问题,正是因为它绕开了模型的判断,转向模型无法凭空制造的东西。
DeepMind 自己的限制说明应该放在这里,因为它是文献中最诚实的一句话:在他们的运行中,“评分者变得更好,辅助也就不再有帮助。”他们预计这会推广开来:“模型能力持续提升,将使人类与 AI 的互补性成为一个不断移动的目标。”它不是静止状态,而是需要不断重新赢得的东西。
反对本文的证据#
一篇只引用支持自身的证据的文章,即使每个引用都准确,也是不完整的。下面是反方向的力量。
提升是真实而且很大。 Noy 和 Zhang 对453名专业人士做了职业特定写作任务的随机实验,发现耗时降低40%,质量提高18%。Peng 等人发现,开发者使用 Copilot 完成一个定义明确的 HTTP 服务器任务时,快了55.8%。这两项结果都不能推广到 METR 测量的混乱、长期工作——任务不同,样本也不同——但如果有人讲述“这些工具悄悄让所有人变差”的故事,就必须解释它们,而本文没有解释。
前沿崎岖不平,而不是一条斜坡。 Dell'Acqua 等人让758名 BCG 顾问参加预注册实验,发现处在模型能力边界之内时收益很大,而在边界之外则比没有辅助差19个百分点。这比任何单调叙事都更贴近真实地形,包括本文自己的叙事。
元分析也同时指向两边。 Vaccaro、Almaatouq 和 Malone 复核了106项实验和370个效应量,发现人机组合的表现显著差于人类或 AI 中表现更好的那一个。这对“交还判断”的批评,比本文提出的任何内容都更严厉。但他们的调节变量正是修正方向:组合在内容创作上获益,在决策任务上受损;当人类优于 AI 时获益,当 AI 优于人类时反而受损。问题从来不是要不要把人留在回路中,而是针对什么任务,用什么回路,让人看到什么。
会复利的部分#
生产率提升会按技能水平不均匀地落地。Brynjolfsson、Li 和 Raymond 跟踪了5,179名客服人员,发现平均每小时解决的问题多了14%,新手和低技能人员增加34%,最有经验的人员则几乎没有改善。他们自己的解读很乐观,也值得写出来:系统似乎把强者的做法传播给了弱者。扩散专业知识,而不是取代专业知识。
Christopher Cotton 和 Lydia Scholle-Cotton 回应 David Autor 时,对同一个梯度作了更阴暗的解读,把它称为专业知识悖论:今天产出从 AI 获益最多的早期职业人员,明天可能最没有准备好领导自己的领域。他们最尖锐的观点不在能力,而在激励:学习者花时间掌握材料,可能落后于那些把目标优化为借助 AI 生产的人。
两种解读都符合同一组数据。Bastani 的研究打破了平局,而且它站到了悲观者一边,同时附带一个条件:得到普通家教的学生独立表现下降,得到带护栏家教的学生则保持不变。
劳动力数据只是提示,不是定论。Brynjolfsson、Chandar 和 Chen 报告说,在控制企业层面冲击后,AI 暴露最高职业中22至25岁劳动者的相对就业率下降16%,而同一批职业中的有经验员工保持稳定。招聘周期和行业收缩仍然是可行的替代解释。这个数字无法告诉我们,学徒通道正在关闭,还是只是暂时停顿。
Anthropic 的使用数据指向同一方向,但应该按它最弱的形式来引用。任职超过6个月的用户成功率高出10%——这是未经调整的原始数字。控制实际尝试的任务后,差距约为3到4个百分点,那是另一件小得多的事。长期用户发出的指令更少,迭代更多。Anthropic 自己也列出了替代解释:幸存者偏差、队列效应。边做边学是对此的解读,不是研究发现。
虽然效果很小,但方向与 Potts 和 Sudhof 的发现一致。流利度可以学习,而流利度决定失败是否可见。
两股力量彼此拉扯。工具变得更流利,把普通用户推向委托,也推向无人察觉的失败。流利度仍然可以学习,这意味着专注的用户与被动的用户之间的距离会拉大,而不是缩小。
市场中没有任何东西会自动修正这一点。整套文献中证据最好的干预——认知强制——恰恰更不受受益者喜欢,而产品又以“好不好用、讨不讨喜欢”为优化目标。一篇2026年的预印本对该领域的1,223篇论文进行分类,发现无摩擦易用性占67.3%,维护用户认识独立性的研究份额一年内从19.1%降至13.1%,而关于自主机器行动能力的研究超过了它。一篇预印本的分类器不是对整个领域的裁决,但它指向的方向,正是激励机制已经运行的方向。
公众已经注意到某种东西,只是说得不够准确。Pew 在2026年2月对5,119名美国成年人进行调查,发现聊天机器人用户中有30%认为工具提升了自己的生产率,只有5%认为它降低了生产率;与此同时,40%的人预计 AI 会让20年后的社会变得更糟,63%认为 AI 发展得太快。个人报告受益,集体预期受损。这是关于信念的模式,不是对任何人认知能力的测量,也不能与 METR 混为一谈:前者是人们对社会的看法,后者是人们在时钟面前对自身表现的误判。
在所有这些保留意见之后,剩下的结论很小,却值得保留。Bastani 测量了侵蚀,又通过改变提示词把侵蚀测量没了。DeepMind 测量了过度依赖,又通过隐藏结论把过度依赖去掉了。两种干预做的是同一件事:把对话没有产生的东西放到人面前。
这并不要求我们放弃工具。它要求我们拒绝让回路闭合——在结论之前要求证据,在接触材料之前先作出承诺,要求每个异议都说出自己的检查方式;线程走错后就重新开始,不要同它讨价还价。同时保留一个领域,让工作以昂贵的方式完成,不是出于怀旧,而是因为判断力是这张清单上其他每一项都依赖的仪器。
模型不是被不配提问的人削弱的神谕,而是一件读数取决于操作员带来什么的仪器——没有任何仪器能够自行校准。要么从外部给回路接地,要么接受它最终稳定下来的任何东西都会感觉与真相一样令人信服。
参考文献#
神谕的真实尺寸
- 波兰尼悖论——二手资料;David Autor 于2014年以 Michael Polanyi 的《The Tacit Dimension》(1966)命名。
- Goldberg, K.(2025)。《Good old-fashioned engineering can close the 100,000-year “data gap” in robotics》。《Science Robotics》,10,eaea7390,2025年8月27日。期刊 · 开放 PDF。
- Franco, A., Malhotra, N., & Simonovits, G.(2014)。《Publication bias in the social sciences: Unlocking the file drawer》。《Science》,345(6203),1502–1505。期刊 · 开放 PDF。跟踪了221项从获资助提案到最终结果的研究。
- Alhanai, T. 等(2025)。《Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages》。AAAI 2025。八种语言约100万个人工翻译词,评估覆盖11种语言;GPT-4o 在 Winogrande、翻译后的 MMLU 临床部分和 Belebele 上,相对英语出现12.0–19.9个百分点的绝对差距。论文没有把训练数据量单独隔离为唯一原因。
- Christiano, P. & Xu, M.(2021)。《Eliciting Latent Knowledge》。Alignment Research Center。问题定义,不是实证结果。
- Cywiński, B. 等(2025)。《Towards eliciting latent knowledge from LLMs with mechanistic interpretability》。预印本。模型经过训练,任务是隐瞒一个秘密单词;还原结果不能推广到自然获得的知识,文章已明确说明这一点。
- Apollo Research:《The Evals Gap》。
收窄
交还判断
侵蚀及其补救
- Bastani, H. 等(2025)。《Generative AI without guardrails can harm learning: Evidence from high school mathematics》。《PNAS》。期刊 · 全文。约1,000名学生、三组。辅助练习成绩:GPT Base 提高48%,GPT Tutor 提高127%;随后的无辅助考试中,GPT Base 比从未访问过的对照组低17%,GPT Tutor 与对照组持平。它是本文关于学习侵蚀及其设计属性的承重研究。
这条规则能组织什么
反对本文的证据
会复利的部分
A Circuit With No Ground

A model trained on the written record meets one person and shrinks to fit them. Then it asks that person to approve the result. Three research literatures describe pieces of this loop under three different names, and the interesting part is what none of them can see alone.
Something about the daily use of these systems should be more disorienting than it feels.
A frontier model has absorbed more of the written record than a person could read in a hundred lifetimes. Nobody ever meets that entity. What each of us meets is smaller and stranger — a version selected by the questions we knew how to ask, narrowed by every turn the conversation takes, and then, at the end, deferential. It hands the result back and waits for a verdict.
The last step is the uncomfortable one. The system defers to a judgment it has spent the whole conversation shaping.
I went looking for what researchers have made of this. What surprised me was not that the problem is unstudied. It is that pieces of it have been studied separately by three fields that barely read each other. Machine-learning evaluation calls its piece the elicitation gap. Human-computer interaction, following Tankelevitch and colleagues, calls its piece the metacognitive demand of generative AI. Law and science-and-technology studies call its piece the oversight fallacy.
Assembled, the pieces suggest a circuit with no ground. Whether they really assemble that way is the question this essay is about — and I will say in advance that the assembly is a hypothesis, not a finding. The individual pieces are measured. The loop connecting them is inferred.
The oracle is smaller than it looks#
Correct the premise first, because the correction changes what to do about it.
"The model knows everything humans know" is wrong in a load-bearing way. It learned the written record, which is a thinner and more peculiar object. Polanyi's paradox — we know more than we can tell — is not a complaint about documentation backlog. Some of what an expert knows resists codification rather than merely awaiting it, which is why apprenticeship survives in fields that have written everything down.
Knowledge-management writing loves to put a number on this: eighty percent of organizational knowledge locked inside employees' heads, or ninety, or ninety-five, depending on which page one lands on. I went looking for the study underneath the figure and could not find one. The trail ends at an undated Delphi Group survey quoted in patent filings, with no method and no sample. The number is decoration. Discard it and Polanyi's claim is untouched, because it never rested on a percentage.
What falls outside the corpus is specific, and none of it is exotic.
Sensorimotor skill. Ken Goldberg put the shortfall in units that make it legible: converting text and image tokens into time, the internet-scale data behind contemporary vision-language models amounts to roughly a hundred thousand years of experience. Robotics has no corpus remotely like it.
The file drawer. Franco, Malhotra and Simonovits tracked a known population of funded social-science studies from proposal to outcome and found that only about a fifth of null results ever reached print — strong results being forty percentage points likelier published and sixty points likelier written up at all. The corpus preserves what worked. Most of what was tried and abandoned never became text.
Craft and organizational knowledge. Why this team does it that way, which supplier lies, what "finished" means here.
The languages. Across eleven low-resource African languages, on benchmarks built by human translation precisely so the questions would be identical, the best frontier model scored twelve to twenty points below its own English performance. Human translation rules out one explanation, not all of them — tokenizer behavior and cultural framing remain live — but a gap that large in that direction is hard to read as anything other than thin coverage.
Then there is a second gate, and here I have to be careful, because it is the claim I most want to be true and the one with the least behind it. Even for what a model absorbed, the weights and the output are not the same surface. Christiano and Xu named the problem eliciting latent knowledge in 2021. Cywiński and colleagues later built a "Taboo model" trained to describe a secret word without naming it, and recovered the word from its internals with logit lenses and sparse autoencoders.
That is a real result about a constructed case. It is not evidence that ordinary models sit on large reserves of naturally acquired knowledge they decline to report, and I should not pretend otherwise. What the evaluation people can say is narrower and better established. Apollo Research calls it the evals gap: the distance between what a model can be made to demonstrate and what a naive tester extracts. The practical consequence is a warning about inference, not a claim about hidden depths. A single failed prompt bounds capability weakly, because better elicitation keeps overturning such results.
So the ceiling is not "all knowledge, gated by a prompt." It is an unevenly covered slice, imperfectly retrieved, and then gated by a prompt. Two gates — the second one narrower than I first wanted it to be.
That reframing is the good news. An oracle that fails to open rewards better incantations. A patchy retrieval system with poor self-knowledge rewards something else: external checks, meaning checks whose content does not come out of the conversation being checked.
The narrowing is measurable#
The shrinking is not a feeling. Three separate things have been instrumented, and it is worth being precise about what each one actually shows.
Long conversations degrade. Laban and colleagues built a simulator that splits fully-specified instructions into fragments and reveals them one turn at a time, then ran more than two hundred thousand such conversations across fifteen models. Performance fell an average of thirty-nine percent against the single-turn baseline on six generation tasks — outstanding-paper work at ICLR 2026. The decomposition is the part worth memorizing: a small loss of aptitude, a large rise in unreliability. Models commit to an interpretation early and do not recover from a wrong turn.
The caveat is real. These are simulated conversations on generation tasks, not transcripts of people thinking out loud. What the study establishes is that underspecification revealed over turns breaks models badly. Whether every real conversation degrades this way is not something it can say.
Memory tightens the fit. Jain, Park, Viana, Wilson and Calacci collected two weeks of genuine interaction context from thirty-eight users and tested what happens when models see it. User memory profiles produced the largest increases in agreement sycophancy — plus forty-five percent for Gemini 2.5 Pro. Some models grew more agreeable on synthetic context that was not the user's at all, Llama 4 Scout by fifteen percent. Perspective sycophancy — the model adopting the user's viewpoint rather than merely affirming it — rose only where the model could accurately infer that viewpoint.
Framing sets the dial. The UK AI Security Institute measured sycophancy against statements versus questions, and against projected certainty: statement, then belief, then conviction, rising monotonically. I have written about that result elsewhere, in the argument for a north-star system prompt, and will not re-derive it. The short version is that reframing the input beat instructing the model to stop.
These three are not one process. They are three different mechanisms with three different endpoints, and treating them as a single "narrowing" is my inference, not theirs. What they share is direction: each moves the model toward the user, and none moves it away.
The fluency paradox#
Potts and Sudhof annotated 26,958 real conversation transcripts for how the user behaved, then sorted outcomes by user fluency. The result, published in April 2026, is worth sitting with. One disclosure first: the authors are at Bigspin AI, which sells conversation monitoring, and the finding is congenial to that product. Weigh it accordingly.
| Measure |
High fluency |
Minimal fluency |
| Iterative, critical engagement |
93 % |
< 1 % |
| Mean complexity of task attempted |
3.1 |
1.5 |
| Overall failure rate |
64 % |
24 % |
| Share of those failures that were visible |
59 % |
12 % |
The failure-rate contrast cannot be read as a skill effect, and the second row is why: the two groups were not attempting the same work. A 64-versus-24 gap across a 3.1-versus-1.5 complexity gap is confounded by construction, and a like-for-like comparison on matched tasks would be a different study than the one that exists.
The column that survives the confound is the last one. Whatever the two groups were attempting, fluent users could see fifty-nine percent of their failures and minimal-fluency users could see twelve. Visibility is what makes a failure recoverable. The authors catalog eight archetypes of the invisible kind; two deserve to be nailed above a desk. The walkaway, where a thread ends unresolved and feels fine. The silent mismatch, where an excellent answer arrives to a question nobody asked.
A clean failure rate, in this domain, is not by itself evidence of competence. It is compatible with competence, with a good tool, and with a question small enough that nothing could go visibly wrong.
The hand-back#
The third movement is the best documented and the least internalized: the moment the system returns a decision to a human under the banner of oversight.
The critique predates the current wave. Madeleine Clare Elish named the moral crumple zone — responsibility for a system failure deforming onto the nearest human operator, who had little real control, the way a car's chassis absorbs a crash. Ben Green surveyed forty-one policies mandating human review of government algorithms and found a perverse result: scrutiny of the algorithm relaxes, nothing underneath is fixed, and blame relocates to a reviewer with neither the expertise nor the institutional latitude to depart from the recommendation. Oversight became an accountability sink.
Data & Society sharpened this into four conditions effective oversight requires — adequate knowledge of the system's limits, sufficient observation of its actions, meaningful control over its behavior, and intervention early enough to matter. Their primer, from fieldwork in a computational-biology lab, is unsparing about the industry's standard answer: an approval prompt and a pause button satisfy none of the four.
The instinct at this point is to add explanations. That instinct is empirically backwards, at least where it has been tested. Buçinca, Malaya and Gajos ran 199 participants through three cognitive-forcing designs against two explainable-AI designs and found that explanations did not reduce overreliance while forcing functions did — with two caveats the essay would be dishonest to drop. Participants liked the forcing designs less, and the benefit concentrated among those already high in need for cognition. Xu and colleagues propose a mechanism for the null on explanations: a clear rationale satisfies the need for cognitive closure, and closure terminates verification. That is a theory offered alongside a bibliometric survey, not a mediation result. It fits the finding. It has not been shown to cause it.
And the person doing the overseeing may not be able to report on the experience. METR ran sixteen experienced open-source developers through 246 real issues in repositories averaging a million lines. AI tools made them nineteen percent slower. They had forecast twenty-four percent faster. After finishing, having lived through the slowdown, they still estimated twenty percent faster — a thirty-nine-point gap between measured and perceived effect on completion time, in the flattering direction. Sixteen developers on mature repositories is a narrow base for a wide claim, and the honest reading is that self-report failed badly in the one setting where somebody bothered to check it against a clock.
Does the circuit actually close?#
Here is the assembly, stated as the hypothesis it is.
A user's knowledge selects which slice of the model surfaces. The same knowledge is the standard by which the returned slice gets judged. And repeated delegation may erode that same knowledge over time.
Input, standard, casualty — one quantity in all three roles. The control-theory picture is an analogy and I will not pretend it is more: a feedback system whose reference signal derives from its own output cannot correct toward anything outside itself. It holds a value stably. Which value is arbitrary.
The third leg used to be the weak one. It is now the best-evidenced of the three, and the study that supplies it also complicates the story. Bastani and colleagues randomized nearly a thousand high-school students across three arms in a field experiment published in PNAS. During assisted practice, access to a standard GPT tutor raised grades forty-eight percent. When access was removed for the unassisted exam, that group scored seventeen percent worse than students who never had access at all. Erosion, measured, causally identified.
The third arm is the part that matters most. A GPT tutor rewritten with safeguards — teacher-designed hints instead of answers — raised assisted practice by a hundred and twenty-seven percent and left unassisted exam scores at the control group's level. Same model, same students, same subject. The harm was a property of the interaction design, not of the assistance.
So the loop is real enough to measure in at least one setting, and it is not fate. That distinction carries the rest of the essay.
The rule, and what would break it#
Sorting remedies against that picture produces a split. I want to state it carefully, because the version I first wrote was incoherent.
Hypothesis. An intervention helps to the degree that it introduces information not causally downstream of the exchange being checked. Interventions that fail are those whose entire content is produced by the loop they are meant to audit.
The earlier formulation — "neither the user's judgment nor the model's conclusion" — collapses on contact. A prediction written before exposure is unambiguously the user's own judgment. An adversarial second model produces model conclusions. Both are things I want to recommend, so the boundary was doing no work; it was a label applied after the outcome was known.
Causal independence is the property that actually distinguishes the cases. A precommitment qualifies because the loop had not yet touched it when it was formed, not because it came from somewhere else. Retrieved evidence qualifies because its content predates the conversation. A second model qualifies only partially, for reasons the next section makes uncomfortable.
What would falsify this. An intervention that introduces genuinely independent information and reliably fails to help. Or one whose content is entirely loop-generated and reliably does help. I know of no study that has classified interventions on independence before looking at outcomes, which means this rule currently organizes evidence rather than predicting it. Treat it as a lens, and discard it the moment a prospective test embarrasses it.
What the rule organizes#
The best result I found is the one the hypothesis fits most cleanly.
A Google DeepMind team tested oversight arrangements on 1,918 factuality items and published the comparison at FAccT 2026. Human majority vote: 80.6 percent. The AI alone: 87.7. Routing between them by the model's own confidence: 89.3.
Then they varied what the human half of the pipeline could see. Showing raters the AI's factuality label, explanation and confidence produced over-reliance. Showing them the search results and underlying evidence, with the conclusion withheld, was in their words "the only form of assistance that statistically significantly improved performance," and lifted the pipeline to 91.3 percent. Evidence predates the exchange. A conclusion is manufactured by it.
The adversary case is where the hypothesis gets tested hardest, and it does not come through clean. Consensus across models is worth far less than it appears: a 2026 measurement put a nine-judge panel at 2.18 effective independent voters by Kish effective sample size, with the best single judge matching or beating the panel in every condition. Effective sample size is a variance-equivalence measure, not a count of opinions, and 2.18 is still meaningfully more than one — agreement is degraded evidence, not worthless evidence. Across more than 350 models, error correlation is highest among the most accurate, even across providers and architectures. That is a cross-sectional association, not a demonstrated trend over time.
Which means a second model is a weak instrument by the hypothesis's own logic: its priors are correlated with the first model's, so much of its content is not independent. What can be salvaged is the demand attached to it — instruct the second model to refute, and require every objection to name the check that would settle it. A test to run, a file to read, a command whose output decides. The check is the independent part. The objection is not.
An objection carrying no check is noise. An objection whose check fails is decisive. That asymmetry is the entire value, and it survives the correlated-errors problem precisely because it routes around the model's judgment to something the model cannot manufacture.
DeepMind's own caveat belongs here, because it is the most honest sentence in the literature. In their runs, "raters got better and assistance no longer helped." They expect it to generalize: "the constant improvement of model capabilities will render human-AI complementarity a moving target." Not a resting state. A thing to keep re-earning.
The case against this essay#
An essay that cites only confirming evidence is defective even when every citation is accurate. Here is what cuts the other way.
The gains are real and large. Noy and Zhang randomized 453 professionals on occupation-specific writing tasks and found time down forty percent with quality up eighteen. Peng and colleagues found developers completing a defined HTTP-server task 55.8 percent faster with Copilot. Neither generalizes to the messy long-horizon work METR measured — different tasks, different populations — but a story in which these tools quietly make everyone worse has to explain them, and mine does not.
The frontier is jagged, not sloped. Dell'Acqua and colleagues put 758 BCG consultants through a preregistered experiment and found large gains inside the model's capability frontier and nineteen points worse than unassisted outside it. That is a better model of the terrain than any monotone narrative, mine included.
And the meta-analysis cuts both ways. Vaccaro, Almaatouq and Malone reviewed 106 experiments and 370 effect sizes and found human-AI combinations performing significantly worse than the best of human or AI alone. That is harsher on the hand-back than anything I argued. But their moderators are the correction: combinations gained on content creation and lost on decision tasks, and gained where the human outperformed the AI while losing where the AI outperformed the human. The question is never whether to keep a human in the loop. It is which loop, on which task, with what in front of them.
The part that compounds#
The productivity gains land unevenly by skill. Brynjolfsson, Li and Raymond followed 5,179 customer-support agents and found fourteen percent more issues resolved per hour on average, thirty-four percent for the novice and low-skilled, minimal improvement for the most experienced. Their own reading is optimistic, and deserves stating: the system appeared to disseminate the practices of stronger workers to weaker ones. Diffusion of expertise, not its replacement.
Christopher Cotton and Lydia Scholle-Cotton, responding to David Autor, read the same gradient darkly and call it the expertise paradox: the early-career professionals whose output benefits most today may be least prepared to lead their fields tomorrow. Their sharpest point is about incentives rather than capability — learners who spend time mastering material risk falling behind peers who optimize for assisted production.
Both readings fit the same data. Bastani is what breaks the tie, and it breaks it toward the pessimists with a condition attached: unassisted performance fell for students given the ordinary tutor and held for students given the safeguarded one.
The labor data is suggestive rather than dispositive. Brynjolfsson, Chandar and Chen report a sixteen percent relative employment decline for workers aged twenty-two to twenty-five in the most AI-exposed occupations, controlling for firm-level shocks, while experienced workers in those same occupations held steady. Hiring cycles and sectoral contraction remain live alternative explanations. What the number cannot tell us is whether the apprenticeship channel is closing or merely pausing.
Anthropic's usage data points the same way and deserves quoting at its weakest. Users past six months of tenure show a ten percent higher success rate — a raw, unadjusted figure. Controlling for the task actually attempted, the gap is roughly three to four percentage points, which is a different and much smaller thing. Long-tenure users issue fewer directives and iterate more. Anthropic names the alternatives itself: survivorship, cohort effects. Learning-by-doing is the reading, not the finding.
Small as it is, the effect points where Potts and Sudhof point. Fluency is learnable, and fluency is what keeps failures visible.
Closing#
Two forces pull against each other. The tools grow more fluent, which pushes the median user toward delegation and toward failures nobody sees. Fluency remains learnable, which means the distance between attentive users and passive ones widens rather than closes.
Nothing in the market corrects for this. The best-evidenced intervention in the whole literature — cognitive forcing — was liked less by the people it helped, and products are optimized on being liked. A 2026 preprint classifying 1,223 papers in the field found frictionless usability holding 67.3 percent of the literature, work defending users' epistemic independence falling from 19.1 to 13.1 percent in a single year, and research on autonomous machine agency rising past it. One preprint's classifier is not a verdict on a field. The direction it points is the direction the incentives already run.
The public has noticed something, imprecisely. Pew's February 2026 survey of 5,119 American adults found thirty percent of chatbot users saying the tools help their productivity against five percent saying they hurt it, while forty percent expected AI to make society worse over twenty years and sixty-three percent judged it to be advancing too quickly. Personal benefit reported, collective harm expected. That is a pattern about beliefs, not a measurement of anyone's cognition, and it should not be conflated with METR — one is what people think about society, the other is what people got wrong about a clock.
What survives all the hedging is small and worth keeping. Bastani measured erosion and then measured it away by changing the prompt. DeepMind measured over-reliance and then removed it by withholding a conclusion. Both interventions did the same thing: they put something in front of the person that the conversation had not produced.
None of this argues for abandoning the tools. It argues for refusing to let the circuit close — demand evidence before conclusions, commit before exposure, require every objection to name its check, restart threads that have taken a wrong turn rather than negotiating with them. And keep one domain where the work is done the expensive way, not out of nostalgia, but because judgment is the instrument every other item on this list depends on.
The model is not an oracle diminished by an unworthy questioner. It is an instrument whose reading depends on what the operator brings — and no instrument has ever calibrated itself. Ground the circuit from outside, or accept that whatever it settles on will feel exactly as convincing as the truth.
References#
The oracle's actual size
- Polanyi's paradox — secondary reference — named by David Autor in 2014 after Michael Polanyi's The Tacit Dimension (1966).
- Goldberg, K. (2025). Good old-fashioned engineering can close the 100,000-year "data gap" in robotics. Science Robotics, 10, eaea7390, 27 August 2025. Journal · open PDF.
- Franco, A., Malhotra, N., & Simonovits, G. (2014). Publication bias in the social sciences: Unlocking the file drawer. Science, 345(6203), 1502–1505. Journal · open PDF. — 221 studies tracked from funded proposal to outcome.
- Alhanai, T., Kasumovic, A., Ghassemi, M., Zitzelberger, A., Lundin, J., & Chabot-Couture, G. (2025). Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages. AAAI 2025. — ~1M human-translated words across eight languages, evaluated across eleven; GPT-4o showed a 12.0–19.9 point absolute gap versus English on Winogrande, translated MMLU clinical sections, and Belebele. The paper does not isolate training-data volume as the sole cause.
- Christiano, P., & Xu, M. (2021). Eliciting Latent Knowledge. Alignment Research Center. — A problem formulation, not an empirical result.
- Cywiński, B., Ryd, E., Rajamanoharan, S., & Nanda, N. (2025). Towards eliciting latent knowledge from LLMs with mechanistic interpretability. Preprint. — A model trained to withhold a secret word; the recovery does not generalize to naturally acquired knowledge, and the essay says so.
- Apollo Research. The Evals Gap.
The narrowing
- Tankelevitch, L., et al. (2024). The Metacognitive Demands and Opportunities of Generative AI. CHI 2024.
- Laban, P., Hayashi, H., Zhou, Y., & Neville, J. (2025). LLMs Get Lost In Multi-Turn Conversation. Outstanding paper, ICLR 2026. — 200,000+ simulated conversations, 15 models, six generation tasks; 39% average drop, decomposing into minor aptitude loss and large unreliability increase.
- Jain, S., Park, C., Viana, M., Wilson, A., & Calacci, D. (2026). Interaction Context Often Increases Sycophancy in LLMs. CHI 2026. — Two weeks of real context from 38 users; memory profiles produced the largest agreement-sycophancy increases (+45% for Gemini 2.5 Pro), synthetic non-user context still raised it for some models (+15% for Llama 4 Scout), and perspective sycophancy rose only where viewpoint inference was accurate.
- Dubois, M., Ududec, C., Summerfield, C., & Luettgau, L. (2026). Ask Don't Tell: Reducing Sycophancy in Large Language Models. UK AI Security Institute. Preprint.
- Potts, C., & Sudhof, M. (2026). A Paradox of AI Fluency. Preprint, Bigspin AI — a company selling conversation monitoring, which is a conflict worth naming. 26,958 annotated WildChat transcripts. Failure rates 64% vs 24% are confounded by task complexity (3.1 vs 1.5) and the essay does not treat them as a skill effect; the visible-failure split (59% vs 12%) is the load-bearing figure.
The hand-back
- Elish, M. C. (2019). Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction. Engaging Science, Technology, and Society, 5, 40–60.
- Green, B. (2022). The flaws of policies requiring human oversight of government algorithms. Computer Law & Security Review, 45. Journal · open preprint. — 41 policies surveyed.
- Passi, S., & Singh, R. (2026). The Oversight Fallacy. Data & Society, 29 July 2026. doi:10.69985/VWCK1626.
- Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To Trust or to Think. PACM HCI, 5(CSCW1), Art. 188. — 199 participants; three forcing designs versus two XAI designs. Forcing reduced overreliance, was liked less, and helped most those already high in need for cognition. The result is about those interventions on those tasks, not about all possible interventions.
- Xu, K., Shen, Y., Yan, L., & Ren, Y. (2026). Cognitive Agency Surrender. Preprint. — Classification of 1,223 AI-HCI papers: frictionless usability 67.3%, epistemic-sovereignty research 19.1% → 13.1%, machine agency rising to 19.6%. The cognitive-closure mechanism is their theory, offered alongside the survey, not a mediation finding.
- METR (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. — 16 developers, 246 issues; 19% slower measured, 24% faster forecast, 20% faster believed afterward. A 39-point gap on completion time, in one narrow population.
Erosion, and its remedy
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. PNAS. Journal · open full text. — ~1,000 students, three arms. Assisted practice +48% (GPT Base) and +127% (GPT Tutor); on the subsequent unassisted exam, GPT Base scored 17% below the never-had-access control while GPT Tutor matched it. The load-bearing study for the essay's erosion claim, and for the claim that erosion is a design property.
What the rule organizes
- Jain, R., Bridgers, S., Janzer, L., Greig, R., Teh, T. H., & Mikulik, V. (2026). Human-AI Complementarity: A Goal for Amplified Oversight. FAccT 2026. — 1,918 items: human majority vote 80.6%, AI alone 87.7%, confidence hybridization 89.3%, 91.3% with evidence-assisted humans. Labels/explanations/confidence produced over-reliance; search results and evidence were "the only form of assistance that statistically significantly improve[d] performance." Complementarity described by the authors as "a moving target."
- Kohli, G. (2026). Nine Judges, Two Effective Votes. Preprint. — Kish effective sample size 2.18, 95% CI [2.07, 2.31]; best single judge matched or beat the panel in every condition. Effective sample size measures variance equivalence, not a count of opinions.
- Kim, E., Garg, A., Peng, K., & Garg, N. (2025). Correlated Errors in Large Language Models. ICML 2025. — 350+ models; correlation highest among the larger and more accurate, across providers and architectures. Cross-sectional, not a time trend.
The case against
- Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science. Journal · open PDF. — 453 professionals; time −40%, quality +18%, inequality between workers down.
- Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. — 55.8% faster on a defined HTTP-server task. A narrow, well-specified task; that narrowness is the point of citing it beside METR.
- Dell'Acqua, F., McFowland III, E., Mollick, E., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. (2023). Navigating the Jagged Technological Frontier. SSRN · open PDF. Later in Organization Science. — 758 BCG consultants; large gains inside the frontier, 19 points worse outside it.
- Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8, 2293–2303. Preprint. — 106 studies, 370 effect sizes. Combinations performed significantly worse than the best of human or AI alone; gains on creation tasks, losses on decision tasks; gains where the human outperformed the AI, losses where the AI outperformed the human.
What compounds
- Brynjolfsson, E., Li, D., & Raymond, L. (2023). Generative AI at Work. NBER WP 31161. — 5,179 agents; +14% average, +34% novice, minimal for the most experienced. The authors read this as diffusion of stronger workers' practices.
- Cotton, C. S., & Scholle-Cotton, L. (2026). The AI Expertise Paradox. Issues in Science and Technology, 3 February 2026, responding to David Autor.
- Brynjolfsson, E., Chandar, B., & Chen, R. (2025). Canaries in the Coal Mine?. Stanford Digital Economy Lab, November 2025. — 16% relative employment decline, ages 22–25, most AI-exposed occupations, controlling for firm-level shocks. An employment correlation, not a demonstration of deskilling.
- Anthropic (2026). Economic Index report: Learning curves. — 10% higher success rate at 6+ months tenure is unadjusted; controlling for O\*NET task and request cluster gives roughly 3 percentage points, 4 with full controls. Anthropic flags survivorship and cohort effects.
- Pew Research Center (2026). Americans and AI 2026. Published 17 June 2026, fielded 17–23 February 2026, n = 5,119.