文学能力是 AI 安全基础设施

作者:李笑来 · 来源:lixiaolai.com · 发布于 2026-04-30 · 原文链接

回复在四十秒后到了——三个利落的段落,不打太极,不留省略号。一位在职专业人士请 AI 助手解释某部小说中某种叙事手法的主题意义,而模型答得像一位研讨课主持人那样流畅:一个关于不可靠叙述者的论断,两处支撑细节,一句关于作者意图的收尾观察。每个句子都能解析。词汇都对。这个论证有着论证该有的形状。

那个自信的错误答案#

回复在四十秒后到了——三个利落的段落,不打太极,不留省略号。一位在职专业人士请 AI 助手解释某部小说中某种叙事手法的主题意义,而模型答得像一位研讨课主持人那样流畅:一个关于不可靠叙述者的论断,两处支撑细节,一句关于作者意图的收尾观察。每个句子都能解析。词汇都对。这个论证有着论证该有的形状。

然而,有什么地方不对劲。这位专业人士——一个有经验的 AI 使用者,一个学会了好好提问、经常核查的人——感觉到了。不是事实错误:人物的名字是对的,情节点也描述得没错。更像是两个想法接合处露出的一道缝。第二段的核心论断,并不完全从第一段推得出来。那个“因此”是承重的,而承重的东西并不在那儿。结论来得太干净了,把几条其实并没有真正编在一起的线捆到了一块。

这位专业人士又读了一遍。那份不安还在。但没有任何程序可以用来缩小它的范围——没有实验可做,没有引文可查,没有反面来源可参照。那份不适是真实的;用来定位其来源的仪器却不在。停顿片刻之后,这个答案被接受了,被转发了,被引进了一份将会送到其他人手上的报告里。

问题不在于 AI 的自信。问题在于评估者的沉默。

接收一个 AI 输出与评估它之间存在一道缺口——而这道缺口,并不会因为你更了解 AI 如何运作就被填上。

共识说对了什么#

如果“把 AI 素养当成一项技术技能”的主张只建立在直觉之上,本文会好写得多。事实并非如此。

哈佛的 Kestin 和同事做过一项随机对照试验,把 194 名本科物理学生分配到定制 AI 导师或主动学习教学两组。1 AI 辅导组的学习增益,比主动学习对照组高出 0.63 到 1.3 个标准差——而且只用了两次课。学生报告了更高的投入度和更强的动机。这项研究经过同行评议,发表于《Scientific Reports》,并且通过使用经准确性审核的预写专家答案来防止幻觉。这些不是会让结果虚高的条件;这些是在相当严格的控制下检验一个真实主张的条件。

Kestin 那项研究是这个领域里最好的实验。一个反对 AI 辅导却绕开这项证据的论证,是没有把功课做完的。诚实的交锋要从承认它所展示的东西开始:设计得当的 AI 辅导,在结构化领域中能高效地传递内容知识。效应量并不轻微。设计是可信的。1

支撑这项证据的机构共识也相当扎实。世界经济论坛 2025 年的《AI 素养框架》把核心能力界定为算法思维、提示词工程、理解 AI 偏见、数据素养,以及对 AI 输出的批判性思考。2 普渡大学在 2025 年 12 月把“AI 工作能力”定为毕业要求,是第一所这样做的主要研究型大学。2 这些框架反映的是真正的政策审议,而不是恐慌。它们代表着那些花了数十年思考“学生需要知道什么”的机构,经过斟酌后的立场。

成本不对称的论证进一步加强了这个主张。Khanmigo 的学生用户在一年之内从约四万增长到七十万。3 一对一的 AI 辅导,对一个没有资源的学生免费可得;而一对一的人类辅导每小时要 50 到 200 美元——支持 AI 工具的公平性论证是真实的,而且这个成本差距并不轻微。

不过,恰恰在这里,Kestin 那几位作者自己写下的话很要紧。在局限性部分,他们指出:“我们并不假定结构化的 AI 辅导在所有情境下都会胜过课堂主动学习,例如那些需要对多个概念做复杂综合、以及需要更高阶批判性思考的情境。”1 这项研究检验的是理解、应用和分析——大致相当于布卢姆分类法的第一到第四层。作者自己划明了它范围之外的东西。

问题不在于 AI 辅导在内容传递上管不管用。它管用。问题在于,内容传递是不是那个卡住脖子的约束。

三项侦测任务#

当有人说“我需要能够评估 AI 输出”时,他描述的其实至少是三个不同的问题。多数关于 AI 素养的讨论都把这一点略过了。这种混同很要紧,因为每个问题所要求的技能确实不同——而其中最难的那一个,得到的关注最少。

第一项任务是文体作者归属侦测​:给你一段文本,它是不是 AI 写的?这是一个模式匹配问题。要问的是,这段文字是否带有 AI 生成的可辨文体签名——那种特有的节奏平坦感,对某些连接短语的偏好,那种既自信又无语气的语域。这项任务已经被直接研究过。

第二项任务是事实捏造侦测​:这条引文存在吗,这个统计数字能追溯到一项真实研究吗,这个历史事件与史料记载相符吗?这是一项核查任务。它所要求的技能是知道怎么去查​——数据库访问、交叉比对、来源追溯。在方法上,它更接近事实核查新闻,而不是文学分析。

第三项任务是推理连贯性评估​:这个论证真的立得住吗?结论真的从前提推得出来吗?这个论断与两段之前所断言的东西在内部一致吗?断言的自信程度,与支撑的质量相称吗?这是另一种活儿——不是模式匹配,不是核查,而是对论证结构本身的评估。

牛津的 Farquhar 和同事发表了目前针对第三项任务最精巧的自动化方法:语义熵——让模型生成多个候选答案,再检查它们是否聚拢在一个一致的含义周围,以此侦测虚构。4 这个方法很优雅。它同时也不可避免地是一项模型侧的操作:它需要访问模型内部的概率分布,也需要多次生成。而一个从聊天界面收到单一自信输出的人,这些一样也没有。

更重要的是,作者点名了一类这个方法够不着的失败:他们称之为“结构性错误”——那些在模型全部输出中都系统而一致地出现的错误,因为它们源自模型在训练中学到的东西。4 而这恰恰是对终端用户最危险的那类失败,因为它们正是模型产出时最自信、最少自我纠正的那些。语义熵侦测不到它们。在自动侦测的天花板之上,剩下的是人的判断。

一项针对这条边界的对照研究——德国的 PMC11914838 实验——很有教益。十三位人文学者和二十二位医学专业人士,试图从一组德语医学生论文中辨认出哪些是 AI 生成的。5 医学专业人士的准确率是 72%;人文学者是 65%。这个差异在统计上无法区分(OR 1.37,95% CI 0.5–3.9)。而那些人文学者,按他们的文学判断,把 ChatGPT 的文字评为在语言上质量更高​——不只是相当,而是更优。5

这项研究发现了它所发现的东西,而值得弄清楚的是它究竟检验了什么:第一项任务。在一个非文学领域、用德语、针对医学生论文文字所做的文体作者归属侦测。人文学者的文学训练,在“单凭表层特征辨认出哪一篇是 AI 生成的”上,并没有带来什么特别的优势。这项研究并没有检验文学训练是否预测“评估一篇医学生论文中的论证是否立得住”的准确率。它本来就不是为此设计的;它测的是别的东西。

这项任务区分——具体说,是“文学训练与第三项任务相关”这个主张——并没有直接实验的支持。那个实验还没有做过。这是本文诚实的空白处:把文学能力当作推理连贯性评估能力的这个结构性论证,其根据是机制、类比与趋同的证据,而不是一项直接测量“文学型读者是否比非文学型读者以更高比率抓到 AI 推理失败”的对照研究。这个论证在呼唤那个实验。它的缺席是一处必须被点名、而不是被掩盖的空白。

这个结构性论证所主张的是:第三项任务,正是一位细读某部论证密集之作的读者所做的事——持续地、隐含地、在数百页之上,作为一种练出来的习惯。评估一个断言是否配得上它的自信,一次转折是否站得住,一个结论是真的被支撑了还是只是被宣称了——这不是把逻辑施加到孤立命题上。这是一种对论证纹理的、被养成的敏感,由长期接触那些在大尺度上要么做到、要么没做到这些品质的写作所建立。

AGI 时代使用 AI 的约束瓶颈不是内容传递。约束瓶颈是第三项任务——而第三项任务,自动化工具没有解决,技术能力框架也没有处理。处理它的是一种人类能力。这种能力有一个名字,而描述它的研究一直明晃晃地藏在错误的那个系里。

阅读所建立、且无法外包的东西#

Maryanne Wolf 在一篇 Medium / Thrive Global 的文章中写道:“深度阅读,和阅读脑回路本身一样,不是天生就有的;它由使用建成,或者因废弃而萎缩。”6 这个观察不是比喻。深度阅读回路——那个涉及语义、音系与语用加工,并连同工作记忆和类比推理的区域网络——是通过练习组装起来的。它不是预装好送来的。那些在合适的发育窗口中被给予持续、有难度的文本的孩子,会建起另一些孩子建不起来、或者只能建得残缺的回路。Wolf 对大学生群体的临床记录——他们进入大学时越来越无法在不丢线索的情况下维持长篇阅读——不是随机对照试验,但它是数十年与这个正处在关键位置的特定人群打交道所积累的专家临床观察。6

那个回路一旦建成,究竟使什么成为可能?不是检索。也不是“知道这些词是什么意思”这种意义上的理解。是某种更具体、也更难命名的东西:对“一个文本正在做真正的论证工作”与“一个文本正在表演论证的自信、却几乎什么也没扛”这两者之间那道被感受到的分界。这正是 Kyle Chayka 在《Behavioral Scientist》上所称的鉴赏力——而他关于算法文化的论证,从一个意想不到的角度照亮了阅读这个问题。7

“我们在算法信息流中于可得性上所获得的东西——随时能扫过一大片材料的即时通道——我们在鉴赏力上失去了,而鉴赏力需要深度和意图。”7 Chayka 写的是音乐、电影和一般意义上的文化对象——是算法推荐系统如何把评价性的辨别力压平成被动消费。但他所命名的机制,与 AI 输出评估中所涉的是同一个:鉴赏力需要的是深度和意图,不是通道。“不得不有意识地累积一份收藏,并把自己究竟最喜欢某位创作者或某个文化体的哪些方面想清楚,这意味着成为一位鉴赏者。”7 关键词是有意识地累积​——正是那种刻意的、累积性的投入,建立起单靠通道无法提供的内部标准。

Wolf 的回路与 Chayka 的鉴赏力,是从不同侧面描述同一个过程:关于“持续投入于困难而优秀的写作会在读者身上建成什么”的神经解释与文化解释。那位在艰深文本上花过许多年的读者——不只是消费它们,而是与它们角力,丢掉线索再找回来,追踪那些配得上自己结论的论证的构造——建成了某种 AI 辅助型读者建不成的东西。不是知识。是一件校准仪器。

认知卸载的文献让这个机制变得精确。Risko 和 Gilbert 的奠基框架,经 Grinschgl 和同事在三个实验、共 516 名被试上确认,确立了:因卸载给外部工具而被释放出来的认知资源,看起来是丢失了​,而不是被腾出来了——除非学习者带着明确的学习目标,否则它们“不会对记忆的形成有所贡献”。8 用 AI 起草论文的学生,并没有把省下的认知努力存起来;他们只是没有做那份功课。同样的逻辑适用于阅读:用 AI 摘要替代持续阅读的学生,并没有为更高阶的分析腾出认知资源。他们没能发育出更高阶分析所依托的那个回路。

Lydiard 那个类比是例示,不是机制。Arthur Lydiard 对长跑的奠基性贡献,在于认识到有氧基础训练——缓慢、持续、大量的努力——能让日后的专项表现训练成为可能。这个结构逻辑映射到阅读上:基底使结构成为可能;把这个比例倒过来,会损伤基底。运动生理学里没有哪个定量比例能干净地映射到认知发展上。这个类比起的是照明作用;它并不构成证明。

上一节点名的那个直接实验还没有做;这里给出的是结构性论证。Wolf 的回路、Chayka 的鉴赏力,以及 Risko 和 Gilbert 的卸载框架,趋同于同一个机制——而它们合在一起,命名了那个无法被外包的东西。

无法被外包的,正是那件校准仪器本身——那个关于论证质量的内部标准。再多的 AI 输出访问权也建不起它,而 AI 中介的阅读还会主动阻止它发育,因为建起它的那份认知工作,要求的是没有脚手架的投入。

教育科技拿走了什么#

Robert Bjork 关于学习的研究,历经数十年和多个实验室的重复验证,趋同于一个既稳健又反直觉的发现:那些容易带来快速表现提升的条件,往往支撑不了长期保持;而那些看起来妨碍即时学习的条件,反倒倾向于固结成持久的能力。9 这就是“合意困难”框架——它的洞见是:有产出的挣扎、交错练习和分散练习,不是学习的障碍,而正是学习的机制。在控制最严格的研究中,提取练习让回忆的表现比重读同样材料高出约 50%。9 在同等任务上,交错练习产生的测验成绩是 63%,而分块练习是 20%。9 机制并不神秘:困难迫使那份能固结记忆、能建立可迁移技能的认知工作发生。轻松测出的是即时表现,却不建立持久的能力。

Khanmigo 这个案例是一个真实的架构反例。可汗学院的 AI 辅导工具包括一个文本难度调节器,为吃力的读者给经典文学与历史文本搭脚手架;一个文本分块功能,把密集段落拆开,并鼓励学生比较简化版与原文;还有一套苏格拉底式提问方法,不给答案,而是引导学生走向推理。3 这些不是摩擦消除器;它们是摩擦调节器。一个让原本被挡在门外的学生得以进入一篇有挑战性文本的脚手架,与一个单纯把挑战拿掉的设计,是两回事。本文批评的不是脚手架本身。批评的是脚手架之后该来的东西——或者说,目前没能来的东西。

按一项针对同行评议文献的系统性范围综述所记录的,自适应学习平台是围绕投入度优化来设计的——实时调整教学策略以维持学习者投入,并随着学生推进而降低认知负荷,是个性化自适应系统明确的设计命题。10 Bjork 的研究预测了当学习被工程化为“心流”时会发生什么:短期收益无法固结成那些更难、更不舒服的条件本可以建起的持久能力。该综述引用的一项研究发现,88% 的学习者在个性化自适应系统中报告了心流状态——而心流状态,恰恰是 Bjork 的合意困难研究所反对的那种条件。心流是舒服的;而“建立持久能力”意义上的学习,是不舒服的。正是自适应系统那个明确的商业主张——无缝调整以维持投入——让它们与耐力所要求的那种有产出的不适处在架构上的冲突之中。逐步撤除的脚手架会拉低投入度指标。而商业模式不奖励撤除。

元分析证据为这个架构论证提供了根基。Alrawashdeh 和同事综合了 12 个国家的 27 项研究,发现个性化自适应学习对阅读素养的总体效应量为 g=0.29。11 温和,但不是零。真正让画面复杂化的发现是:在元分析数据中,无教师在场的条件比有教师在场的条件显示出更大的效应。11 这是反直觉的——如果自适应学习是对教师教学的补充,那么教师在场本该有帮助。这个模式提示,被测到的收益集中在低阶的、操练式的领域,也就是自动化反馈擅长、而教师引导的讨论并非必需的地方。而更高阶的理解——那种建立耐力、要求持续困难、涉及长程论证追踪的理解——恰恰是教师在场最要紧的领域;它同时也是元分析无法分离出效应的领域,因为各研究报告结果的方式不一致。

Cheung 和 Slavin 专门考察了部署最广的那些配置。11 综合性的独立教育科技项目——像 READ 180 和 Fast ForWord 这样在学校中作为独立阅读干预大规模部署的项目——产出的效应量是 0.04 到 0.06。接近于零。部署最广的工具,在最常见的学校部署模式下,效应贴着地板。这个模式与 Bjork 的预测严丝合缝:自适应学习显示出的收益,集中在技术反复操练低阶、易测量技能的地方。而那些更高阶的、由教师引导的、带有有产出的不适的工作——也就是建立阅读耐力的那部分工作——正是技术系统性地从学习环境中拿走的东西,而正是这次拿走,对 AGI 时代所要求的那项能力至关重要。

投入度指标与耐力指标拉向相反的方向。而商业模式选中的是错的那一个。

基底本来就在侵蚀#

阅读的衰退比它如今所叠加的 AI 更古老。

Bone 和同事分析了美国时间使用调查中 236,270 名受访者、跨越二十年的数据,追踪 2004 到 2023 年的休闲阅读。12 在某一天有阅读行为的美国人比例,从 28% 降到 16%——相对降幅 43%,年均下降约 3%(单一年龄组中降幅最陡的是 66 岁及以上的成年人;ATUS 不统计线上和屏幕阅读,因此这个指标专门覆盖的是传统书与电子阅读器格式)。社会经济地位的差距明显拉大:到 2023 年,受过研究生教育的读者每日阅读的可能性,是受过高中教育者的 2.79 倍。12

青少年数据来自另一个互补的来源。Twenge 和同事分析了“监测未来”调查四十年的数据——每年约五万名学生,四十年间超过一百万名青少年——发现报告每日阅读的美国十二年级学生,从 1970 年代末的约 60% 降到 2016 年的 16%。13 约有三分之一的人报告,在调查前的一年里没有为消遣读过任何书。数据窗口终止于 2016 年——比 ChatGPT 发布早六年。这场崩塌的直接推手是智能手机;而智能手机在青少年群体中于 2010 到 2014 年间变得无处不在。AI 继承了这个趋势;它不是这个趋势的源头。13

AI 加到一个已被削弱的基底之上的,是一种不同的成本结构。当基底强健时,把 AI 用在错的任务上是一个生产力错误。当基底虚弱时,把 AI 用在错的任务上是一场认知危机——使用者抓不到他从前抓得到的东西,而且他并不总是知道自己抓不到。

MIT 媒体实验室那项研究点出了认知机制,并作了恰当的保留。Kosmyna 和同事在三种写作条件下——无辅助、搜索引擎辅助、LLM 辅助——测量了 54 名被试的 EEG 脑连接。14 层级很清楚:纯用脑的条件产生了最强的 alpha 与 beta 网络连接;LLM 辅助条件产生的最弱。关键在于,当被分配到 LLM 条件的被试在第四次实验中改为无辅助写作时,他们的脑连接仍然低于纯用脑组——一种在工具被撤走之后仍然持续的“投入不足”。83% 这个数字比连接强度的发现更稳健:83% 使用 LLM 的被试,无法引用自己几分钟前写下的文章内容。14 作者造了“认知负债”一词来描述这个模式。这项研究 n=54(其中 n=18 完成了交叉实验),截至撰稿时仍未经同行评议;方向性的发现在这些限制之下仍然成立,但具体量级不应被当作定论。14

Gerlich 对 666 名被试的调查发现,频繁使用 AI 工具与批判性思维得分之间存在显著负相关,并以认知卸载为中介,其中 17–25 岁这一群体的 AI 依赖度最高、得分最低——这是一个相关性模式,不是因果证明,但与方向性的论证是一致的。15

基底因非 AI 的原因已经削弱了二十年。AGI 时代抬高了这份虚弱的代价。而教育科技的主流部署方式,被设计成消除有产出的不适、优化投入度,同时加速了这两者——而且它恰恰是在基底缺席变得代价最高的那一刻这样做的。

阅读作为公共基础设施#

本文所辩护的这项能力,是可以训练的。

一项针对 1,962 名台湾老年人的 14 年纵向研究发现,每周至少阅读一次,在随访第 6 年、第 10 年和第 14 年都独立预测了低 46% 的认知衰退风险,即便在控制了其他认知活动——电视、广播、游戏、社交——之后仍然成立。16 校正后比值比为 0.54(95% CI:0.34–0.86),在各教育水平上都一致。这个发现的适用范围很要紧:人群是老年人(64 岁及以上),而这个发现针对的是认知维持,不是青少年期的认知发展。剂量—反应信号是阅读所特有的,而且是真实的。16

家庭藏书的数据更为宏阔。Sikora 和同事分析了 PIAAC 数据集——31 个社会中的 162,955 名成年人,采样于 2011 到 2015 年间。17 在青少年时期的家中拥有 80 本以上藏书,独立预测了成年后更高的读写能力、算术能力和 ICT 问题解决能力——而且是在父母教育程度或个人自身受教育程度的影响之外​。这个效应是对数线性的,藏书量较小时回报最大,而且它跨越三大洲一直延续到工作年龄。17 它起作用的渠道既不神秘也不是比喻:在大多数机构化教学之前就早早接触书籍,建起了此后一切认知工作所依托的那个基底。

这一点的神经生物学根据,在两项针对儿童的独立研究中得到了确认。PMC9588575 记录了与社会经济地位相关的、阅读相关脑区上的差异——左半球结构性与功能性连接降低,阅读网络的皮层表面积减小——并把它们与书籍可及性的减少和亲子共读频率的降低联系起来。18 PMC12309101 在 1,534 名一至六年级中国学生身上确认了这条中介路径:家中藏书数量和开始阅读的年龄,中介了社会经济地位与阅读能力之间的关系。18 因果方向是从发育逻辑推断的,而不是从实验得来的——但书籍贫乏环境的生物学相关物,在童年期就是可观察的,不是作为猜测,而是作为神经影像数据。

这些是带有认知后果的公平性发现。阅读环境的梯度很陡。受过研究生教育的美国人每日阅读的比率,是受过高中教育的美国人的 2.79 倍。12 富裕家庭的孩子更可能在书籍丰富的家中长大。机构化的阅读——被规定的、有难度的、有教师在场来协助处理这份难度的——曾是少数几条渠道之一,能让那些家中书籍不丰的孩子被要求去接触困难文本。站得住的主张是接触​,而不是拉平​。机构化的阅读课程并没有拉平结果;社会经济地位的差距在五十年的强制教学中一直存在。接触困难文本是必要的,但不是充分的;差距之所以持续,是因为机构化的那个版本不够好,而不是因为接触本身错了。

布迪厄式的批评——文学教育在历史上一直充当着阶层分化的机制,把文化能力标记为归属感的代理——值得交锋,而不是被打发掉。有难度的阅读课程在机构中被部署的历史记录,让这个批评在描述层面是准确的。本文的主张不是要保住某个经典书目,而是要保住那份把持续的困难阅读当作发育必需品的承诺​。谁的书、用什么语言、由哪些作者写——那些是关于实施的下游问题。上游的问题是:阅读所建立的那项能力,是否应当继续是一项机构优先事项。这个问题与“由哪些文本来承担训练负荷”无关。

教育科技的公平性论证是志向性的:结果“在低收入环境中仍无定论”,而且很少有大规模研究考察社会经济地位如何与教育科技的有效性交互。19 机构化阅读的论证有另一个毛病——它有拉平失败的记录——但存在一个重要的不对称:在布卢姆第二到第四层上免费提供的 AI 辅导,是一项真实的干预,而支持它的成本论证也是真实的。可是,如果本文的论点成立——如果 AGI 时代的约束瓶颈,是 AI 辅导建不起来的那项更高阶评估能力——那么,瞄准了错误能力的免费辅导,就不是它看上去的那种公平性胜利。辅导触及了更多学生;而最需要被发育出来的那项技能,仍然没有被碰到。

普林斯顿大学选定 Maryanne Wolf 的《Reader, Come Home》作为 2030 届的入学预读书,于 2026 年 4 月公布。20 校长 Eisgruber 写道,深度的、沉浸式的阅读“处在普林斯顿教育的核心”——而有挑战性的书为学生提供了“某种独特、宝贵且不可替代的东西”。机构层面的认可正在到来。问题在于它是否来得及,以及它是否只到达那些学生本来就有书香家庭、父母本来就读书的机构。

而这个问题伸向教育之外。追踪一个长论证直到它的结论、辨认出自信在什么时候没有挣来、察觉一个看起来像推理的结构其实只是在表演推理——这些是形成有见识的集体判断的认知前提:为了跟上一场公共卫生紧急事件的推理,为了评估一场政策辩论中的取舍,为了把一个构造精良的民主论证与一个构造精良的、听起来像民主的论证区分开。一个无法在自信而不确定的条件下评估推理的选民群体,不只是受教育程度更低。它更容易被操纵——不是被 AI 操纵,而是被那个控制着 AI 自信地说些什么的人操纵。

文学能力,阅读耐力,Wolf 所描述的那个回路和 Chayka 所命名的那份鉴赏力——这些是真正评估 AI 输出所依赖的基础设施。是保护它们还是抛弃它们,是一个关于“下一代将继承什么样的认知公地”的决定——以及那片公地是否有能力承担 AGI 时代比以往任何时代都更沉重地要求于它的工作。

延伸阅读#

  • Wolf, Maryanne. Reader, Come Home: The Reading Brain in a Digital World​. Harper, 2018.——深度阅读回路论证的第一手来源;SC1-B 取材于 Wolf 的临床观察,以及这本书所作的更宽泛的神经科学论证。正文部分仅据图书馆本;这本书的论点是本文根基性的思想背景。
  • Carr, Nicholas. The Shallows: What the Internet Is Doing to Our Brains​. W. W. Norton, 2010.——那个奠基性的通俗论证:持续注意力与深度阅读正被超文本和扫读习惯结构性地削弱。本文从 Carr 停下的地方接着走,并把基底论证专门应用到 AI 时代的评估问题上。
  • Newport, Cal. Deep Work: Rules for Focused Success in a Distracted World​. Grand Central Publishing, 2016.——注意力经济批判在职业场景中的应用;对本文主要读者而言是相关背景,他们很可能已经接触过 Newport 的框架,而本文正建立在他们既有的词汇之上。
  • Chayka, Kyle. Filterworld: How Algorithms Flattened Culture​. Doubleday, 2024.——《Behavioral Scientist》那些引文(SC3-C)所出自的完整专著论证;鉴赏力这一论点在书里比在被引的那篇文章里展开得更充分。
  • Bjork, R.A., & Bjork, E.L. "Desirable difficulties in theory and practice." Journal of Applied Research in Memory and Cognition​, 9(4), 475–479, 2020.——合意困难这一研究纲领最新的一次综合;与 SC4-A(2011 年那一章)互补,而且对想读第一手文献的读者更易获取。
  • Cheung, A.C.K., & Slavin, R.E. "How Features of Educational Technology Applications Affect Student Reading Outcomes: A Meta-Analysis." Educational Research Review​, 2012.(84 项研究,6 万余名 K-12 被试)——独立项目那组数字(SC4-B)所出自的、覆盖面更广的元分析;2012 年这次综合所涵盖的证据基础,比 2013 年那篇聚焦吃力读者的论文更宽。
  • PMC11047126. "Cognitive reserve over the life course and risk of dementia: a systematic review and meta-analysis." Frontiers in Aging Neuroscience​, 2024.——证据档案中作为储备证据提及的那项认知储备元分析(SC5-B)。它包含早年与晚年认知储备的比较;本文没有引用它,是因为“阅读把痴呆风险降低 18%”这种表述站不住(阅读只是广义认知储备的一个组成部分),但这项元分析支撑了本文关于“为什么早期认知投入重要”的论证。

Footnotes

  1. Kestin, G., et al. "AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting." Scientific Reports​, 2025. doi: 10.1038/s41598-025-97652-6. https://www.nature.com/articles/s41598-025-97652-6 2 3

  2. World Economic Forum. "Why AI literacy is now a core competency in education." May 2025. https://www.weforum.org/stories/2025/05/why-ai-literacy-is-now-a-core-competency-in-education/ | Purdue University. "Purdue Unveils Comprehensive AI Strategy; Trustees Approve AI Working Competency Graduation Requirement." December 2025. https://www.purdue.edu/newsroom/2025/Q4/purdue-unveils-comprehensive-ai-strategy-trustees-approve-ai-working-competency-graduation-requirement/ 2

  3. Khan Academy. "ELA AI Teacher Tools." Khan Academy Blog​, 2025. https://blog.khanacademy.org/ela-ai-teacher-tools/ 2

  4. Farquhar, S., Kossen, J., Kuhn, L., & Gal, Y. "Detecting hallucinations in large language models using semantic entropy." Nature​, 630, 625–630, 2024. https://www.nature.com/articles/s41586-024-07421-0 2

  5. "Detecting Artificial Intelligence–Generated Versus Human-Written Medical Student Essays: Semirandomized Controlled Study." PMC11914838​, 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC11914838/ 2

  6. Wolf, Maryanne. "Skim Reading Is the New Normal: The Effect on Society." Medium / Thrive Global​, 2018. 2

  7. Chayka, K. "How to Cultivate Taste in the Age of Algorithms." Behavioral Scientist​, 2024. https://behavioralscientist.org/how-to-cultivate-taste-in-the-age-of-algorithms/ 2 3

  8. Risko, E.F., & Gilbert, S.J. "Cognitive Offloading." Trends in Cognitive Sciences​, 20(9), 676–688, 2016. https://pubmed.ncbi.nlm.nih.gov/27542527/ | Grinschgl, S., Papenmeier, F., & Meyerhoff, H.S. "Consequences of cognitive offloading: Boosting performance but diminishing memory." PMC8358584​. https://pmc.ncbi.nlm.nih.gov/articles/PMC8358584/

  9. Bjork, E.L., & Bjork, R.A. "Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning." 收于 Psychology and the Real World​, 2011. https://bjorklab.psych.ucla.edu/wp-content/uploads/sites/13/2016/04/EBjork_RBjork_2011.pdf | Bjork, R.A., & Bjork, E.L. "Desirable difficulties in theory and practice." Journal of Applied Research in Memory and Cognition​, 9(4), 475–479, 2020. 2 3

  10. "Personalized adaptive learning in higher education: A scoping review of key characteristics and impact on academic performance and engagement." PMC11544060​, 2024. https://pmc.ncbi.nlm.nih.gov/articles/PMC11544060/

  11. Alrawashdeh, G., Fyffe, S., Azevedo, R., & Castillo, C. "Exploring the impact of personalized and adaptive learning technologies on reading literacy: A meta-analysis." Educational Research Review​, 2024. | Cheung, A.C.K., & Slavin, R.E. "The effectiveness of educational technology applications for enhancing reading achievement in K-12 classrooms: A meta-analysis." Educational Research Review​, 2012. 2 3

  12. Bone, J.K., et al. "The decline in reading for pleasure over 20 years of the American Time Use Survey." iScience​, 2025. 2 3

  13. Twenge, J.M., Martin, G.N., & Spitzberg, B.H. "Trends in U.S. Adolescents' Media Use, 1976–2016: The Rise of Digital Media, the Decline of TV, and the (Near) Demise of Print." Psychology of Popular Media Culture​, 8(4), 329–345, 2019. 2

  14. Kosmyna, N., Hauptmann, E., et al. "Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task." MIT Media Lab, 2025. arXiv 2506.08872.(n=54,其中 n=18 完成交叉实验;截至撰稿时未经同行评议。) 2 3

  15. Gerlich, M. "AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking." Societies​, 15(1), 6, 2025.

  16. "Reading activity prevents long-term decline in cognitive function in older people: evidence from a 14-year longitudinal study." PMC8482376​, International Psychogeriatrics​, 2020. https://pmc.ncbi.nlm.nih.gov/articles/PMC8482376/ 2

  17. Sikora, J., Evans, M.D.R., & Kelley, J. "Scholarly culture: How books in adolescence enhance adult literacy, numeracy and technology skills in 31 societies." Social Science Research​, 77, 1–15, 2019. 2

  18. "Socioeconomic status and reading outcomes: Neurobiological and behavioral correlates." PMC9588575​. https://pmc.ncbi.nlm.nih.gov/articles/PMC9588575/ | "Influence of socioeconomic status on children's reading abilities: the mediating role of home learning environment." PMC12309101​. https://pmc.ncbi.nlm.nih.gov/articles/PMC12309101/ 2

  19. "Bridging EdTech gaps: Examining learning equity in low-income educational settings." ScienceDirect​, 2025. https://www.sciencedirect.com/science/article/abs/pii/S0738059325001968 | Wiley. "Adaptive, but Equitable? Exploring the Impact of Machine Learning-Based Adaptive Support on Educational Debts in Undergraduate Chemistry." Science Education​, 2025. https://onlinelibrary.wiley.com/doi/10.1002/sce.70042

  20. Princeton University. "Reader, Come Home by Maryanne Wolf selected as Princeton Pre-read." April 8, 2026. https://www.princeton.edu/news/2026/04/08/reader-come-home-maryanne-wolf-selected-princeton-pre-read

Literary Competence Is AI-Safety Infrastructure

The Confident Wrong Answer#

The message came back in forty seconds — three crisp paragraphs, no hedging, no ellipses. A working professional had asked an AI assistant to explain the thematic significance of a particular narrative technique in a novel, and the model answered with the fluency of a seminar leader: a claim about unreliable narration, two supporting details, a closing observation about the author's intention. Every sentence parsed. The vocabulary was right. The argument had the shape of an argument.

And yet something was off. The professional — an experienced AI user, someone who had learned to prompt well and verify often — sensed it. Not a factual error: the character's name was right, the plot point was correctly described. Something more like a seam showing at the join between two ideas. The second paragraph's central claim did not quite follow from the first. The "therefore" was load-bearing, and the load was not there. The conclusion arrived too cleanly, tying together threads that had not actually been braided.

The professional read it again. The unease persisted. But there was no procedure for narrowing it — no experiment to run, no citation to check, no counter-source to consult. The discomfort was real; the instrument for locating its source was missing. After a pause, the answer was accepted, forwarded, cited in a report that would reach other people.

The problem was not the AI's confidence. The problem was the evaluator's silence.

There is a gap between receiving an AI output and evaluating it — and that gap is not closed by knowing more about how AI works.

What the Consensus Gets Right#

If the case for AI literacy as technical skill rested on intuition, this article would be easier to write. It does not.

Kestin and colleagues at Harvard ran a randomized controlled trial in which 194 undergraduate physics students were assigned either to a custom AI tutor or to active-learning instruction.1 The AI tutoring group showed learning gains 0.63 to 1.3 standard deviations larger than the active-learning control — in two sessions. Students reported higher engagement and higher motivation. The study was peer-reviewed, published in Scientific Reports​, and designed to prevent hallucination by using pre-written expert answers vetted for accuracy. These are not conditions that inflate results; they are conditions that test a real claim under reasonably rigorous controls.

The Kestin study is the best experiment in the field. An argument against AI tutoring that passes over this evidence has not done its work. The honest engagement begins with granting what it shows: AI tutoring, properly designed, efficiently delivers content knowledge in structured domains. The effect sizes are not marginal. The design is credible.1

The institutional consensus behind this evidence is substantial. The World Economic Forum's 2025 AI Literacy Framework identifies core competencies as algorithmic thinking, prompt engineering, understanding AI bias, data literacy, and critical thinking about AI outputs.2 Purdue University enacted an AI Working Competency graduation requirement in December 2025, the first major research university to do so.2 These frameworks reflect genuine policy deliberation, not panic. They represent the considered position of institutions that have spent decades thinking about what students need to know.

The cost-asymmetry argument reinforces the case. Khanmigo grew from roughly 40,000 to 700,000 student users in a single year.3 One-on-one AI tutoring, available at no cost to a student without resources, placed against one-on-one human tutoring at $50 to $200 per hour — the equity argument for AI tools is real, and the cost differential is not marginal.

What the Kestin authors themselves wrote, however, matters precisely here. In their limitations section, they noted: "we do not presume that structured AI tutoring will always outperform in-class active learning in all contexts, for example, those requiring complex synthesis of multiple concepts and higher-order critical thinking."1 The study tested understanding, applying, and analyzing — roughly Bloom's taxonomy levels one through four. The authors specified what remained outside its scope.

The question is not whether AI tutoring works at content delivery. It does. The question is whether content delivery is the binding constraint.

Three Detection Tasks#

When someone says "I need to be able to evaluate AI outputs," they are describing at least three different problems. Most discussions of AI literacy elide this. The conflation matters, because the skills required for each problem are genuinely different — and the one that is hardest is the one that receives the least attention.

The first task is stylistic authorship detection​: given a piece of text, was it written by an AI? This is a pattern-matching problem. The question is whether the prose has recognizable stylistic signatures of AI generation — the specific rhythmic flatness, the preference for certain connective phrases, the register that is simultaneously confident and toneless. This task has been studied directly.

The second task is factual fabrication detection​: does this citation exist, does this statistic trace to a real study, does this historical event match the historical record? This is a verification task. The skill it requires is knowing how to check — database access, cross-referencing, source tracing. Methodologically, it is closer to fact-checking journalism than to literary analysis.

The third task is reasoning-coherence evaluation​: does this argument actually hold together? Does the conclusion follow from the premises? Is this claim internally consistent with what was asserted two paragraphs earlier? Is the confidence of the assertion matched by the quality of the support? This is a different kind of work — not pattern-matching, not verification, but the evaluation of argumentative structure itself.

Farquhar and colleagues at Oxford published what is currently the most sophisticated automated approach to the third task: semantic entropy, which detects confabulation by having the model generate multiple candidate answers and checking whether they cluster around a consistent meaning.4 The method is elegant. It is also, unavoidably, a model-side operation: it requires access to the model's internal probability distributions and multiple generations. A person who receives a single confident output from a chatbot interface has none of this.

More importantly, the authors named a class of failures the method cannot reach: what they call "structural errors" — mistakes that are systematic and consistent across all of the model's outputs because they originate in what the model learned during training.4 These are precisely the failures most dangerous to end-users, because they are the ones the model produces with the most confidence and least self-correction. Semantic entropy cannot detect them. What is left, above the ceiling of automated detection, is human judgment.

A controlled study of this boundary — the German PMC11914838 experiment — is instructive. Thirteen humanities scholars and twenty-two medical professionals attempted to identify which of a set of German-language medical student essays were AI-generated.5 The medical professionals achieved 72% accuracy; the humanities scholars achieved 65%. The difference was statistically indistinguishable (OR 1.37, 95% CI 0.5–3.9). The humanities scholars rated ChatGPT prose as linguistically higher quality than the human student writing — not merely comparable: superior, by their literary judgment.5

The study found what it found, and it is worth understanding precisely what it tested: task one. Stylistic authorship detection in a non-literary domain, in German, on medical student essay prose. The literary training of the humanities scholars conferred no particular advantage at recognizing which text was AI-generated from surface features alone. The study did not test whether literary training predicts accuracy at evaluating whether the argument in a medical student essay holds together. It was not designed to; it measured something different.

The task distinction — and specifically the claim that literary training is relevant to task three — is not supported by a direct experiment. That experiment has not been run. This is the article's honest empty space: the structural argument for literary competence as the capacity for reasoning-coherence evaluation is grounded in mechanism, analogy, and convergent evidence, not in a controlled study that directly measures whether literary readers catch AI reasoning failures at higher rates than non-literary readers. The argument calls for that experiment. The absence of it is a gap that must be named, not concealed.

What the structural argument proposes is this: task three is what a careful reader of a densely argued work does — constantly, implicitly, over hundreds of pages — as a matter of practised habit. The evaluation of whether an assertion earns its confidence, whether a transition is warranted, whether a conclusion is genuinely supported or merely claimed, is not logic applied to isolated propositions. It is a cultivated sensitivity to argumentative texture, built by sustained exposure to writing that either demonstrates or fails to demonstrate these qualities at scale.

The binding constraint in AGI-era AI use is not content delivery. The binding constraint is task three — and task three is unsolved by automated tools and unaddressed by the technical-competency frameworks. What addresses it is a human capacity. The capacity has a name, and the research that describes it has been hiding in plain sight in the wrong department.

What Reading Builds That Cannot Be Outsourced#

Maryanne Wolf, in a Medium / Thrive Global article, wrote: "Deep reading, like the reading brain circuit itself, is not a given; it is built by use, or it atrophies from disuse."6 The observation is not metaphorical. The deep reading circuit — the network of regions involving semantic, phonological, and pragmatic processing, along with working memory and analogical reasoning — assembles through practice. It does not arrive pre-built. Children who are given sustained, difficult text in the appropriate developmental windows build circuits that children who are not given such text do not build, or build incompletely. Wolf's clinical documentation of this in college students — who arrive at university increasingly unable to sustain long-form reading without losing the thread — is not an RCT, but it is expert clinical observation accumulated over decades of working with the specific population at stake.6

What does that circuit, once built, actually enable? Not retrieval. Not comprehension in the sense of knowing what the words mean. Something more specific and harder to name: the felt distinction between a text that is doing genuine argumentative work and a text that is performing argumentative confidence while carrying almost nothing. This is what Kyle Chayka, writing in Behavioral Scientist​, calls connoisseurship — and his argument about algorithmic culture illuminates the reading problem from an unexpected angle.7

"What we gain with algorithmic feeds in terms of availability — having instant access to a broad range of material to be scanned at will — we lose in connoisseurship, which requires depth and intention."7 Chayka is writing about music, film, cultural objects generally — the way algorithmic recommendation systems flatten evaluative discrimination into passive consumption. But the mechanism he names is the same one at stake in AI-output evaluation: connoisseurship requires depth and intention, not access. "Having to consciously accrue a collection and think through what you enjoy most about a particular creator or body of culture means becoming a connoisseur."7 The keyword is consciously accrue — the deliberate, cumulative engagement that builds an internal standard that access alone cannot provide.

Wolf's circuit and Chayka's connoisseurship are describing the same process from different sides: the neural and the cultural account of what sustained engagement with difficult, excellent writing builds in the reader. The reader who has spent years with demanding texts — not merely consuming them, but wrestling with them, losing the thread and finding it again, tracing the architecture of arguments that earn their conclusions — builds something that an AI-assisted reader does not. Not knowledge. A calibration instrument.

The cognitive offloading literature makes the mechanism precise. Risko and Gilbert's foundational framework, confirmed by Grinschgl and colleagues across 516 participants in three experiments, established that cognitive resources released by offloading to external tools appear lost​, not freed — they "do not contribute to the formation of memory" unless the learner has an explicit goal to learn.8 The student who uses an AI to draft an essay does not bank the cognitive effort; they simply do not do the work. The same logic applies to reading: a student who uses AI summaries in place of sustained reading does not free cognitive resources for higher-order analysis. They do not develop the circuit that higher-order analysis sits on.

The Lydiard analogy is illustration, not mechanism. Arthur Lydiard's foundational contribution to distance running was the recognition that base aerobic conditioning — slow, sustained, high-volume effort — enables specific performance work later. The structural logic maps to reading: substrate enables structure; inverting the ratio damages the substrate. No quantitative ratio from sports physiology maps cleanly to cognitive development. The analogy illuminates; it does not prove.

The direct experiment named in the previous section has not been run; what follows here is the structural argument. Wolf's circuit, Chayka's connoisseurship, and Risko and Gilbert's offloading framework converge on the same mechanism — and together they name what cannot be outsourced.

What cannot be outsourced is the calibration instrument itself — the internal standard for argumentative quality that no amount of access to AI outputs will build, and that AI-mediated reading actively prevents from developing, because the cognitive work that builds it requires engagement without the scaffold.

What EdTech Removes#

Robert Bjork's research on learning, replicated across decades and laboratories, converges on a finding that is both robust and counterintuitive: conditions that lend themselves to rapid performance gains often fail to support long-term retention, whereas conditions that seem to impede immediate learning tend to consolidate into durable capacity.9 This is the desirable-difficulties framework — the insight that productive struggle, interleaving, and spacing are not obstacles to learning but the mechanism of it. Retrieval practice, in the most carefully controlled studies, improves recall roughly 50% better than restudying the same material.9 Interleaved practice produces test performance of 63% versus 20% compared to blocked practice on equivalent tasks.9 The mechanism is not mysterious: difficulty forces the cognitive work that consolidates memory and builds transferable skill. Ease measures immediate performance without building durable capacity.

The Khanmigo case is a genuine architectural counter-instance. Khan Academy's AI tutoring tools include a Text Releveler that scaffolds canonical literary and historical texts for struggling readers, a Chunk Text feature that breaks down dense passages and encourages comparison between the simplified and original versions, and a Socratic-question approach that withholds answers and guides students toward reasoning rather than providing it.3 These are not friction-removers; they are friction-mediators. A scaffold that makes a challenging text accessible to a student who would otherwise be locked out is a different design from one that simply removes challenge. The article's critique is not of scaffolding as such. It is of what follows scaffolding — or rather, what currently fails to follow it.

Adaptive learning platforms, as documented in a systematic scoping review of the peer-reviewed literature, are designed around engagement optimization — real-time adjustment of teaching strategies to maintain learner engagement, and reduction of cognitive load as students progress, are the explicit design propositions of personalized adaptive systems.10 Bjork's research predicts what happens when learning is engineered for flow: short-term gains that fail to consolidate into the durable capacity that the harder, less comfortable conditions would have built. One study cited in the review found 88% of learners reporting flow states with personalized adaptive systems — and flow state is precisely the condition Bjork's desirable-difficulties research argues against. Flow is comfortable; learning, in the sense that builds durable capacity, is not comfortable. It is the adaptive systems' explicit business claim — seamless adjustment to maintain engagement — that puts them in architectural conflict with the productive discomfort that stamina requires. Faded scaffolds reduce engagement metrics. The business model does not reward fading.

The meta-analytic evidence grounds the architectural argument. Alrawashdeh and colleagues synthesized 27 studies across 12 countries and found an overall effect size of g=0.29 for personalized adaptive learning on reading literacy.11 Modest, but not zero. The finding that complicates the picture: teacher-absent conditions showed larger effects than teacher-present conditions in the meta-analytic data.11 This is counterintuitive — if adaptive learning supplements teacher instruction, teacher presence should help. The pattern suggests the measured gains are concentrated in lower-order, drill-and-practice domains where automated feedback excels and teacher-facilitated discussion is not needed. Higher-order comprehension — the kind that builds stamina, requires sustained difficulty, involves extended argument-tracking — is the domain where teacher presence matters most; it is also the domain where the meta-analysis cannot isolate effects because studies report outcomes inconsistently.

Cheung and Slavin examined the most widely deployed configurations specifically.11 Comprehensive standalone EdTech programs — programs like READ 180 and Fast ForWord, deployed at scale in schools as standalone reading interventions — produced effect sizes of 0.04 to 0.06. Near-zero. The most widely deployed tools, in the most common school deployment pattern, show effects at the floor. The pattern fits the Bjork prediction exactly: the gains adaptive learning shows are concentrated where the technology drills lower-order, easily measurable skills. The higher-order, teacher-facilitated, productive-discomfort work — the work that builds reading stamina — is what the technology systematically removes from the learning environment, and it is that removal that matters for the capacity the AGI era requires.

The engagement metric and the stamina metric pull in opposite directions. The business model has selected for the wrong one.

The Substrate Was Already Eroding#

The reading decline is older than the AI it now compounds.

Bone and colleagues analyzed the American Time Use Survey across 236,270 respondents over twenty years, tracking reading for pleasure from 2004 to 2023.12 The proportion of Americans reading on a given day fell from 28% to 16% — a 43% relative decline, occurring at a rate of roughly 3% per year (the steepest single age-group decline was among adults 66 and older; the ATUS excludes online and screen reading, so the measure covers traditional and e-reader formats specifically). The SES gap widened substantially: postgraduate-educated readers were 2.79 times more likely to read daily than high-school-educated readers by 2023.12

The adolescent data comes from a different and complementary source. Twenge and colleagues analyzed four decades of the Monitoring the Future survey — roughly 50,000 students per year, over a million teenagers across forty years — and found that American 12th-graders reporting daily reading fell from approximately 60% in the late 1970s to 16% by 2016.13 Approximately one in three reported reading no books for pleasure in the year preceding the survey. The data window ends in 2016 — six years before the release of ChatGPT. The smartphone was the proximate driver of this collapse; the smartphone became ubiquitous between 2010 and 2014 in the adolescent population. AI inherits this trend; it did not originate it.13

What AI adds to a substrate already weakened is a different cost structure. When the substrate was strong, using AI for the wrong task was a productivity error. When the substrate is weak, using AI for the wrong task is an epistemic crisis — the user cannot catch what they used to be able to catch, and they do not always know they cannot.

The MIT Media Lab study names the cognitive mechanism, with appropriate hedging. Kosmyna and colleagues measured EEG brain connectivity in 54 participants across three writing conditions — unaided, search-engine-assisted, and LLM-assisted.14 The hierarchy was clear: brain-only condition produced the strongest alpha and beta network connectivity; LLM-assisted condition produced the weakest. Critically, when participants who had been assigned to the LLM condition switched to writing unaided in a fourth session, their brain connectivity remained below that of the brain-only group — "under-engagement" that persisted even after the tool was removed. The 83% figure is more robust than the connectivity-magnitude finding: 83% of LLM-using participants were unable to quote from essays they had written minutes before.14 The authors coined "cognitive debt" to describe the pattern. The study used n=54 (with n=18 completing the crossover session) and remains unpeer-reviewed as of this writing; the directional finding survives these constraints, but specific magnitudes should not be treated as settled.14

Gerlich's survey of 666 participants found a significant negative correlation between frequent AI tool use and critical thinking scores, mediated by cognitive offloading, with the 17–25 cohort showing the highest AI dependence and the lowest scores — a correlational pattern, not a causal demonstration, but consistent with the directional argument.15

The substrate had been weakening for two decades from non-AI causes. The AGI era raises the cost of the weakness. EdTech's dominant deployment, designed to remove productive discomfort and optimize engagement, accelerates both — and it does so at exactly the moment when the substrate's absence becomes most expensive.

Reading as Public Infrastructure#

The capacity this article defends is trainable.

A 14-year longitudinal study of 1,962 Taiwanese older adults found that reading at least once per week independently predicted a 46% lower risk of cognitive decline at 6, 10, and 14 years of follow-up, even after controlling for other cognitive activities — television, radio, games, socializing.16 The AOR was 0.54 (95% CI: 0.34–0.86), consistent across all education levels. The scope of the finding matters: the population is older adults (age 64+), and the finding addresses cognitive maintenance, not cognitive development in youth. The dose-response signal is specific to reading and genuine.16

The home library data is more sweeping. Sikora and colleagues analyzed the PIAAC dataset — 162,955 adults across 31 societies, sampled between 2011 and 2015.17 Growing up with 80 or more books in the adolescent home independently predicted higher adult literacy, numeracy, and ICT problem-solving skills — beyond the effect of parental education or the individual's own educational attainment. The effect was loglinear, with the greatest returns at smaller library sizes, and it persisted into working age across three continents.17 The channel by which this operates is neither mystery nor metaphor: early exposure to books, before most institutional instruction, builds the substrate that all subsequent cognitive work draws on.

The neurobiological grounding for this was confirmed in two separate studies of children. PMC9588575 documented SES-associated differences in reading-relevant brain regions — reduced left-hemisphere structural and functional connectivity, reduced cortical surface area in the reading network — linked to reduced access to books and lower parent-child reading frequency.18 PMC12309101 confirmed the mediation pathway in 1,534 Chinese students across grades one through six: the number of books at home and the age at which reading was initiated mediated the relationship between SES and reading ability.18 The causal direction is inferred from developmental logic, not from an experiment — but the biological correlates of book-poor environments are observable in childhood, not as speculation but as neuroimaging data.

These are equity findings with cognitive consequences. The reading-environment gradient is steep. Postgraduate-educated Americans read daily at 2.79 times the rate of high-school-educated Americans.12 Wealthy children are more likely to grow up with book-rich homes. Institutional reading — mandated, demanding, with teachers present to facilitate the difficulty — was one of the few channels by which children without book-rich homes were required to engage with difficult text. The defensible claim is exposure​, not equalization​. Institutional reading curricula did not equalize outcomes; the SES gap persisted through fifty years of mandated instruction. Exposure to demanding text is necessary, not sufficient; the gap persisted because the institutional version was inadequate, not because the exposure itself was misguided.

The Bourdieu critique — that literary education has functioned historically as a class-stratifying mechanism, marking cultural competence as a proxy for belonging — deserves engagement rather than dismissal. The record of how demanding reading curricula have been deployed institutionally makes the critique descriptively accurate. The article's claim is not to preserve the canon but to preserve the commitment to sustained difficult reading as a developmental requirement. Whose books, in what languages, by what authors — those are downstream questions about implementation. The upstream question is whether the capacity the reading builds should remain an institutional priority. That question is independent of which texts carry the training load.

The EdTech equity case is aspirational: results "remain inconclusive for low-income settings," and few large-scale studies examine how SES interacts with EdTech effectiveness.19 The institutional reading case has a different problem — it has documented failures at equalization — but an important asymmetry: free AI tutoring at Bloom levels two through four is a genuine intervention, and the cost argument for it is real. If the article's thesis holds — if the binding constraint in the AGI era is the higher-order evaluation capacity that AI tutoring does not build — then free tutoring targeted at the wrong capacity is not the equity win it appears to be. The tutoring reaches more students; the skill that most needs developing is left untouched.

Princeton University selected Maryanne Wolf's Reader, Come Home as its Pre-read for the Class of 2030, announced in April 2026.20 President Eisgruber wrote that deep, immersive reading was "at the heart of a Princeton education" — and that challenging books offer students "something distinctive, valuable, and irreplaceable." The institutional recognition is arriving. The question is whether it arrives in time, and whether it arrives only at institutions whose students already have book-rich homes and parents who read.

That question reaches past education. The capacity to track a long argument to its conclusion, to recognize when confidence is unearned, to notice when a structure that looks like reasoning is only performing it — these are the cognitive preconditions for informed collective judgment: for following the reasoning of a public health emergency, for evaluating the trade-offs in a policy debate, for distinguishing a well-constructed democratic argument from a well-constructed democratic-sounding argument. An electorate that cannot evaluate reasoning under conditions of confident uncertainty is not merely less educated. It is more manipulable — not by the AI, but by whoever controls what the AI confidently says.

Literary competence, reading stamina, the circuit Wolf describes and the connoisseurship Chayka names — these are the infrastructure on which genuine evaluation of AI output depends. The decision to protect them, or to abandon them, is a decision about what kind of epistemic commons the next generation will inherit — and whether that commons will be capable of the work that the AGI era, more than any previous era, requires of it.

References#

  1. Kestin, G., et al. "AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting." Scientific Reports​, 2025. doi: 10.1038/s41598-025-97652-6. https://www.nature.com/articles/s41598-025-97652-6 ↩︎ ↩︎ ↩︎

  2. World Economic Forum. "Why AI literacy is now a core competency in education." May 2025. https://www.weforum.org/stories/2025/05/why-ai-literacy-is-now-a-core-competency-in-education/ | Purdue University. "Purdue Unveils Comprehensive AI Strategy; Trustees Approve AI Working Competency Graduation Requirement." December 2025. https://www.purdue.edu/newsroom/2025/Q4/purdue-unveils-comprehensive-ai-strategy-trustees-approve-ai-working-competency-graduation-requirement/ ↩︎ ↩︎

  3. Khan Academy. "ELA AI Teacher Tools." Khan Academy Blog​, 2025. https://blog.khanacademy.org/ela-ai-teacher-tools/ ↩︎ ↩︎

  4. Farquhar, S., Kossen, J., Kuhn, L., & Gal, Y. "Detecting hallucinations in large language models using semantic entropy." Nature​, 630, 625–630, June 2024. doi: 10.1038/s41586-024-07421-0. https://www.nature.com/articles/s41586-024-07421-0 ↩︎ ↩︎

  5. "Detecting Artificial Intelligence–Generated Versus Human-Written Medical Student Essays: Semirandomized Controlled Study." PMC11914838​, 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC11914838/ ↩︎ ↩︎

  6. Wolf, Maryanne. "Skim Reading Is the New Normal: The Effect on Society." Medium / Thrive Global​, 2018. (Quote confirmed verbatim from this piece: "Deep reading, like the reading brain circuit itself, is not a given; it is built by use, or it atrophies from disuse.") See also: Wolf, Maryanne. Reader, Come Home: The Reading Brain in a Digital World​. Harper, 2018. Well, T. "The Reading Crisis in College." Psychology Today​, March 11, 2025. https://www.psychologytoday.com/us/blog/the-clarity/202503/the-reading-crisis-in-college ↩︎ ↩︎

  7. Chayka, K. "How to Cultivate Taste in the Age of Algorithms." Behavioral Scientist​, 2024. https://behavioralscientist.org/how-to-cultivate-taste-in-the-age-of-algorithms/ See also: Chayka, K. Filterworld: How Algorithms Flattened Culture​. Doubleday, 2024. ↩︎ ↩︎ ↩︎

  8. Risko, E.F., & Gilbert, S.J. "Cognitive Offloading." Trends in Cognitive Sciences​, 20(9), 676–688, 2016. https://pubmed.ncbi.nlm.nih.gov/27542527/ | Grinschgl, S., Papenmeier, F., & Meyerhoff, H.S. "Consequences of cognitive offloading: Boosting performance but diminishing memory." PMC8358584​. https://pmc.ncbi.nlm.nih.gov/articles/PMC8358584/ ↩︎

  9. Bjork, E.L., & Bjork, R.A. "Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning." In Psychology and the Real World​, 2011. https://bjorklab.psych.ucla.edu/wp-content/uploads/sites/13/2016/04/EBjork_RBjork_2011.pdf | Bjork, R.A., & Bjork, E.L. "Desirable difficulties in theory and practice." Journal of Applied Research in Memory and Cognition​, 9(4), 475–479, 2020. ↩︎ ↩︎ ↩︎

  10. "Personalized adaptive learning in higher education: A scoping review of key characteristics and impact on academic performance and engagement." PMC11544060​, 2024. https://pmc.ncbi.nlm.nih.gov/articles/PMC11544060/ ↩︎

  11. Alrawashdeh, G., Fyffe, S., Azevedo, R., & Castillo, C. "Exploring the impact of personalized and adaptive learning technologies on reading literacy: A global meta-analysis." Educational Research Review​, 2023. doi: 10.1016/j.edurev.2023.100541. https://www.sciencedirect.com/science/article/abs/pii/S1747938X23000805 | Cheung, A.C.K., & Slavin, R.E. "Effects of Educational Technology Applications on Reading Outcomes for Struggling Readers: A Best-Evidence Synthesis." Reading Research Quarterly​, 2013. doi: 10.1002/rrq.50 ↩︎ ↩︎ ↩︎

  12. Bone, J.K., et al. "The decline in reading for pleasure over 20 years of the American Time Use Survey." iScience​, vol. 28, no. 9, 2025, art. 113288. doi: 10.1016/j.isci.2025.113288. https://pmc.ncbi.nlm.nih.gov/articles/PMC12496190/ ↩︎ ↩︎ ↩︎

  13. Twenge, J.M., Martin, G.N., & Spitzberg, B.H. "Trends in U.S. Adolescents' Media Use, 1976–2016: The Rise of Digital Media, the Decline of TV, and the (Near) Demise of Print." Psychology of Popular Media Culture​, 8, 329–345, 2018. doi: 10.1037/ppm0000203. https://psycnet.apa.org/record/2018-41062-001 ↩︎ ↩︎

  14. Kosmyna, N., Hauptmann, E., et al. "Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task." arXiv:2506.08872​, June 2025. https://arxiv.org/abs/2506.08872 ↩︎ ↩︎ ↩︎

  15. Gerlich, M. "AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking." Societies​, 15(1), 6, January 2025. doi: 10.3390/soc15010006. https://www.mdpi.com/2075-4698/15/1/6 ↩︎

  16. "Reading activity prevents long-term decline in cognitive function in older people: evidence from a 14-year longitudinal study." PMC8482376​, International Psychogeriatrics​, 2020. https://pmc.ncbi.nlm.nih.gov/articles/PMC8482376/ ↩︎ ↩︎

  17. Sikora, J., Evans, M.D.R., & Kelley, J. "Scholarly culture: How books in adolescence enhance adult literacy, numeracy and technology skills in 31 societies." Social Science Research​, 77, January 2019. doi: 10.1016/j.ssresearch.2018.10.003. https://www.sciencedirect.com/science/article/abs/pii/S0049089X18300607 ↩︎ ↩︎

  18. "Socioeconomic status and reading outcomes: Neurobiological and behavioral correlates." PMC9588575​. https://pmc.ncbi.nlm.nih.gov/articles/PMC9588575/ | "Influence of socioeconomic status on children's reading abilities: the mediating role of home learning environment." PMC12309101​. https://pmc.ncbi.nlm.nih.gov/articles/PMC12309101/ ↩︎ ↩︎

  19. "Bridging EdTech gaps: Examining learning equity in low-income educational settings." ScienceDirect​, 2025. https://www.sciencedirect.com/science/article/abs/pii/S0738059325001968 | Wiley. "Adaptive, but Equitable? Exploring the Impact of Machine Learning-Based Adaptive Support on Educational Debts in Undergraduate Chemistry." Science Education​, 2025. https://onlinelibrary.wiley.com/doi/10.1002/sce.70042 ↩︎

  20. Princeton University. "Reader, Come Home by Maryanne Wolf selected as Princeton Pre-read." April 8, 2026. https://www.princeton.edu/news/2026/04/08/reader-come-home-maryanne-wolf-selected-princeton-pre-read ↩︎

Further Reading#

  • Wolf, Maryanne. Reader, Come Home: The Reading Brain in a Digital World​. Harper, 2018. — The primary source for the deep reading circuit argument; SC1-B draws on Wolf's clinical observations and the broader neuroscientific argument the book makes. Library-only for the interior; the book's thesis is the article's foundational intellectual context.

  • Carr, Nicholas. The Shallows: What the Internet Is Doing to Our Brains​. W. W. Norton, 2010. — The foundational popular argument that sustained attention and deep reading are being structurally degraded by hypertext and scanning habits. The present article picks up where Carr left off and applies the substrate argument specifically to AI-era evaluation.

  • Newport, Cal. Deep Work: Rules for Focused Success in a Distracted World​. Grand Central Publishing, 2016. — The professional application of the attention-economy critique; relevant background for the primary audience, who have likely encountered Newport's framework and whose existing vocabulary the article builds on.

  • Chayka, Kyle. Filterworld: How Algorithms Flattened Culture​. Doubleday, 2024. — The full book-length argument from which the Behavioral Scientist quotes (SC3-C) are drawn; the connoisseurship argument is more fully developed here than in the cited article.

  • Bjork, R.A., & Bjork, E.L. "Desirable difficulties in theory and practice." Journal of Applied Research in Memory and Cognition​, 9(4), 475–479, 2020. — The most recent synthesis of the desirable difficulties research program; complements SC4-A (the 2011 chapter) and is more directly available to readers wanting to engage the primary source.

  • Cheung, A.C.K., & Slavin, R.E. "How Features of Educational Technology Applications Affect Student Reading Outcomes: A Meta-Analysis." Educational Research Review​, 2012. (84 studies, 60,000+ K-12 participants) — The broader meta-analysis from which the standalone-program figures (SC4-B) are drawn; the 2012 synthesis covers a wider evidence base than the 2013 struggling-reader focus paper.

  • PMC11047126. "Cognitive reserve over the life course and risk of dementia: a systematic review and meta-analysis." Frontiers in Aging Neuroscience​, 2024. — The cognitive reserve meta-analysis (SC5-B) mentioned in the evidence dossier as reserve evidence. Contains the early-life vs. late-life cognitive reserve comparison; the article does not cite it because the "reading reduces dementia risk by 18%" formulation cannot be supported (reading is one component of broadly-defined cognitive reserve), but the meta-analysis informs the argument about why early cognitive investment matters.