Mind · In / Out · In · 报告

科学中的 AI:早期洞见

AI in Science: Early Insights

Google ATLAS、Google DeepMind、MIT FutureTech 合作 · Google AI 官方 PDF · 2026-09-15

「验证不可压缩」迄今最硬的一手实证:AI 压掉生成之后,瓶颈整个下移到了验证。

Indigo 的结论

「验证不可压缩」迄今最硬的一手实证:AI 压掉生成,瓶颈下移到验证和物理实验,89% 的省时者拿超过一成的时间去核查 AI 的输出。但这是 Gemini 自证加问卷自报,数字当 Google 口径的早期快照。

怎么读这篇 Google 和 DeepMind 自己发的机构报告,替自家说话的成分明确:秀 AI 做科学,秀 Gemini 的采用。但数据是真的、量大,1500 万次 Gemini 交互加 637 位科学家的问卷,而且罕见地把负面如实写了出来:下游瓶颈、假设积压、转向更安全的问题。数字当 Google 口径的早期快照,方向可信。

需要记住的几件事

  1. 瓶颈下移到验证:AI 压掉上游的生成,约束移到验证和物理实验;41% 假设积压;89% 的省时者花一成以上时间核查 AI 输出。
  2. LLM 与专用模型是经济互补,最细一层几乎不重叠,指向「LLM 调度专用模型」的系统:价值在调度、专用模型和数据。
  3. 49% 被推向更安全、更渐进的问题,对比 28% 更敢冒险:可验证的领域在加速,代价可能是少碰冒险问题。
  4. 每周省近 7 小时、主要投回研究、十分之八产出更高,但这是问卷自报的感知,不是实测。
  5. 三处打折:只看 Gemini,问卷是自报的非概率样本,分类管线最细一层准确率只有 28%。方向可信,数字当快照。

拆解 · 6 步

  1. 01

    摘要:四个发现

    三个数据源:1500 万次 Gemini 交互、2600 多个专用模型、600 多位科学家。发现:采用广泛、两类模型互补、每周省近 7 小时、瓶颈下移到验证。 读这一段原文 →

  2. 02

    为什么做,怎么做

    乐观和担忧都缺实证;报告把日志、模型清单和问卷映射到同一套科学任务分类上,从能力、采用一路追到对科研方式的影响。 读这一段原文 →

  3. 03

    采用广泛,两类模型互补

    科学类职业用 AI 的可能是就业基线的 2.7 倍;LLM 和专用模型越细分越不重叠;约四分之三的人省了时间,平均每周近 7 小时。 读这一段原文 →

  4. 04

    瓶颈下移到验证

    超过十分之四的人说主要约束移到了下游;41% 的人假设积压;省下的时间有相当一部分拿去核查 AI 输出;49% 被推向更安全的问题。 读这一段原文 →

  5. 05

    从略的数据,和作者自认的局限

    只看 Gemini、不含企业数据;模型清单不具代表性;问卷可能有选择偏差;分类有误差;只是 2026 年的快照,只能说明相关,不能说明因果。 读这一段原文 →

  6. 06

    讨论:验证税与第 37 手

    AI 工具在按任务分工;约一半人被推向更安全的问题;验证税很高;下一步可能是 LLM 调度专用模型,人负责选问题、定证据标准。 读这一段原文 →

什么会让我改口

没有廉价验证器的开放式科研,在物理验证没跟上的情况下,生产率和发现的斜率照样拐上去。

怎么读这篇

Google 和 DeepMind 自己发的机构报告,替自家说话的成分明确:秀 AI 做科学,秀 Gemini 的采用。但数据是真的、量大,1500 万次 Gemini 交互加 637 位科学家的问卷,而且罕见地把负面如实写了出来:下游瓶颈、假设积压、转向更安全的问题。数字当 Google 口径的早期快照,方向可信。

拆解 · 6 步
  1. 摘要:四个发现
  2. 为什么做,怎么做
  3. 采用广泛,两类模型互补
  4. 瓶颈下移到验证
  5. 从略的数据,和作者自认的局限
  6. 讨论:验证税与第 37 手
01

摘要:四个发现

三个数据源:1500 万次 Gemini 交互、2600 多个专用模型、600 多位科学家。发现:采用广泛、两类模型互补、每周省近 7 小时、瓶颈下移到验证。

摘要

科学进步是经济增长和繁荣的关键驱动力。人们对 AI 给科学带来的影响既非常兴奋,也有担忧,但到目前为止数据很少。我们从三个数据源给出早期洞见:1500 万次 Gemini 交互的样本、覆盖各学科的 2600 多个专用 AI 模型清单,以及对 600 多位科学家的问卷调查。我们把这些数据映射到一套新的科学任务分类体系上,研究科学家如何使用 AI。主要发现有四条。第一,采用广泛、覆盖面广:科学家使用 AI 的程度高于大多数其他职业;专用 AI 模型覆盖的学科很广,被引用次数很高;接近一半受访科学家报告每天都在使用某种形式的 AI。第二,我们发现证据表明,LLM(以 Gemini 的使用为代理)和专用模型是互补的:LLM 用于通用分析、写代码和撰写论文,专用模型则提供特定领域的预测、数据生成和分类。第三,科学家报告使用 AI 带来了很大的生产率提升:每周节省将近 7 小时,这些时间主要被重新投入到更多研究中。最后,我们表明 AI 已经在改变科研过程。随着科研的某些阶段变得更容易,瓶颈向下游转移。科学家报告未经检验的假设越积越多,对产出进行核查的需求也很大。我们的发现表明,AI 在提高科研生产率方面潜力巨大。然而,与其他行业一样,它的最终影响将取决于复杂的任务相互依赖,以及为消除新出现的瓶颈所做的投资。

02

为什么做,怎么做

乐观和担忧都缺实证;报告把日志、模型清单和问卷映射到同一套科学任务分类上,从能力、采用一路追到对科研方式的影响。

1 引言

科学发现是长期进步和经济增长的主要驱动力之一(Romer, 1990; Mokyr, 1992; Aghion 和 Howitt, 1992; Jones, 1995; Akcigit 等, 2021)。然而近年来,人们担心研究生产率在下降,「好点子越来越难找」(Bloom 等, 2020; Park 等, 2023)。AI 的快速发展,让人们乐观地认为这项技术可能扭转这一趋势、加快科学进步(Wang 等, 2023; Amodei, 2024; Curto Millet, 2025; Hassabis 和 Manyika, 2026; Agrawal 等, 2026)。DeepMind 的 AlphaFold 2 这样的专用 AI 模型(Jumper 等, 2021; Hassabis, 2022)已经被公认为诺贝尔奖级别的突破。与此同时,通用 LLM 正在迅速渗入日常的研究工作流程(Van Noorden 和 Perkel, 2023; Wiley, 2025)。

如果 AI 是一种「发明方法的发明」(IMI)(Aghion 等, 2017; Cockburn 等, 2018; Crafts, 2021; Cunningham 等, 2026),它就会让生产率增长长期持续加速。AI 对经济增长和进步的潜力关系重大,因为不断增加的政府债务和人口老龄化,正威胁着美国过去 150 年人均年增长约 2% 的平衡增长路径(Jones, 2026);AI 在科学中的作用,可能是克服这些逆风、加快增长的关键。

然而,对 AI 提升科学进步潜力的乐观并非人人都有。有人担心,AI 生成内容、幻觉和「理解的错觉」的泛滥,会让同行评审和验证不堪重负(Birhane 等, 2023; Messeri 和 Crockett, 2024)。另一些人认为,在现有文献上优化出来的模型,会把研究者引向渐进的、熟知的范式,而不是高风险、高回报的发现(Duede, 2025);或者,高昂的算力和资本门槛会拉大资金充裕的实验室与资源不足的机构之间的差距(Thompson 等, 2022; Ahmed 等, 2023; Besiroglu 等, 2024)。也有人担心,AI 可能加速产出只是看起来像科学的低质量研究(Luo 等, 2025; Gyevnár 等, 2026)。但到目前为止,这些争论几乎没有实证证据。

以往关于 AI 采用及其对科学影响的学术研究,主要依赖公开的科学计量记录,比如论文发表和引用趋势、专利申请,以及 LLM 辅助写作留下的痕迹(Kusumegi 等, 2025; Renault 等, 2026; Trišović 等, 2025; Duede 等, 2024)。这些都是最终产出,有滞后,对采用情况的追踪也不完美。前沿 AI 实验室掌握着研究者实际如何使用模型的实时数据,但基于这类日志的研究,到目前为止关注的是整体的劳动力市场格局(Chatterji 等, 2025; Appel 等, 2025; Handa 等, 2025; Iscenko 等, 2026),对科学关注相对很少。

本文开启一个关于 AI 对科学影响的新研究项目,结合了三个新数据源的洞见:1500 万次 Gemini 交互的样本;覆盖各学科(从蛋白质结构预测到材料发现和天气预报)的 2600 多个专用 AI 模型的文献计量数据;以及对 600 多位科学家的问卷调查。我们把三个数据源都映射到 MIT FutureTech 开发的一套新的科学任务分类体系上,它能在 OpenAlex 的各个科学子领域中细到任务层面(Emmens 等, 2026)。

分析揭示了四个发现。第一,科学家在 AI 使用上领先于其他职业;相对于其就业占比,科学类职业在 Gemini 使用中占比偏高。接近一半受访科学家报告每天都把 AI 作为工作流程的一部分。但只看 LLM 的使用会漏掉很大一块:受访科学家使用专用模型的程度,和使用编程助手差不多。这些模型在学术工作中也被大量引用。科学中的 AI 采用在地理上也很集中,使用量随一个国家的科研人员规模而增长。第二,LLM 和专用模型是经济上的互补品,两者分工明确:前者主要广泛用于写代码和写作等各类任务,后者用于更专门的生成、预测和模拟,两类模型的任务重叠有限。第三,科学家报告 AI 带来了可观的生产率提升,平均每周节省将近 7 小时,这些时间大多被投回研究。研究者还报告获得了更多跨领域、跨学科的洞见。第四,虽然有些任务(比如定量和计算工作)可能被加速了,但科学产出的瓶颈转移到了下游的物理实验和验证。科学家报告未经检验的假设在积压,大量时间花在核查 AI 的输出上,并且出现了转向更安全问题的倾向。

再详细介绍一下方法。我们的第一个数据源来自 Google ATLAS 1.0 项目(Iscenko 等, 2026),包括会话界面(Gemini App 和 AI Mode)以及 Gemini API 上约 1500 万次匿名交互。科研工作通过一个三阶段流程分离出来:首先,一个分类器去掉与工作无关和教学类的交互;然后,一个过滤器把样本限定在科研集中的细分职业;最后,由于这些职业里常有非科学家,我们基于 OECD 的弗拉斯卡蒂手册(2015)和英国研究与创新署(2025)构建了一个新的科学分类器,根据交互内容判断它们是否可能属于科学家的工作流程。这样留下了 36 万次科学交互。

但如何把 LLM 交互映射到具体任务上?职业任务(O*NET)和时间使用(ATUS)都有标准分类体系,但用在科学研究上是不够的。「分析数据」「写报告」这样笼统的类别,抹平了定义科研的那些领域特定活动,从提出假设到设计分子构建体、优化检测方法。衡量 AI 对科学的影响,需要一个共同的尺度,分别刻画研究者在各学科中做什么,以及其中哪些任务得到了 AI 工具的帮助、赋能或自动化。为此,我们用自己定制的分层发现与分类引擎「观察聚类与分类体系组织」(OCTO,Iscenko 等 (2026)),把科学类 LLM 交互映射到两套基于科学的分类体系上:OpenAlex 科学学科,以及 MIT FutureTech 新任务分类体系的三个层级。

第二个数据源是新编制的一份清单,收录 2012 年以来发布的 2600 多个专用模型。它们来自许多学科,包括生物学(例如预测蛋白质折叠的 AlphaFold(Hassabis, 2022))、材料科学(例如预测稳定晶体结构的 GNoME(Merchant 等, 2023))、化学(例如用于设计无机化合物的生成式扩散模型 MatterGen(Zeni 等, 2025))等。我们把在网络资源、出版物和代码仓库上自下而上的深度智能体搜索,与 Epoch 的 AI 模型数据库(Epoch AI, 2026)结合起来编成这份清单,再用 OpenAlex 的元数据加以充实。然后用 OCTO 提取每个模型所支持的研究任务(例如预测蛋白质结构),并像处理日志数据那样,映射到 FutureTech 和 OpenAlex 分类体系中相应的任务上。这样,LLM 和专用模型的使用就被放到了同一个尺度上,让我们能从实证上研究这两类模型之间的分工。我们追踪了这些模型的 46 万条独立引用,来追溯它们在各科学领域间的扩散和知识流动。

最后一个数据源,是我们于 2026 年 7 月到 8 月对美国和英国 637 位在职研究者所做的原创问卷调查。样本包括学术界的课题负责人(PI)、业界研究人员和青年科学家,覆盖四大科学领域。问卷让我们能填补日志数据和发表数据都说明不了的空白,比如自我感知的节省时间、瓶颈、优先事项,以及科研工作流程的变化。三个数据源合在一起,让我们能从能力、到采用、再到它对科研方式的影响,一路追踪 AI。

03

采用广泛,两类模型互补

科学类职业用 AI 的可能是就业基线的 2.7 倍;LLM 和专用模型越细分越不重叠;约四分之三的人省了时间,平均每周近 7 小时。

从 LLM 交互数据看,AI 在科学中的渗透很广。科学领域对 LLM 的使用,相对于其在美国就业中的占比是偏高的:例如在美国,标准职业分类(SOC)第 19 大类(包括生命科学、物理科学和社会科学类职业)使用 AI 的可能性,大约是就业基线的 2.7 倍。计算机科学和工程等更偏定量的领域在采用上领先,但政治学、医学和农业科学等其他专业也是 AI 的重要使用者。专用模型存在于大多数学科,但占比因研究领域差别很大。它们的被引用次数也很高:发布专用模型的论文里,接近一半进入了各自领域被引用最多的前 1%。科学中的使用并不稀疏,接近一半受访科学家把 AI 当作日常工作流程的一部分。从采用的地域差异看,一个国家占全球研究者份额每增加 1 个百分点,它在科学类 LLM 交互中的份额就增加 0.9 个百分点;专用模型的开发集中在美国、中国、英国、欧盟、韩国和加拿大。一些低收入国家大量引用这些模型,却不开发它们,这表明即使在没有能力建造的地方,开放模型也能广泛扩散。

重要的是,LLM 和专用模型似乎是经济上的互补品,彼此补足、可能还相互增强对方的能力。在任务分类体系的最高层,两类模型看起来很像,都以定量数据分析和建模为主;但一旦把任务进一步细分,相似性就消失了。两类模型在各任务上的使用相关性,从最粗的任务层级到最细的层级单调下降,到最细一层几乎没有重叠。LLM 广泛用于写作和调试研究代码、统计分析、文献综述、协助起草文档等任务。专用模型在健康和生命科学中相对更常见,用于预测疾病结局、构建分子、运行复杂模拟。因此,这两类模型似乎在科学知识的生产中相互强化,这限制了目前一种技术能在多大程度上替代另一种。问卷数据也说明了同样的情况:科学家对 LLM 和专用模型都有使用。

科学家还报告看到了 AI 带来的真实生产率提升。大约四分之三的研究者报告节省了时间,平均每周节省将近 7 小时。这些时间回到了研究中,变成了更多产出。大约十分之八的人报告过去三年实验室产出更高,89% 预计还会继续增加。收益也不只在速度上:大约 68% 的科学家报告更容易获得其他学科的洞见,这与人们的一个期望相符,即 AI 也许能减轻把研究者推向越来越窄专业的「知识负担」(Jones, 2009)。

04

瓶颈下移到验证

超过十分之四的人说主要约束移到了下游;41% 的人假设积压;省下的时间有相当一部分拿去核查 AI 输出;49% 被推向更安全的问题。

但瓶颈依然存在,这可以解释为什么局部的生产率提升还没有转化为新发现和新应用的爆发。超过十分之四的受访科学家报告,过去两年里他们的主要约束已经转移到下游:实验室执行、临床验证和实地数据采集。AI 工具可能增加了可行理论和假设的数量,却不一定增加了检验它们的手段。这导致 41% 的人报告未经检验的假设积压越来越多。验证吃掉了红利的一大块:节省了时间的人里,89% 把其中超过十分之一的时间花在核查 AI 的输出上,46% 花了超过四分之一。从长远看最令人担忧的是,49% 的科学家说 AI 把他们推向更安全、更渐进的项目,那里基准已经建立、结果可靠;而说 AI 让他们敢于挑战更有风险问题的只有 28%。如果 AI 主要降低的是渐进式工作的成本,它可能会增加论文的数量,却不能实质性地推进科学前沿。

本文是对 AI 此刻如何影响科学的一次早期快照。LLM 与专用模型之间的分工,指出了未来可能的方向:一个 LLM 调度者规划实验,调用 AlphaFold 或材料模型之类的专用模型,并解读结果。这样的系统还在开发中。要把它做出来并兑现其潜力,需要解决科学家指出的那些瓶颈,把投资引向物理实验和验证能力、验证工具,以及能让 AI 去碰更有风险而不是更安全问题的数据和基准。

本文其余部分安排如下。第 2 节详细介绍数据源和测量方法。第 3 节用 Google ATLAS 数据(Iscenko 等, 2026),经 MIT FutureTech 科学任务分类体系(Emmens 等, 2026)映射到科研工作流程和领域上,展示 LLM 在科学中采用的一些早期信号。第 4 节分析专用科学 AI 模型所在的领域、与下游任务的关联,以及溢出效应的早期信号。第 5 节给出关于研究者时间分配、感知到的瓶颈和预期生产率影响的问卷结果。最后,第 6 节说明局限并讨论结果。

05

从略的数据,和作者自认的局限

只看 Gemini、不含企业数据;模型清单不具代表性;问卷可能有选择偏差;分类有误差;只是 2026 年的快照,只能说明相关,不能说明因果。

第 2–5 节:数据与结果

[编者注] 第 2 至第 5 节详细介绍三个数据源和测量方法,并用图表呈现结果;关键数字已在上面的引言中概述。这几节此处从略,请看原文 PDF。

6.1 局限

我们必须承认数据、测量和范围上的若干局限。第一,我们的日志分析(第 3 节)所考察的 1500 万次交互样本,完全取自 Google 的 ATLAS 数据集(Iscenko 等, 2026)。企业数据不在本次研究之内。推广结论时应当谨慎,因为 Gemini 用户可能偏向某些特定学科,或者使用模式与其他闭源和开放模型的用户不同。例如,健康科学等监管更严领域的科学家,可能相对更倾向于使用企业账户或专门的解决方案,而这些我们看不到。还值得指出,我们的分析不包括 Antigravity 这类智能体式 AI 工具,科学家可以在其中借助 Gemini 和其他 LLM 调用专用模型的能力(例如 Applebaum 等, 2026);把智能体数据纳入分析,是我们研究项目重要的下一步。

第二,我们的专用模型清单是收录科学领域大量特定领域 AI 模型的第一步,但目前不应被视为有代表性。例如,我们目前聚焦于那些更容易通过网络搜索、文献计量平台和代码仓库发现的知名模型,而遗漏了在单项研究中基于开源框架构建、或微调开放权重模型得到的定制模型,以及商业模型。专用模型清单及基于它的引用分析也天然滞后,这意味着我们捕捉的是一个动态格局中的快照。随着 AI 辅助论文的同行评审标准发生变化,长期的引用模式可能会有不同的演变。

我们的第三个数据源,也就是科学家问卷,可能存在选择偏差。尤其是,那些已经对 AI 充满热情或经常使用 AI 的科学家,可能更愿意回答关于科学中 AI 的问卷。这可能会高估真实的采用率和平均每周节省的时间。我们的样本量也限制了我们在一些关键问题上探究学科差异的能力,比如下游验证(例如物理实验和临床试验)中瓶颈的转移,这可能因学科而异。

[脚注 26] 化学和医学在物理验证上可能面临很大的资本和时间门槛,而计算机科学或数学这样的学科可以完全在计算机中(通过计算)验证假设。因此,我们的总体结果可能高估了「湿」科学的物理瓶颈,相对于「干」的计算科学而言。

我们依靠 MIT FutureTech 的科学任务分类体系,把对 LLM 交互日志和专用模型清单数据的分析整合起来。从招聘启事中推导任务分类体系有其局限。在科学研究与更广泛的理工科经济之间划清界线并不容易,所用的科研岗位招聘启事(以及由此提取的任务)中难免有漏判和误判。随着我们对刻画各科学子领域所执行任务的能力做进一步精调,这套分类体系会继续改进。

我们用来把日志以及从专用模型中提取的任务归入 MIT FutureTech 科学任务分类体系的映射过程,也有其局限。在不知道用户身份的情况下,用 LLM 来区分哪些任务属于科学(以及所有下游的领域和任务分类),我们的分类器目前存在这类做法固有的问题。例如,我们明确只在对话内容似乎表明是科学工作时,才(由 LLM 来分类)打上「科学」标签。因此,我们可能会在不同任务类型上有差别地、系统性地漏掉那些隐含的或表述松散的科学工作(例如安排会议时语境可能较少,而为论文做数据分析时语境较多)。

此外,撇开把「野外」那些嘈杂的科研活动描述归入一个大型分类体系中模糊、有时还相互重叠的标签所带来的分类误差不谈,我们的分类过程把每条日志或模型任务只映射到一个 MIT 任务上,忽略了这样的事实:例如,某些专门任务可能与多项下游科研活动相关,或者一次 Gemini 对话中可能包含多项任务。科学家也可能为开源模型「找到自己的用法」,把它们用于开发者没有预料到、也没有在模型说明中提到的任务。我们计划在今后的工作中处理这些局限,改进分类和过滤流程。

我们还要指出,本文捕捉的是 2026 年上半年到年中这段时间 AI 采用情况的快照,这可能影响我们的部分发现。例如,我们发现这段时间里 LLM 和专用模型之间分工明确,但这可能会变。能力越来越强、能解决数学、基因组学或生命科学等领域难题的新一代前沿 LLM(以及专用模型)的出现(例如 Anthropic, 2026; Callaway, 2026; OpenAI, 2026; Romera-Paredes 等, 2023),让不同类型模型之间的分工未来如何演变变得不确定。我们计划在后续各期工作中追踪这一点。

最后,我们的结果确立的是相关关系,而不是因果关系。需要今后的工作来分离出使用 AI 工具对科研生产率的因果影响。

06

讨论:验证税与第 37 手

AI 工具在按任务分工;约一半人被推向更安全的问题;验证税很高;下一步可能是 LLM 调度专用模型,人负责选问题、定证据标准。

6.2 讨论

尽管有这些局限,我们认为我们的发现为了解 AI 如何重塑科学提供了一次重要的早期观察。我们的证据指向的,不是劳动力被简单替代、或自动化发现带来生产率失控飙升的故事,而是一幅更细致的图景。

AI 的使用遍及各个科学学科,所以科学家在 AI 使用上相对于其他职业占比偏高,也许并不意外。当然,科学中 AI 的扩散,与计算资本或无形资本的其他分布(以及投资能力,例如 Brynjolfsson 等 (2021))密切相关。Gemini 的使用量大致与各国科研人员数量成比例,而专用 AI 模型的开发主要集中在少数发达经济体。前沿的科学 AI 往往需要算力可得性和特定领域人力资本这类互补条件。这意味着一个风险:如果没有政策行动,资源不足或发展中的地区可能会落后。AI 增强科学的前沿,可能需要投资来缓解全球科研能力的差距。

在任务分析上,专用 AI 模型和 Gemini 的使用看起来相似,两者的主要用途都是定量建模和数据分析。但当我们深入到具体的分析类型,就开始看到模型专门化的证据,以及科学生产函数内部的分工。专用模型的作用类似一种特定领域的资本:生成合成数据、预测分子性质、执行模拟,这些任务看起来(至少目前)更适合由专用模型而不是通用 LLM 来完成。与此同时,通用模型承担了更广泛的任务,比如写代码、综合文献和日常事务。就像研究者会按学科或工作流程环节分工一样,AI 工具可能也在专门化。

在被采用的地方,AI 似乎带来了可观的生产率红利。研究者报告了可观的净节省时间,并把它们投到其他项目上。然而,AI 在科学中带来的这些变化不会只发生在「做多少」的层面。大约一半受访研究者报告,AI 把他们引向更安全、更渐进的问题,那里有干净的数据和基准。只有 28% 转而去做风险更高的项目,尽管有证据表明,最重要的科学进步依赖高风险的探索(Azoulay 等, 2011)和新颖的假设(Uzzi 等, 2013)。AI 模型究竟会把研究者引向更多渐进式工作、还是促成「登月」项目,抑或两者兼有,还有待观察;今后的工作应当研究 AI 使用与科学家风险偏好之间的关系。也许 AI 会带来一种「路灯效应」(Hoelzemann 等, 2024; Nagaraj 和 Tranchero, 2023):做渐进式工作的成本相对于高风险项目下降了。至少在短期内,被 AI 变得容易的渐进式工作,在全部研究产出中的比重可能会上升。另一方面,科学家报告 AI 促成了更多跨学科工作,可能像一个翻译者,帮助研究者更有效地重组想法(Fang 和 Evans, 2026),并帮助克服知识负担(Jones, 2009)。

[脚注 28] 这不一定是坏结果。例如,加深对疾病机理理解的渐进式「常规科学」,也在为药物开发积累证据(McNamee 等, 2017)。

尽管任务层面的时间节省很普遍,宏观层面的科学发现速度仍受工作流程中任务依赖关系的约束。这些隐藏的科研「组织」互补性,与 AI 对一般工作的影响类似(Demirer 等, 2026; Gans 和 Goldfarb, 2026; Kremer, 1993)。例如,计算建模和数据分析也许被 AI 工具加速了,专用 AI 模型也许提供了更有希望去测试的生物结构,但科研工作这一整包任务中剩下的部分更难规模化,会吸收掉节省下来的时间。科研工作的这种捆绑和任务互锁意味着,更广泛的生产率提升可能需要重新划定岗位边界,而某些类型的工作捆绑得比其他工作更紧(Garicano 等, 2026)。例如,数据分析生产率的提升,可能把排队的工作推向难以自动化的物理阶段,使之成为瓶颈(Jones, 2025)。

此外,AI 的输出并不一定是「拿来即用」的。在报告节省了时间的研究者中,接近 90% 把节省时间中相当一部分用于核查 AI 的输出,约 46% 花在这上面的时间超过节省时间的四分之一。就像在软件工程中一样,采用 AI 工具的科学家可能会被推向审计、调试和质量筛查。尤其在科研工作中,正确是最重要的。降低提出假设、生成代码、分析或测试的边际成本,会让核实结果(或者也许是人的判断,例如 Agrawal 等 (2019))的互补性边际价值变得很高。这种高昂的「验证税」,很可能源于科学中可靠、正确的产出价值很高。

我们记录的这种分工,也暗示了不久之后的科学家可能是什么样子。今天,研究者在工具之间手动切换:用聊天机器人写代码和起草,用 AlphaFold 或材料模型做预测,中间夹着电子表格或实验记录本。下一步可能是一个 LLM 直接调度这些专用模型:规划实验,在每一步调用合适的模型,交叉核对各个输出,再把一个候选答案交给科学家判断。人的工作转向选择问题、设定证据标准、决定什么值得付出物理实验的成本。我们的数据表明,我们还没走到那一步:验证仍在消耗 AI 节省下来的大部分时间,未经检验的假设积压在增加。归根到底,人们的期望不止于节省时间。2016 年 AlphaGo(Silver 等, 2016)对李世石下出第 37 手时,那一步棋按它自己的模型估计,人类大约一万次里才会下一次,而它赢了。更先进的 AI 系统在科学中的前景,就在于它们终将下出那样的棋:做出超越人类研究者自己会去尝试的极限的发现。

我们的工作对科学中 AI 工具的采用、其潜在影响以及可能限制这种影响的瓶颈,给出了一些初步的测量。随着这些局限浮现,跨学科工作的新机会、以及计算密集型领域的新研究方向也随之出现。科学生产与其他各类生产有许多共同特征。科学家对 AI 工具的快速整合和采用,加上突破性发现的巨大重要性,使科学成为理解 AI 所带来机遇的理想领域。通过长期持续这项工作,我们可以看着实验者做实验,从科学的成功和失败中学习,释放 AI 在科研以及经济其他领域带来的生产率提升。

判断收口延伸

Indigo 的结论

「验证不可压缩」迄今最硬的一手实证:AI 压掉生成,瓶颈下移到验证和物理实验,89% 的省时者拿超过一成的时间去核查 AI 的输出。但这是 Gemini 自证加问卷自报,数字当 Google 口径的早期快照。

需要记住的几件事

  1. 瓶颈下移到验证:AI 压掉上游的生成,约束移到验证和物理实验;41% 假设积压;89% 的省时者花一成以上时间核查 AI 输出。
  2. LLM 与专用模型是经济互补,最细一层几乎不重叠,指向「LLM 调度专用模型」的系统:价值在调度、专用模型和数据。
  3. 49% 被推向更安全、更渐进的问题,对比 28% 更敢冒险:可验证的领域在加速,代价可能是少碰冒险问题。
  4. 每周省近 7 小时、主要投回研究、十分之八产出更高,但这是问卷自报的感知,不是实测。
  5. 三处打折:只看 Gemini,问卷是自报的非概率样本,分类管线最细一层准确率只有 28%。方向可信,数字当快照。

放回主线

证实

验证不可压缩 这条线迄今最硬的一手实证支柱:Google 用 1500 万次交互和 600 多位科学家量出「压上游、瓶颈下移、验证吃红利」。

冲突

可验证域能否泛化 49% 转向更安全的问题,是负面苗头:可验证的领域被攻下,不等于前沿被推进。

证实

你拥有的不是模型 「LLM 调度专用模型,加上数据和基准」,把价值放在调度、独有的专用模型和数据上,而不是通用 LLM。

补充

开放权重的安全政治学 低收入国家大量引用却不开发专用模型,是开放模型广泛扩散的一手数据。

证实

DeepMind Co-Scientist 同一种诚实:AI 帮科研是真的,但终审验证和物理世界恰是瓶颈。

补充

Andrew Ng:工作末日不来因 AI 还不够好 Ng 讲宏观的互补,这份报告给出科学场景的微观数据:两类模型互补,省下的时间投回研究。

什么会让我改口

没有廉价验证器的开放式科研,在物理验证没跟上的情况下,生产率和发现的斜率照样拐上去。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Report

AI in Science: Early Insights

Google ATLAS, Google DeepMind, MIT FutureTech · ai.google · 2026-09-15

The hardest first-hand evidence yet that verification is incompressible: once AI crushes generation, the whole bottleneck moves to verification.

Indigo's conclusion

The hardest first-hand evidence yet that verification is incompressible: AI crushes generation, the bottleneck moves to verification and physical experiment, and 89% of time-savers spend over a tenth of the saved time checking AI output. But it's Gemini grading itself plus a self-reported survey; read the numbers as an early Google snapshot.

How to read this An institutional report from Google and DeepMind with an obvious interest: show off AI for science and Gemini adoption. But the data is real and large, 15 million Gemini interactions plus a survey of 637 scientists, and it's unusually honest about the bad news: downstream bottlenecks, a backlog of untested hypotheses, a drift toward safer problems. Treat the numbers as an early Google-side snapshot; trust the direction.

What to remember

  1. The bottleneck moves to verification: AI crushes upstream generation, constraints move to verification and physical experiments; 41% report a hypothesis backlog; 89% of time-savers spend over a tenth of it checking AI output.
  2. LLMs and specialized models are economic complements with almost no overlap at the finest level, pointing to LLMs orchestrating specialized models: value lies in orchestration, specialized models and data.
  3. 49% are pushed toward safer, incremental problems versus 28% toward riskier ones: verifiable domains speed up, possibly at the cost of risky questions.
  4. Nearly 7 hours saved a week, mostly reinvested in research, 8 in 10 report higher output; but it's self-reported perception, not measurement.
  5. Three discounts: Gemini only, a self-reported non-probability survey, and a classifier only 28% accurate at the finest level. Trust the direction; treat the numbers as a snapshot.

Breakdown · 6 steps

  1. 01

    Abstract: four findings

    Three data sources: 15 million Gemini interactions, 2,600+ specialized models, 600+ scientists. Findings: broad adoption, complementary model types, nearly 7 hours saved a week, bottlenecks shifting to verification. Read this part →

  2. 02

    Why and how

    Both optimism and worry lack evidence; the report maps logs, a model inventory and a survey onto one taxonomy of scientific tasks, following AI from capability to adoption to its effect on how science gets done. Read this part →

  3. 03

    Broad adoption, complementary models

    Science occupations are 2.7 times as likely to use AI as the employment baseline; LLMs and specialized models overlap less the finer the tasks; about three quarters save time, nearly 7 hours a week on average. Read this part →

  4. 04

    The bottleneck moves to verification

    More than 4 in 10 say their main constraint has moved downstream; 41% report a backlog of hypotheses; much of the saved time goes to checking AI output; 49% are pushed toward safer problems. Read this part →

  5. 05

    Omitted data, and the authors' own limits

    Gemini only, no enterprise data; the model inventory isn't representative; the survey may be self-selected; classification has errors; it's a 2026 snapshot showing association, not causation. Read this part →

  6. 06

    Discussion: the verification tax and move 37

    AI tools are dividing the labor by task; about half are pushed toward safer problems; the verification tax is high; next may be an LLM orchestrating specialized models, with people choosing questions and setting the standard of evidence. Read this part →

What would change my mind

open-ended science with no cheap verifier showing the same inflection in productivity and discovery without physical verification catching up.

How to read this

An institutional report from Google and DeepMind with an obvious interest: show off AI for science and Gemini adoption. But the data is real and large, 15 million Gemini interactions plus a survey of 637 scientists, and it's unusually honest about the bad news: downstream bottlenecks, a backlog of untested hypotheses, a drift toward safer problems. Treat the numbers as an early Google-side snapshot; trust the direction.

Breakdown · 6 steps
  1. Abstract: four findings
  2. Why and how
  3. Broad adoption, complementary models
  4. The bottleneck moves to verification
  5. Omitted data, and the authors' own limits
  6. Discussion: the verification tax and move 37
01

Abstract: four findings

Three data sources: 15 million Gemini interactions, 2,600+ specialized models, 600+ scientists. Findings: broad adoption, complementary model types, nearly 7 hours saved a week, bottlenecks shifting to verification.

Abstract

Scientific progress is a key driver of economic growth and prosperity. There is great excitement - but also concerns - about the impacts of AI on science, but so far little data. We provide early insights on this from three data sources: a sample of 15 million Gemini interactions, an inventory of over 2,600 specialized AI models across disciplines, and a survey of over 600 scientists. We map these data to a new taxonomy of scientific tasks to study how scientists are using AI. Four main findings emerge. First, we find broad adoption and coverage: scientists use AI more than most other occupations. Specialized AI models have broad disciplinary coverage and are highly cited. Nearly half of the scientists surveyed report using some form of AI every day. Second, we document evidence that LLMs (proxied through Gemini usage) and specialized models act as complements—LLMs are used for general analysis, coding, and manuscript preparation, while specialized models provide domain-specific predictions, data generation and classification. Third, scientists report large productivity gains from using AI: a saving of nearly 7 hours per week, time which is primarily re-invested in more research. Finally, we show that AI is already changing the scientific process. As some stages of scientific research become easier, bottlenecks shift downstream. Scientists report an increased backlog of untested hypotheses and substantial demand for output verification. Our findings suggest that AI holds significant potential to increase scientific productivity. However, as with other sectors, its ultimate impact will be governed by complex task interdependencies and investment into the elimination of emerging bottlenecks.

02

Why and how

Both optimism and worry lack evidence; the report maps logs, a model inventory and a survey onto one taxonomy of scientific tasks, following AI from capability to adoption to its effect on how science gets done.

1 Introduction

Scientific discovery is one of the primary drivers of long-run progress and economic growth (Romer, 1990; Mokyr, 1992; Aghion and Howitt, 1992; Jones, 1995; Akcigit et al., 2021). Yet concerns of decreased research productivity or “ideas getting harder to find” have been raised in recent years (Bloom et al., 2020; Park et al., 2023). The rapid development of AI has fueled optimism that the technology could reverse this trend and speed up scientific progress (Wang et al., 2023; Amodei, 2024; Curto Millet, 2025; Hassabis and Manyika, 2026; Agrawal et al., 2026). Specialized AI models, such as DeepMind’s AlphaFold 2 (Jumper et al., 2021; Hassabis, 2022) have already been recognized as Nobel-caliber breakthroughs. Meanwhile, general-purpose LLMs are rapidly diffusing across day-to-day research workflows (Van Noorden and Perkel, 2023; Wiley, 2025).

If AI acts as an “IMI”, an Invention of a Method of Invention (Aghion et al., 2017; Cockburn et al., 2018; Crafts, 2021; Cunningham et al., 2026), it would create a durable acceleration to productivity growth over time. Much is riding on AI’s potential for economic growth and progress, as increasing government debt and aging populations threaten to disrupt the balanced growth path of roughly 2 percent annual expansion per capita over the last 150 years in the US (Jones, 2026); AI’s role in science may be key for overcoming these headwinds and accelerating growth.

Yet the optimism around AI’s potential to increase scientific progress is not universal. Some worry that a deluge of AI-generated content, hallucinations, and “illusions of understanding” will strain peer review and validation (Birhane et al., 2023; Messeri and Crockett, 2024). Others argue that models optimized on the existing literature will steer researchers toward incremental, well-known paradigms rather than high-risk, high-reward discoveries (Duede, 2025), or that steep compute and capital requirements will widen the gap between well-funded laboratories and under-resourced institutions (Thompson et al., 2022; Ahmed et al., 2023; Besiroglu et al., 2024). Some also worry that AI might accelerate the production of low-quality research that merely appears scientific (Luo et al., 2025; Gyevnár et al., 2026). But these debates have thus far evolved with little empirical evidence.

Prior academic research about AI adoption and its impact on science has primarily relied on open scientometric records such as publication and citation trends, patent filings, and traces of LLM-assisted writing (Kusumegi et al., 2025; Renault et al., 2026; Trišović et al., 2025; Duede et al., 2024). These are finalized outputs that arrive with a lag and track adoption imperfectly. Frontier AI labs hold real-time telemetry on how researchers actually use models, but the studies built on such logs have so far focused on aggregate labor market patterns (Chatterji et al., 2025; Appel et al., 2025; Handa et al., 2025; Iscenko et al., 2026), with relatively little focus on science.

This paper inaugurates a new research program on the impact of AI on science by combining insights from three new data sources: a sample of 15 million Gemini interactions, bibliometrics data for more than 2,600 specialized AI models across academic disciplines (from protein structure prediction to materials discovery and weather forecasting), and a survey of more than 600 scientists. We map all three data sources to a new taxonomy of scientific tasks developed by MIT FutureTech, which gives task-level granularity across OpenAlex scientific subfields (Emmens et al., 2026).

The analysis reveals four findings. First, scientists lead other occupations in the use of AI; scientific occupations are overrepresented in Gemini usage relative to their share of employment. Nearly half of surveyed scientists report using AI every day as part of their workflow. But looking at LLM use alone would miss much of the picture. Surveyed scientists use specialized models about as much as coding assistants. These models are also heavily cited in scholarly work. AI adoption in science is also geographically concentrated, with usage scaling with a country’s scientific workforce. Second, LLMs and specialized models are economic complements with a clear division of labor between them: the former are primarily used broadly across tasks including coding and writing, while the latter are used for more specialized generation, prediction, and simulation, with limited overlap in tasks between the two model categories. Third, scientists report substantial productivity gains from AI, with average time saving of just below 7 hours per week. The time is mostly put back into research. Researchers also report greater cross-field and interdisciplinary insights. Fourth, while some tasks—for example, quantitative, computational work—may be accelerated, the bottlenecks to scientific output shift downstream into physical experimentation and validation. Scientists report a backlog of untested hypotheses, substantial time spent verifying AI outputs, and a tilt toward safer questions.

To give more detail on the methodology, our first data source comes from the Google ATLAS 1.0 project (Iscenko et al., 2026) which consists of about 15 million anonymized interactions across the conversational surfaces (Gemini App and AI Mode), and Gemini API. Scientific work is isolated through a three-stage pipeline. First, a classifier removes non-work and educational interactions; then, a filter restricts the sample to the detailed occupations where research is concentrated; finally, as these occupations often contain non-scientists, a new science classifier is built on the OECD Frascati (2015) and UK Research and Innovation (2025) using the content of interactions to classify if they are likely part of a scientist workflow. This leaves us with 360,000 science interactions.

But how does one map LLM interactions to specific tasks? Standard taxonomies exist for occupational tasks (O*NET) and time use (ATUS). But when applied to scientific research, these tools are insufficient. Generic categories like “analyzing data” or “writing reports” flatten the domain-specific activities that define research, from hypothesis formulation to molecular construct design or assay optimization. Measuring AI’s impact on science requires a common metric that separately captures what researchers do across disciplines and which of those tasks AI tools help, enable, or automate. To accomplish this, we used our custom hierarchical discovery and classification engine - Observation Clustering and Taxonomy Organisation (OCTO, Iscenko et al. (2026)) - to map scientific LLM interactions to two science-based taxonomies: the OpenAlex scientific disciplines and the three levels of the new MIT FutureTech task taxonomy.

The second data source comes from a newly-compiled inventory of over 2,600 specialized models published since 2012. These include models from many disciplines including biology (e.g., AlphaFold (Hassabis, 2022) predicting protein folding), material science (e.g., GNoME (Merchant et al., 2023) predicting stable crystal structures), chemistry (e.g., MatterGen (Zeni et al., 2025), a generative diffusion model for inorganic compound design) and more. We assembled the inventory by combining a bottom-up deep agentic search over web sources, publications, and code repositories with Epoch’s AI model database (Epoch AI, 2026), then enriched the output using metadata from OpenAlex. OCTO was then used to extract the research tasks that each model enables (e.g., predicting protein structures), which were then mapped to the respective tasks in the FutureTech and OpenAlex taxonomies, as done with the log data. This exercise puts the LLM and specialized model usage on a common metric, which allows us to empirically study the division of labor between these categories of models. We have tracked 460,000 unique citations to these models to trace diffusion and knowledge flows across scientific fields.

Our final data source is an original survey of 637 active researchers in the US and UK carried out in July-August 2026. The sample included Principal Investigators (PIs) in academia, industry researchers, and early career scientists across four major scientific domains. The survey allows us to fill the gaps that neither log telemetry nor publication data can speak to, such as perceived time saved, bottlenecks, priorities, and changes in the scientific workflow. Together, the three data sources let us follow AI from capability, to adoption, to its effect on how science gets done.

03

Broad adoption, complementary models

Science occupations are 2.7 times as likely to use AI as the employment baseline; LLMs and specialized models overlap less the finer the tasks; about three quarters save time, nearly 7 hours a week on average.

Looking at the LLM interaction data, we find extensive penetration of AI in science. Use of LLMs in science is overrepresented relative to its share of US employment; for example, in the US, the Standard Occupational Classification (SOC) 19 job category roles (containing a set of Life, Physical and Social Sciences occupations) are about 2.7 times more likely to use AI compared to employment baseline. While more quantitative fields such as computer science and engineering lead in adoption, other specialties such as political science, medicine, and agricultural sciences are also significant adopters of AI. Specialized models are present across most disciplines, though the share varies substantially by field of study. They are also highly cited. Nearly half of the papers that introduce specialized models are in the top 1% of citations within their respective fields. Usage in science is not sparse, with almost half of surveyed scientists using AI as part of their daily workflow. Looking at geographic variation in adoption, a one percentage point increase in a country’s share of the world’s researchers is associated with a 0.9 percentage point increase in its share of science LLM interactions, and specialized model development is concentrated in the U.S., China, the UK, the EU, Korea, and Canada. Some lower-income countries cite these models heavily without developing them, which suggests that open models can diffuse widely even where the capacity to build them does not.

Importantly, LLMs and specialized models appear to act as economic complements, filling out and potentially enhancing the capabilities of the other. While at the top level of the task taxonomy the two types of models look alike—quantitative data analysis and modeling dominates both—the similarity disappears once tasks are disaggregated further. The correlation in usage across tasks falls monotonically between the most aggregated and most granular tasks, and at the finest level there is almost no overlap. LLMs are used broadly across tasks such as writing and troubleshooting research code, statistical analysis, literature review, and helping with drafting documents. Specialized models are relatively more common in health and life sciences and are used to predict disease outcomes, engineer molecular constructs, and run complex simulations. Therefore both model families seem to reinforce one another in the production of scientific knowledge, limiting the extent to which one technology can at this point substitute for the other. Survey data tells a similar story: scientists split usage in both LLMs and specialized models.

Scientists also report seeing real productivity gains from AI. Around three quarters of researchers report time savings, with the average amount clocking in at almost 7 hours saved a week. That time goes back into research—into more output. About 8 in 10 report higher lab output over the past three years and 89% expect further increases. The gains extend beyond speed. About 68% of scientists report more access to insights from other disciplines, consistent with hopes that AI may lower the burden of knowledge that pushes researchers into ever-narrower specialties (Jones, 2009).

04

The bottleneck moves to verification

More than 4 in 10 say their main constraint has moved downstream; 41% report a backlog of hypotheses; much of the saved time goes to checking AI output; 49% are pushed toward safer problems.

But bottlenecks remain, which can explain why localized productivity gains have not yet translated into an explosion of new discoveries and applications. More than 4 in 10 of surveyed scientists report that their primary constraint has moved downstream over the past two years, into lab execution, clinical validation, and field data collection. AI tools may have increased the number of viable theories and hypotheses but not necessarily the means to test them. That has led 41% to report a growing backlog of untested hypotheses. Verification absorbs a large share of the dividend: 89% of those who save time spend more than a tenth of it checking AI outputs, and 46% spend more than a quarter. Most concerning for the long run, 49% of scientists say AI pushes them toward safer, more incremental projects where benchmarks are established and results are reliable, against 28% who say it lets them take on riskier questions. If AI mainly lowers the cost of incremental work, it could raise the volume of papers while not meaningfully advancing the scientific frontier.

This paper represents an early snapshot into how AI is affecting science right now. The division of labor between LLMs and specialized models points to where this could go: an LLM orchestrator that plans experiments, calls specialized models such as AlphaFold or a materials model, and interprets the results. That system is still work in progress. Delivering it and realizing its promise requires solving the bottlenecks that scientists identify and directing investment to physical experimentation and validation capacity, to verification tools, and to the data and benchmarks that would let AI take on riskier questions rather than safer ones.

The remainder of this paper is structured as follows. Section 2 details our data sources and measurement. Section 3 shows some early signals of LLM adoption in science using Google ATLAS data (Iscenko et al., 2026) mapped to scientific workflows and fields via the MIT FutureTech Scientific Task Taxonomy (Emmens et al., 2026). Section 4 analyzes the fields, downstream task relevance, and early signals of spillovers of the specialized scientific AI models. Section 5 presents survey findings on researchers’ time allocation, perceived bottlenecks, and expected productivity impacts. Finally, Section 6 outlines our limitations and discusses the results.

05

Omitted data, and the authors' own limits

Gemini only, no enterprise data; the model inventory isn't representative; the survey may be self-selected; classification has errors; it's a 2026 snapshot showing association, not causation.

Sections 2–5: data and results

[Editor's note] Sections 2 to 5 describe the three data sources and the measurement in detail and present the results in figures and tables; the key numbers are summarized in the introduction above. They are omitted here; see the original PDF.

6.1 Limitations

Several limitations in our data, measurement, and scope must be acknowledged. First, our log analysis (section 3) considered a sample of 15 million interactions drawn exclusively from Google’s ATLAS (Iscenko et al., 2026) dataset. Enterprise data was excluded from this investigation. Generalization should be done with caution as Gemini users may skew toward specific scientific disciplines or have distinct use patterns from other proprietary and open models. For example, scientists in more regulated areas like health sciences may be relatively more likely to use enterprise accounts or specialized solutions, which we cannot observe. It is also worth noting that our analysis excludes agentic AI tools like Antigravity where scientists can use Gemini and other LLMs to leverage specialized model capabilities (e.g., Applebaum et al., 2026)—incorporating agentic data into our analysis is an important next step for our research program.

Second, our specialized model inventory is an initial step towards capturing the wealth of domain-specific AI models in science, but at this point should not be considered representative. For example, we currently focus on notable models that are easier to detect through web searches, bibliometric platforms and coding repositories, but miss bespoke models designed during individual studies building on open source frameworks or fine-tuning open weight models, as well as commercial models. The specialized model inventory and the citation analysis building on it are also inherently lagging which means we are capturing a snapshot in a dynamic landscape. It could be that long-term citation patterns evolve differently as peer-review standards for AI-assisted papers change.

Our third data source—the scientist survey—could suffer from selection bias. In particular those scientists who are already enthusiastic about or frequently use AI might be more likely to respond to a survey about AI in science. This could potentially overestimate the real adoption rate and the average weekly time savings. Our sample size also limits our ability to explore disciplinary differences in key questions such as shifting bottlenecks in downstream validation (e.g., physical experimentation and clinical trials), which could vary by discipline.

[Footnote 26] While chemistry and medicine could face large capital/temporal barriers in physical validation, disciplines like computer science or mathematics can validate hypotheses entirely in-silico (computationally). Our aggregate findings, as a result, may over-represent the physical bottlenecks of “wet” sciences compared to “dry” computational sciences.

We rely on MIT FutureTech’s Scientific Task Taxonomy to integrate our analysis of LLM interaction log and specialized model inventory data. Deriving a task taxonomy from job postings has limitations. Defining the line between science research and the broader STEM economy is challenging, and there will be some false negatives and positives in the science research job postings used (and hence extracted tasks). The taxonomy will continue to be improved as we fine-tune the ability to characterize the tasks performed across scientific subfields.

The mapping procedure we use to classify logs and tasks extracted from specialized models into the MIT FutureTech Scientific Task Taxonomy also has limitations. Our classifier currently has the inherent issues associated with trying to disambiguate what tasks are science (and all downstream field/task classifications) using LLMs, without knowing user identity. For example, we are explicitly trying to assign a “science” label (through the LLM making this classification) only when the content of the conversation seems to suggest it. As a result we may systematically omit implicit or loosely framed scientific work differentially across task types (e.g. context may be less prevalent when scheduling meetings, and more prevalent when doing data analysis for a paper).

Additionally, and leaving aside classification error emerging from the need to assign noisy descriptions of scientific activity “in the wild” to ambiguous and in some case overlapping labels in a large taxonomy, our classification procedure maps each log/model task to a single MIT task, ignoring the fact that, for example, some specialized tasks might be relevant for multiple downstream scientific activities, or multiple tasks might be performed in a single Gemini conversation. It might also be the case that scientists “find their own use” for open source models, leveraging them for tasks that the developers had not anticipated or mentioned in the model descriptions. We plan to address these limitations and improve our classification and filtering procedure in future work.

We also note our paper captures a snapshot of AI adoption through the early and mid-period of 2026. This may matter for some of our findings. For example, we found a clear division of labor between LLMs and specialized models in this period. However, this may be subject to change. The emergence of new frontier LLMs (and specialized models) increasingly able to solve difficult problems in fields like mathematics, genomics or life sciences (e.g., Anthropic, 2026; Callaway, 2026; OpenAI, 2026; Romera-Paredes et al., 2023) creates uncertainty about the future evolution in the division of labor between different types of model. We plan to track this in future iterations of our work.

Finally, our results establish associations rather than causality. Future work is needed to isolate the causal impact of using AI tools on scientific productivity.

06

Discussion: the verification tax and move 37

AI tools are dividing the labor by task; about half are pushed toward safer problems; the verification tax is high; next may be an LLM orchestrating specialized models, with people choosing questions and setting the standard of evidence.

6.2 Discussion

Notwithstanding these limitations, we believe that our findings provide an important early look into how AI is reshaping science. Rather than a simple story of labor substitution or runaway productivity gains from automated discovery, our evidence points to a more nuanced picture.

AI usage is widespread across scientific disciplines, so it is perhaps unsurprising that scientists appear more over-represented in terms of AI usage relative to other occupations. Of course, the diffusion of AI in science is closely related to other distributions (and capacity to make investments, e.g. Brynjolfsson et al. (2021)) in computational or intangible capital. Gemini usage scales roughly in proportion to national researcher populations and the creation of specialized AI models is mostly concentrated in a small group of developed economies. Cutting-edge scientific AI often requires complements like compute availability and domain specific human capital. This suggests a risk that without policy action under-resourced or developing regions could fall behind. The frontier of AI-augmented science might require investment to mitigate global disparities in research capacity.

In terms of task analyses, specialized AI models and Gemini usage appear similar. Both have the primary use category in quantitative modeling and data analysis. But when we dig into specific types of analysis, we start seeing evidence of model specialization and a division of labor within the scientific production function. Specialized models function as a kind of domain-specific capital—generating synthetic data, predicting molecular properties, and executing simulations—tasks that seem better served by a specialized model instead of a general LLM (at least for now). Meanwhile, general-purpose models absorb a wider range of tasks like writing code, synthesizing literature and operations. Similar to the way in which researchers specialize across disciplines or workflow components, AI tools may be specializing as well.

Where adopted, AI seems to deliver meaningful productivity dividends. Researchers report substantial net time savings that they recycle back into other projects. However, these changes enabled by AI in science will not only be on the extensive margin. About half of surveyed researchers report that AI directs them toward safer, incremental questions where clean data and benchmarks exist. Only 28 percent instead pursue riskier projects–despite evidence that most important scientific progress relies on high-risk exploration (Azoulay et al., 2011) and novel hypotheses (Uzzi et al., 2013). It remains to be seen whether AI models will lead researchers toward more incremental work, enable “moonshot” projects, or possibly both; future work should examine the relationship between AI usage and the risk profile of scientists. It might be that AI delivers a “Streetlight Effect” (Hoelzemann et al., 2024; Nagaraj and Tranchero, 2023) where the costs of executing incremental work drops relative to higher-risk projects. Incremental work made easy by AI could grow as a proportion of research output overall, at least in the short-run. On the other hand, scientists report that AI enables more interdisciplinary work, potentially acting as a translator helping researchers recombine ideas more effectively (Fang and Evans, 2026) and helping overcome the burden of knowledge (Jones, 2009).

[Footnote 28] This is not necessarily a negative outcome—for example, incremental “normal science” that enhances the understanding of disease mechanisms contributes to evidence on drug development (McNamee et al., 2017).

Despite widespread task-level time savings, macro-level scientific discovery rates remain bound by workflow task dependencies. These hidden scientific “organizational” complementarities are similar to AI’s impact on work in a general sense (Demirer et al., 2026; Gans and Goldfarb, 2026; Kremer, 1993). For example, it may be that computational modeling and data analysis are accelerated by AI tooling or that specialized AI models provide more promising biological structures to test, but the residual tasks in the bundle of scientific work are harder to scale and absorb the time savings. This bundling and task interlock for scientific work means that broader productivity gains may require redrawing job boundaries, and for some types of work the task bundling is stronger than others (Garicano et al., 2026). Enhanced productivity in data analysis might, for example, push the queue of work into hard-to-automate, physical stages that become a bottleneck (Jones, 2025).

Furthermore, AI outputs are not necessarily “turn-key”. Almost 90 percent of researchers who report time savings spend a meaningful share of their saved time verifying AI outputs, and about 46 percent spend more than a quarter of their saved time doing so. As with software engineering, scientists who adopt AI tools may be pushed toward auditing, debugging, and quality screening. Especially in the case of scientific work, being correct is of paramount importance. Lowering the marginal cost of hypothesis generation, code generation, analysis, or testing will create a high complementary marginal value of verifying results (or, perhaps, human judgment e.g., Agrawal et al. (2019)). This high “verification tax” likely arises from the high value of reliable and correct output in science.

The division of labor we document also suggests what the scientist of the near future might look like. Today a researcher moves between tools by hand: a chatbot for code and drafting, AlphaFold or a materials model for prediction, a spreadsheet or a lab notebook in between. The next step may be an LLM that orchestrates the specialized models directly, planning the experiment, calling the right model for each step, checking outputs against each other, and returning a candidate answer for the scientist to judge. The human’s job shifts toward choosing the question, setting the standard of evidence, deciding what is worth the cost of a physical test. Our data say we are not there yet: verification still consumes a large share of the time AI saves and the backlog of untested hypotheses is growing. The hope, in the end, is larger than time savings. When AlphaGo (Silver et al., 2016) played its 37th move against Lee Sedol in 2016, it played a move that its own model estimated a human would make about one time in ten thousand, and it won. The promise of more advanced AI systems in science is that they will eventually make moves of that kind: discoveries beyond the limit of what human researchers would have tried on their own.

Our work here has provided some initial measurements of the adoption of AI tools in science, its potential impact and bottlenecks that could limit this impact. As these limitations arise, so do new opportunities for cross-disciplinary work and new directions of inquiry in computationally demanding domains. Scientific production shares many of the features of other kinds of production. The rapid integration and adoption of AI tools by scientists combined with the substantive importance of breakthrough discoveries makes this an ideal domain to understand opportunities posed by AI. By continuing this work over time, we can watch the experimenters experiment, learning from scientific successes and failures to unlock AI-driven productivity gains in scientific R&D and elsewhere in the economy.

Where Indigo landsFurther

Indigo's conclusion

The hardest first-hand evidence yet that verification is incompressible: AI crushes generation, the bottleneck moves to verification and physical experiment, and 89% of time-savers spend over a tenth of the saved time checking AI output. But it's Gemini grading itself plus a self-reported survey; read the numbers as an early Google snapshot.

What to remember

  1. The bottleneck moves to verification: AI crushes upstream generation, constraints move to verification and physical experiments; 41% report a hypothesis backlog; 89% of time-savers spend over a tenth of it checking AI output.
  2. LLMs and specialized models are economic complements with almost no overlap at the finest level, pointing to LLMs orchestrating specialized models: value lies in orchestration, specialized models and data.
  3. 49% are pushed toward safer, incremental problems versus 28% toward riskier ones: verifiable domains speed up, possibly at the cost of risky questions.
  4. Nearly 7 hours saved a week, mostly reinvested in research, 8 in 10 report higher output; but it's self-reported perception, not measurement.
  5. Three discounts: Gemini only, a self-reported non-probability survey, and a classifier only 28% accurate at the finest level. Trust the direction; treat the numbers as a snapshot.

Back on the long-running theses

confirms

Verification is incompressible The hardest first-hand pillar yet: Google measured, with 15 million interactions and 600+ scientists, that AI compresses the upstream, the bottleneck moves down, and verification eats the dividend.

conflicts

Can we scale beyond easy-to-verify fields? 49% moving to safer problems is an early warning: conquering verifiable domains isn't advancing the frontier.

confirms

You don't own the model LLMs orchestrating specialized models, plus data and benchmarks, place the value in orchestration, unique specialized models and data, not the general LLM.

adds to

The politics of open weights Lower-income countries citing specialized models without developing them is first-hand evidence of open models diffusing widely.

confirms

DeepMind Co-Scientist The same honesty: AI helping research is real, but final verification and the physical world are the bottlenecks.

adds to

Andrew Ng: no job apocalypse, because AI isn't good enough yet Ng describes complementarity at the macro level; this report gives the micro data for science: complementary model types, with saved time reinvested in research.

What would change my mind

open-ended science with no cheap verifier showing the same inflection in productivity and discovery without physical verification catching up.

Finished. Indigo's take on this piece is in two places: