AI 工具在按任务分工;约一半人被推向更安全的问题;验证税很高;下一步可能是 LLM 调度专用模型,人负责选问题、定证据标准。 读这一段原文 →
什么会让我改口
没有廉价验证器的开放式科研,在物理验证没跟上的情况下,生产率和发现的斜率照样拐上去。
怎么读这篇
Google 和 DeepMind 自己发的机构报告,替自家说话的成分明确:秀 AI 做科学,秀 Gemini 的采用。但数据是真的、量大,1500 万次 Gemini 交互加 637 位科学家的问卷,而且罕见地把负面如实写了出来:下游瓶颈、假设积压、转向更安全的问题。数字当 Google 口径的早期快照,方向可信。
科学进步是经济增长和繁荣的关键驱动力。人们对 AI 给科学带来的影响既非常兴奋,也有担忧,但到目前为止数据很少。我们从三个数据源给出早期洞见:1500 万次 Gemini 交互的样本、覆盖各学科的 2600 多个专用 AI 模型清单,以及对 600 多位科学家的问卷调查。我们把这些数据映射到一套新的科学任务分类体系上,研究科学家如何使用 AI。主要发现有四条。第一,采用广泛、覆盖面广:科学家使用 AI 的程度高于大多数其他职业;专用 AI 模型覆盖的学科很广,被引用次数很高;接近一半受访科学家报告每天都在使用某种形式的 AI。第二,我们发现证据表明,LLM(以 Gemini 的使用为代理)和专用模型是互补的:LLM 用于通用分析、写代码和撰写论文,专用模型则提供特定领域的预测、数据生成和分类。第三,科学家报告使用 AI 带来了很大的生产率提升:每周节省将近 7 小时,这些时间主要被重新投入到更多研究中。最后,我们表明 AI 已经在改变科研过程。随着科研的某些阶段变得更容易,瓶颈向下游转移。科学家报告未经检验的假设越积越多,对产出进行核查的需求也很大。我们的发现表明,AI 在提高科研生产率方面潜力巨大。然而,与其他行业一样,它的最终影响将取决于复杂的任务相互依赖,以及为消除新出现的瓶颈所做的投资。
以往关于 AI 采用及其对科学影响的学术研究,主要依赖公开的科学计量记录,比如论文发表和引用趋势、专利申请,以及 LLM 辅助写作留下的痕迹(Kusumegi 等, 2025; Renault 等, 2026; Trišović 等, 2025; Duede 等, 2024)。这些都是最终产出,有滞后,对采用情况的追踪也不完美。前沿 AI 实验室掌握着研究者实际如何使用模型的实时数据,但基于这类日志的研究,到目前为止关注的是整体的劳动力市场格局(Chatterji 等, 2025; Appel 等, 2025; Handa 等, 2025; Iscenko 等, 2026),对科学关注相对很少。
本文开启一个关于 AI 对科学影响的新研究项目,结合了三个新数据源的洞见:1500 万次 Gemini 交互的样本;覆盖各学科(从蛋白质结构预测到材料发现和天气预报)的 2600 多个专用 AI 模型的文献计量数据;以及对 600 多位科学家的问卷调查。我们把三个数据源都映射到 MIT FutureTech 开发的一套新的科学任务分类体系上,它能在 OpenAlex 的各个科学子领域中细到任务层面(Emmens 等, 2026)。
分析揭示了四个发现。第一,科学家在 AI 使用上领先于其他职业;相对于其就业占比,科学类职业在 Gemini 使用中占比偏高。接近一半受访科学家报告每天都把 AI 作为工作流程的一部分。但只看 LLM 的使用会漏掉很大一块:受访科学家使用专用模型的程度,和使用编程助手差不多。这些模型在学术工作中也被大量引用。科学中的 AI 采用在地理上也很集中,使用量随一个国家的科研人员规模而增长。第二,LLM 和专用模型是经济上的互补品,两者分工明确:前者主要广泛用于写代码和写作等各类任务,后者用于更专门的生成、预测和模拟,两类模型的任务重叠有限。第三,科学家报告 AI 带来了可观的生产率提升,平均每周节省将近 7 小时,这些时间大多被投回研究。研究者还报告获得了更多跨领域、跨学科的洞见。第四,虽然有些任务(比如定量和计算工作)可能被加速了,但科学产出的瓶颈转移到了下游的物理实验和验证。科学家报告未经检验的假设在积压,大量时间花在核查 AI 的输出上,并且出现了转向更安全问题的倾向。
再详细介绍一下方法。我们的第一个数据源来自 Google ATLAS 1.0 项目(Iscenko 等, 2026),包括会话界面(Gemini App 和 AI Mode)以及 Gemini API 上约 1500 万次匿名交互。科研工作通过一个三阶段流程分离出来:首先,一个分类器去掉与工作无关和教学类的交互;然后,一个过滤器把样本限定在科研集中的细分职业;最后,由于这些职业里常有非科学家,我们基于 OECD 的弗拉斯卡蒂手册(2015)和英国研究与创新署(2025)构建了一个新的科学分类器,根据交互内容判断它们是否可能属于科学家的工作流程。这样留下了 36 万次科学交互。
但如何把 LLM 交互映射到具体任务上?职业任务(O*NET)和时间使用(ATUS)都有标准分类体系,但用在科学研究上是不够的。「分析数据」「写报告」这样笼统的类别,抹平了定义科研的那些领域特定活动,从提出假设到设计分子构建体、优化检测方法。衡量 AI 对科学的影响,需要一个共同的尺度,分别刻画研究者在各学科中做什么,以及其中哪些任务得到了 AI 工具的帮助、赋能或自动化。为此,我们用自己定制的分层发现与分类引擎「观察聚类与分类体系组织」(OCTO,Iscenko 等 (2026)),把科学类 LLM 交互映射到两套基于科学的分类体系上:OpenAlex 科学学科,以及 MIT FutureTech 新任务分类体系的三个层级。
科学家还报告看到了 AI 带来的真实生产率提升。大约四分之三的研究者报告节省了时间,平均每周节省将近 7 小时。这些时间回到了研究中,变成了更多产出。大约十分之八的人报告过去三年实验室产出更高,89% 预计还会继续增加。收益也不只在速度上:大约 68% 的科学家报告更容易获得其他学科的洞见,这与人们的一个期望相符,即 AI 也许能减轻把研究者推向越来越窄专业的「知识负担」(Jones, 2009)。
04
瓶颈下移到验证
超过十分之四的人说主要约束移到了下游;41% 的人假设积压;省下的时间有相当一部分拿去核查 AI 输出;49% 被推向更安全的问题。
但瓶颈依然存在,这可以解释为什么局部的生产率提升还没有转化为新发现和新应用的爆发。超过十分之四的受访科学家报告,过去两年里他们的主要约束已经转移到下游:实验室执行、临床验证和实地数据采集。AI 工具可能增加了可行理论和假设的数量,却不一定增加了检验它们的手段。这导致 41% 的人报告未经检验的假设积压越来越多。验证吃掉了红利的一大块:节省了时间的人里,89% 把其中超过十分之一的时间花在核查 AI 的输出上,46% 花了超过四分之一。从长远看最令人担忧的是,49% 的科学家说 AI 把他们推向更安全、更渐进的项目,那里基准已经建立、结果可靠;而说 AI 让他们敢于挑战更有风险问题的只有 28%。如果 AI 主要降低的是渐进式工作的成本,它可能会增加论文的数量,却不能实质性地推进科学前沿。
本文是对 AI 此刻如何影响科学的一次早期快照。LLM 与专用模型之间的分工,指出了未来可能的方向:一个 LLM 调度者规划实验,调用 AlphaFold 或材料模型之类的专用模型,并解读结果。这样的系统还在开发中。要把它做出来并兑现其潜力,需要解决科学家指出的那些瓶颈,把投资引向物理实验和验证能力、验证工具,以及能让 AI 去碰更有风险而不是更安全问题的数据和基准。
本文其余部分安排如下。第 2 节详细介绍数据源和测量方法。第 3 节用 Google ATLAS 数据(Iscenko 等, 2026),经 MIT FutureTech 科学任务分类体系(Emmens 等, 2026)映射到科研工作流程和领域上,展示 LLM 在科学中采用的一些早期信号。第 4 节分析专用科学 AI 模型所在的领域、与下游任务的关联,以及溢出效应的早期信号。第 5 节给出关于研究者时间分配、感知到的瓶颈和预期生产率影响的问卷结果。最后,第 6 节说明局限并讨论结果。
第二,我们的专用模型清单是收录科学领域大量特定领域 AI 模型的第一步,但目前不应被视为有代表性。例如,我们目前聚焦于那些更容易通过网络搜索、文献计量平台和代码仓库发现的知名模型,而遗漏了在单项研究中基于开源框架构建、或微调开放权重模型得到的定制模型,以及商业模型。专用模型清单及基于它的引用分析也天然滞后,这意味着我们捕捉的是一个动态格局中的快照。随着 AI 辅助论文的同行评审标准发生变化,长期的引用模式可能会有不同的演变。
我们的第三个数据源,也就是科学家问卷,可能存在选择偏差。尤其是,那些已经对 AI 充满热情或经常使用 AI 的科学家,可能更愿意回答关于科学中 AI 的问卷。这可能会高估真实的采用率和平均每周节省的时间。我们的样本量也限制了我们在一些关键问题上探究学科差异的能力,比如下游验证(例如物理实验和临床试验)中瓶颈的转移,这可能因学科而异。
我们依靠 MIT FutureTech 的科学任务分类体系,把对 LLM 交互日志和专用模型清单数据的分析整合起来。从招聘启事中推导任务分类体系有其局限。在科学研究与更广泛的理工科经济之间划清界线并不容易,所用的科研岗位招聘启事(以及由此提取的任务)中难免有漏判和误判。随着我们对刻画各科学子领域所执行任务的能力做进一步精调,这套分类体系会继续改进。
我们用来把日志以及从专用模型中提取的任务归入 MIT FutureTech 科学任务分类体系的映射过程,也有其局限。在不知道用户身份的情况下,用 LLM 来区分哪些任务属于科学(以及所有下游的领域和任务分类),我们的分类器目前存在这类做法固有的问题。例如,我们明确只在对话内容似乎表明是科学工作时,才(由 LLM 来分类)打上「科学」标签。因此,我们可能会在不同任务类型上有差别地、系统性地漏掉那些隐含的或表述松散的科学工作(例如安排会议时语境可能较少,而为论文做数据分析时语境较多)。
此外,撇开把「野外」那些嘈杂的科研活动描述归入一个大型分类体系中模糊、有时还相互重叠的标签所带来的分类误差不谈,我们的分类过程把每条日志或模型任务只映射到一个 MIT 任务上,忽略了这样的事实:例如,某些专门任务可能与多项下游科研活动相关,或者一次 Gemini 对话中可能包含多项任务。科学家也可能为开源模型「找到自己的用法」,把它们用于开发者没有预料到、也没有在模型说明中提到的任务。我们计划在今后的工作中处理这些局限,改进分类和过滤流程。
最后,我们的结果确立的是相关关系,而不是因果关系。需要今后的工作来分离出使用 AI 工具对科研生产率的因果影响。
06
讨论:验证税与第 37 手
AI 工具在按任务分工;约一半人被推向更安全的问题;验证税很高;下一步可能是 LLM 调度专用模型,人负责选问题、定证据标准。
6.2 讨论
尽管有这些局限,我们认为我们的发现为了解 AI 如何重塑科学提供了一次重要的早期观察。我们的证据指向的,不是劳动力被简单替代、或自动化发现带来生产率失控飙升的故事,而是一幅更细致的图景。
AI 的使用遍及各个科学学科,所以科学家在 AI 使用上相对于其他职业占比偏高,也许并不意外。当然,科学中 AI 的扩散,与计算资本或无形资本的其他分布(以及投资能力,例如 Brynjolfsson 等 (2021))密切相关。Gemini 的使用量大致与各国科研人员数量成比例,而专用 AI 模型的开发主要集中在少数发达经济体。前沿的科学 AI 往往需要算力可得性和特定领域人力资本这类互补条件。这意味着一个风险:如果没有政策行动,资源不足或发展中的地区可能会落后。AI 增强科学的前沿,可能需要投资来缓解全球科研能力的差距。
在任务分析上,专用 AI 模型和 Gemini 的使用看起来相似,两者的主要用途都是定量建模和数据分析。但当我们深入到具体的分析类型,就开始看到模型专门化的证据,以及科学生产函数内部的分工。专用模型的作用类似一种特定领域的资本:生成合成数据、预测分子性质、执行模拟,这些任务看起来(至少目前)更适合由专用模型而不是通用 LLM 来完成。与此同时,通用模型承担了更广泛的任务,比如写代码、综合文献和日常事务。就像研究者会按学科或工作流程环节分工一样,AI 工具可能也在专门化。
在被采用的地方,AI 似乎带来了可观的生产率红利。研究者报告了可观的净节省时间,并把它们投到其他项目上。然而,AI 在科学中带来的这些变化不会只发生在「做多少」的层面。大约一半受访研究者报告,AI 把他们引向更安全、更渐进的问题,那里有干净的数据和基准。只有 28% 转而去做风险更高的项目,尽管有证据表明,最重要的科学进步依赖高风险的探索(Azoulay 等, 2011)和新颖的假设(Uzzi 等, 2013)。AI 模型究竟会把研究者引向更多渐进式工作、还是促成「登月」项目,抑或两者兼有,还有待观察;今后的工作应当研究 AI 使用与科学家风险偏好之间的关系。也许 AI 会带来一种「路灯效应」(Hoelzemann 等, 2024; Nagaraj 和 Tranchero, 2023):做渐进式工作的成本相对于高风险项目下降了。至少在短期内,被 AI 变得容易的渐进式工作,在全部研究产出中的比重可能会上升。另一方面,科学家报告 AI 促成了更多跨学科工作,可能像一个翻译者,帮助研究者更有效地重组想法(Fang 和 Evans, 2026),并帮助克服知识负担(Jones, 2009)。
尽管任务层面的时间节省很普遍,宏观层面的科学发现速度仍受工作流程中任务依赖关系的约束。这些隐藏的科研「组织」互补性,与 AI 对一般工作的影响类似(Demirer 等, 2026; Gans 和 Goldfarb, 2026; Kremer, 1993)。例如,计算建模和数据分析也许被 AI 工具加速了,专用 AI 模型也许提供了更有希望去测试的生物结构,但科研工作这一整包任务中剩下的部分更难规模化,会吸收掉节省下来的时间。科研工作的这种捆绑和任务互锁意味着,更广泛的生产率提升可能需要重新划定岗位边界,而某些类型的工作捆绑得比其他工作更紧(Garicano 等, 2026)。例如,数据分析生产率的提升,可能把排队的工作推向难以自动化的物理阶段,使之成为瓶颈(Jones, 2025)。
此外,AI 的输出并不一定是「拿来即用」的。在报告节省了时间的研究者中,接近 90% 把节省时间中相当一部分用于核查 AI 的输出,约 46% 花在这上面的时间超过节省时间的四分之一。就像在软件工程中一样,采用 AI 工具的科学家可能会被推向审计、调试和质量筛查。尤其在科研工作中,正确是最重要的。降低提出假设、生成代码、分析或测试的边际成本,会让核实结果(或者也许是人的判断,例如 Agrawal 等 (2019))的互补性边际价值变得很高。这种高昂的「验证税」,很可能源于科学中可靠、正确的产出价值很高。
我们记录的这种分工,也暗示了不久之后的科学家可能是什么样子。今天,研究者在工具之间手动切换:用聊天机器人写代码和起草,用 AlphaFold 或材料模型做预测,中间夹着电子表格或实验记录本。下一步可能是一个 LLM 直接调度这些专用模型:规划实验,在每一步调用合适的模型,交叉核对各个输出,再把一个候选答案交给科学家判断。人的工作转向选择问题、设定证据标准、决定什么值得付出物理实验的成本。我们的数据表明,我们还没走到那一步:验证仍在消耗 AI 节省下来的大部分时间,未经检验的假设积压在增加。归根到底,人们的期望不止于节省时间。2016 年 AlphaGo(Silver 等, 2016)对李世石下出第 37 手时,那一步棋按它自己的模型估计,人类大约一万次里才会下一次,而它赢了。更先进的 AI 系统在科学中的前景,就在于它们终将下出那样的棋:做出超越人类研究者自己会去尝试的极限的发现。
我们的工作对科学中 AI 工具的采用、其潜在影响以及可能限制这种影响的瓶颈,给出了一些初步的测量。随着这些局限浮现,跨学科工作的新机会、以及计算密集型领域的新研究方向也随之出现。科学生产与其他各类生产有许多共同特征。科学家对 AI 工具的快速整合和采用,加上突破性发现的巨大重要性,使科学成为理解 AI 所带来机遇的理想领域。通过长期持续这项工作,我们可以看着实验者做实验,从科学的成功和失败中学习,释放 AI 在科研以及经济其他领域带来的生产率提升。
判断收口延伸
Indigo 的结论
「验证不可压缩」迄今最硬的一手实证:AI 压掉生成,瓶颈下移到验证和物理实验,89% 的省时者拿超过一成的时间去核查 AI 的输出。但这是 Gemini 自证加问卷自报,数字当 Google 口径的早期快照。
需要记住的几件事
瓶颈下移到验证:AI 压掉上游的生成,约束移到验证和物理实验;41% 假设积压;89% 的省时者花一成以上时间核查 AI 输出。
Google ATLAS, Google DeepMind, MIT FutureTech · ai.google · 2026-09-15
The hardest first-hand evidence yet that verification is incompressible: once AI crushes generation, the whole bottleneck moves to verification.
Indigo's conclusion
The hardest first-hand evidence yet that verification is incompressible: AI crushes generation, the bottleneck moves to verification and physical experiment, and 89% of time-savers spend over a tenth of the saved time checking AI output. But it's Gemini grading itself plus a self-reported survey; read the numbers as an early Google snapshot.
How to read this An institutional report from Google and DeepMind with an obvious interest: show off AI for science and Gemini adoption. But the data is real and large, 15 million Gemini interactions plus a survey of 637 scientists, and it's unusually honest about the bad news: downstream bottlenecks, a backlog of untested hypotheses, a drift toward safer problems. Treat the numbers as an early Google-side snapshot; trust the direction.
What to remember
The bottleneck moves to verification: AI crushes upstream generation, constraints move to verification and physical experiments; 41% report a hypothesis backlog; 89% of time-savers spend over a tenth of it checking AI output.
LLMs and specialized models are economic complements with almost no overlap at the finest level, pointing to LLMs orchestrating specialized models: value lies in orchestration, specialized models and data.
49% are pushed toward safer, incremental problems versus 28% toward riskier ones: verifiable domains speed up, possibly at the cost of risky questions.
Nearly 7 hours saved a week, mostly reinvested in research, 8 in 10 report higher output; but it's self-reported perception, not measurement.
Three discounts: Gemini only, a self-reported non-probability survey, and a classifier only 28% accurate at the finest level. Trust the direction; treat the numbers as a snapshot.
Breakdown · 6 steps
01
Abstract: four findings
Three data sources: 15 million Gemini interactions, 2,600+ specialized models, 600+ scientists. Findings: broad adoption, complementary model types, nearly 7 hours saved a week, bottlenecks shifting to verification. Read this part →
02
Why and how
Both optimism and worry lack evidence; the report maps logs, a model inventory and a survey onto one taxonomy of scientific tasks, following AI from capability to adoption to its effect on how science gets done. Read this part →
03
Broad adoption, complementary models
Science occupations are 2.7 times as likely to use AI as the employment baseline; LLMs and specialized models overlap less the finer the tasks; about three quarters save time, nearly 7 hours a week on average. Read this part →
04
The bottleneck moves to verification
More than 4 in 10 say their main constraint has moved downstream; 41% report a backlog of hypotheses; much of the saved time goes to checking AI output; 49% are pushed toward safer problems. Read this part →
05
Omitted data, and the authors' own limits
Gemini only, no enterprise data; the model inventory isn't representative; the survey may be self-selected; classification has errors; it's a 2026 snapshot showing association, not causation. Read this part →
06
Discussion: the verification tax and move 37
AI tools are dividing the labor by task; about half are pushed toward safer problems; the verification tax is high; next may be an LLM orchestrating specialized models, with people choosing questions and setting the standard of evidence. Read this part →
What would change my mind
open-ended science with no cheap verifier showing the same inflection in productivity and discovery without physical verification catching up.
How to read this
An institutional report from Google and DeepMind with an obvious interest: show off AI for science and Gemini adoption. But the data is real and large, 15 million Gemini interactions plus a survey of 637 scientists, and it's unusually honest about the bad news: downstream bottlenecks, a backlog of untested hypotheses, a drift toward safer problems. Treat the numbers as an early Google-side snapshot; trust the direction.
Three data sources: 15 million Gemini interactions, 2,600+ specialized models, 600+ scientists. Findings: broad adoption, complementary model types, nearly 7 hours saved a week, bottlenecks shifting to verification.
Abstract
Scientific progress is a key driver of economic growth and prosperity. There is great excitement - but also concerns - about the impacts of AI on science, but so far little data. We provide early insights on this from three data sources: a sample of 15 million Gemini interactions, an inventory of over 2,600 specialized AI models across disciplines, and a survey of over 600 scientists. We map these data to a new taxonomy of scientific tasks to study how scientists are using AI. Four main findings emerge. First, we find broad adoption and coverage: scientists use AI more than most other occupations. Specialized AI models have broad disciplinary coverage and are highly cited. Nearly half of the scientists surveyed report using some form of AI every day. Second, we document evidence that LLMs (proxied through Gemini usage) and specialized models act as complements—LLMs are used for general analysis, coding, and manuscript preparation, while specialized models provide domain-specific predictions, data generation and classification. Third, scientists report large productivity gains from using AI: a saving of nearly 7 hours per week, time which is primarily re-invested in more research. Finally, we show that AI is already changing the scientific process. As some stages of scientific research become easier, bottlenecks shift downstream. Scientists report an increased backlog of untested hypotheses and substantial demand for output verification. Our findings suggest that AI holds significant potential to increase scientific productivity. However, as with other sectors, its ultimate impact will be governed by complex task interdependencies and investment into the elimination of emerging bottlenecks.
02
Why and how
Both optimism and worry lack evidence; the report maps logs, a model inventory and a survey onto one taxonomy of scientific tasks, following AI from capability to adoption to its effect on how science gets done.
1 Introduction
Scientific discovery is one of the primary drivers of long-run progress and economic growth (Romer, 1990; Mokyr, 1992; Aghion and Howitt, 1992; Jones, 1995; Akcigit et al., 2021). Yet concerns of decreased research productivity or “ideas getting harder to find” have been raised in recent years (Bloom et al., 2020; Park et al., 2023). The rapid development of AI has fueled optimism that the technology could reverse this trend and speed up scientific progress (Wang et al., 2023; Amodei, 2024; Curto Millet, 2025; Hassabis and Manyika, 2026; Agrawal et al., 2026). Specialized AI models, such as DeepMind’s AlphaFold 2 (Jumper et al., 2021; Hassabis, 2022) have already been recognized as Nobel-caliber breakthroughs. Meanwhile, general-purpose LLMs are rapidly diffusing across day-to-day research workflows (Van Noorden and Perkel, 2023; Wiley, 2025).
If AI acts as an “IMI”, an Invention of a Method of Invention (Aghion et al., 2017; Cockburn et al., 2018; Crafts, 2021; Cunningham et al., 2026), it would create a durable acceleration to productivity growth over time. Much is riding on AI’s potential for economic growth and progress, as increasing government debt and aging populations threaten to disrupt the balanced growth path of roughly 2 percent annual expansion per capita over the last 150 years in the US (Jones, 2026); AI’s role in science may be key for overcoming these headwinds and accelerating growth.
Yet the optimism around AI’s potential to increase scientific progress is not universal. Some worry that a deluge of AI-generated content, hallucinations, and “illusions of understanding” will strain peer review and validation (Birhane et al., 2023; Messeri and Crockett, 2024). Others argue that models optimized on the existing literature will steer researchers toward incremental, well-known paradigms rather than high-risk, high-reward discoveries (Duede, 2025), or that steep compute and capital requirements will widen the gap between well-funded laboratories and under-resourced institutions (Thompson et al., 2022; Ahmed et al., 2023; Besiroglu et al., 2024). Some also worry that AI might accelerate the production of low-quality research that merely appears scientific (Luo et al., 2025; Gyevnár et al., 2026). But these debates have thus far evolved with little empirical evidence.
Prior academic research about AI adoption and its impact on science has primarily relied on open scientometric records such as publication and citation trends, patent filings, and traces of LLM-assisted writing (Kusumegi et al., 2025; Renault et al., 2026; Trišović et al., 2025; Duede et al., 2024). These are finalized outputs that arrive with a lag and track adoption imperfectly. Frontier AI labs hold real-time telemetry on how researchers actually use models, but the studies built on such logs have so far focused on aggregate labor market patterns (Chatterji et al., 2025; Appel et al., 2025; Handa et al., 2025; Iscenko et al., 2026), with relatively little focus on science.
This paper inaugurates a new research program on the impact of AI on science by combining insights from three new data sources: a sample of 15 million Gemini interactions, bibliometrics data for more than 2,600 specialized AI models across academic disciplines (from protein structure prediction to materials discovery and weather forecasting), and a survey of more than 600 scientists. We map all three data sources to a new taxonomy of scientific tasks developed by MIT FutureTech, which gives task-level granularity across OpenAlex scientific subfields (Emmens et al., 2026).
The analysis reveals four findings. First, scientists lead other occupations in the use of AI; scientific occupations are overrepresented in Gemini usage relative to their share of employment. Nearly half of surveyed scientists report using AI every day as part of their workflow. But looking at LLM use alone would miss much of the picture. Surveyed scientists use specialized models about as much as coding assistants. These models are also heavily cited in scholarly work. AI adoption in science is also geographically concentrated, with usage scaling with a country’s scientific workforce. Second, LLMs and specialized models are economic complements with a clear division of labor between them: the former are primarily used broadly across tasks including coding and writing, while the latter are used for more specialized generation, prediction, and simulation, with limited overlap in tasks between the two model categories. Third, scientists report substantial productivity gains from AI, with average time saving of just below 7 hours per week. The time is mostly put back into research. Researchers also report greater cross-field and interdisciplinary insights. Fourth, while some tasks—for example, quantitative, computational work—may be accelerated, the bottlenecks to scientific output shift downstream into physical experimentation and validation. Scientists report a backlog of untested hypotheses, substantial time spent verifying AI outputs, and a tilt toward safer questions.
To give more detail on the methodology, our first data source comes from the Google ATLAS 1.0 project (Iscenko et al., 2026) which consists of about 15 million anonymized interactions across the conversational surfaces (Gemini App and AI Mode), and Gemini API. Scientific work is isolated through a three-stage pipeline. First, a classifier removes non-work and educational interactions; then, a filter restricts the sample to the detailed occupations where research is concentrated; finally, as these occupations often contain non-scientists, a new science classifier is built on the OECD Frascati (2015) and UK Research and Innovation (2025) using the content of interactions to classify if they are likely part of a scientist workflow. This leaves us with 360,000 science interactions.
But how does one map LLM interactions to specific tasks? Standard taxonomies exist for occupational tasks (O*NET) and time use (ATUS). But when applied to scientific research, these tools are insufficient. Generic categories like “analyzing data” or “writing reports” flatten the domain-specific activities that define research, from hypothesis formulation to molecular construct design or assay optimization. Measuring AI’s impact on science requires a common metric that separately captures what researchers do across disciplines and which of those tasks AI tools help, enable, or automate. To accomplish this, we used our custom hierarchical discovery and classification engine - Observation Clustering and Taxonomy Organisation (OCTO, Iscenko et al. (2026)) - to map scientific LLM interactions to two science-based taxonomies: the OpenAlex scientific disciplines and the three levels of the new MIT FutureTech task taxonomy.
The second data source comes from a newly-compiled inventory of over 2,600 specialized models published since 2012. These include models from many disciplines including biology (e.g., AlphaFold (Hassabis, 2022) predicting protein folding), material science (e.g., GNoME (Merchant et al., 2023) predicting stable crystal structures), chemistry (e.g., MatterGen (Zeni et al., 2025), a generative diffusion model for inorganic compound design) and more. We assembled the inventory by combining a bottom-up deep agentic search over web sources, publications, and code repositories with Epoch’s AI model database (Epoch AI, 2026), then enriched the output using metadata from OpenAlex. OCTO was then used to extract the research tasks that each model enables (e.g., predicting protein structures), which were then mapped to the respective tasks in the FutureTech and OpenAlex taxonomies, as done with the log data. This exercise puts the LLM and specialized model usage on a common metric, which allows us to empirically study the division of labor between these categories of models. We have tracked 460,000 unique citations to these models to trace diffusion and knowledge flows across scientific fields.
Our final data source is an original survey of 637 active researchers in the US and UK carried out in July-August 2026. The sample included Principal Investigators (PIs) in academia, industry researchers, and early career scientists across four major scientific domains. The survey allows us to fill the gaps that neither log telemetry nor publication data can speak to, such as perceived time saved, bottlenecks, priorities, and changes in the scientific workflow. Together, the three data sources let us follow AI from capability, to adoption, to its effect on how science gets done.
03
Broad adoption, complementary models
Science occupations are 2.7 times as likely to use AI as the employment baseline; LLMs and specialized models overlap less the finer the tasks; about three quarters save time, nearly 7 hours a week on average.
Looking at the LLM interaction data, we find extensive penetration of AI in science. Use of LLMs in science is overrepresented relative to its share of US employment; for example, in the US, the Standard Occupational Classification (SOC) 19 job category roles (containing a set of Life, Physical and Social Sciences occupations) are about 2.7 times more likely to use AI compared to employment baseline. While more quantitative fields such as computer science and engineering lead in adoption, other specialties such as political science, medicine, and agricultural sciences are also significant adopters of AI. Specialized models are present across most disciplines, though the share varies substantially by field of study. They are also highly cited. Nearly half of the papers that introduce specialized models are in the top 1% of citations within their respective fields. Usage in science is not sparse, with almost half of surveyed scientists using AI as part of their daily workflow. Looking at geographic variation in adoption, a one percentage point increase in a country’s share of the world’s researchers is associated with a 0.9 percentage point increase in its share of science LLM interactions, and specialized model development is concentrated in the U.S., China, the UK, the EU, Korea, and Canada. Some lower-income countries cite these models heavily without developing them, which suggests that open models can diffuse widely even where the capacity to build them does not.
Importantly, LLMs and specialized models appear to act as economic complements, filling out and potentially enhancing the capabilities of the other. While at the top level of the task taxonomy the two types of models look alike—quantitative data analysis and modeling dominates both—the similarity disappears once tasks are disaggregated further. The correlation in usage across tasks falls monotonically between the most aggregated and most granular tasks, and at the finest level there is almost no overlap. LLMs are used broadly across tasks such as writing and troubleshooting research code, statistical analysis, literature review, and helping with drafting documents. Specialized models are relatively more common in health and life sciences and are used to predict disease outcomes, engineer molecular constructs, and run complex simulations. Therefore both model families seem to reinforce one another in the production of scientific knowledge, limiting the extent to which one technology can at this point substitute for the other. Survey data tells a similar story: scientists split usage in both LLMs and specialized models.
Scientists also report seeing real productivity gains from AI. Around three quarters of researchers report time savings, with the average amount clocking in at almost 7 hours saved a week. That time goes back into research—into more output. About 8 in 10 report higher lab output over the past three years and 89% expect further increases. The gains extend beyond speed. About 68% of scientists report more access to insights from other disciplines, consistent with hopes that AI may lower the burden of knowledge that pushes researchers into ever-narrower specialties (Jones, 2009).
04
The bottleneck moves to verification
More than 4 in 10 say their main constraint has moved downstream; 41% report a backlog of hypotheses; much of the saved time goes to checking AI output; 49% are pushed toward safer problems.
But bottlenecks remain, which can explain why localized productivity gains have not yet translated into an explosion of new discoveries and applications. More than 4 in 10 of surveyed scientists report that their primary constraint has moved downstream over the past two years, into lab execution, clinical validation, and field data collection. AI tools may have increased the number of viable theories and hypotheses but not necessarily the means to test them. That has led 41% to report a growing backlog of untested hypotheses. Verification absorbs a large share of the dividend: 89% of those who save time spend more than a tenth of it checking AI outputs, and 46% spend more than a quarter. Most concerning for the long run, 49% of scientists say AI pushes them toward safer, more incremental projects where benchmarks are established and results are reliable, against 28% who say it lets them take on riskier questions. If AI mainly lowers the cost of incremental work, it could raise the volume of papers while not meaningfully advancing the scientific frontier.
This paper represents an early snapshot into how AI is affecting science right now. The division of labor between LLMs and specialized models points to where this could go: an LLM orchestrator that plans experiments, calls specialized models such as AlphaFold or a materials model, and interprets the results. That system is still work in progress. Delivering it and realizing its promise requires solving the bottlenecks that scientists identify and directing investment to physical experimentation and validation capacity, to verification tools, and to the data and benchmarks that would let AI take on riskier questions rather than safer ones.
The remainder of this paper is structured as follows. Section 2 details our data sources and measurement. Section 3 shows some early signals of LLM adoption in science using Google ATLAS data (Iscenko et al., 2026) mapped to scientific workflows and fields via the MIT FutureTech Scientific Task Taxonomy (Emmens et al., 2026). Section 4 analyzes the fields, downstream task relevance, and early signals of spillovers of the specialized scientific AI models. Section 5 presents survey findings on researchers’ time allocation, perceived bottlenecks, and expected productivity impacts. Finally, Section 6 outlines our limitations and discusses the results.
05
Omitted data, and the authors' own limits
Gemini only, no enterprise data; the model inventory isn't representative; the survey may be self-selected; classification has errors; it's a 2026 snapshot showing association, not causation.
Sections 2–5: data and results
[Editor's note] Sections 2 to 5 describe the three data sources and the measurement in detail and present the results in figures and tables; the key numbers are summarized in the introduction above. They are omitted here; see the original PDF.
6.1 Limitations
Several limitations in our data, measurement, and scope must be acknowledged. First, our log analysis (section 3) considered a sample of 15 million interactions drawn exclusively from Google’s ATLAS (Iscenko et al., 2026) dataset. Enterprise data was excluded from this investigation. Generalization should be done with caution as Gemini users may skew toward specific scientific disciplines or have distinct use patterns from other proprietary and open models. For example, scientists in more regulated areas like health sciences may be relatively more likely to use enterprise accounts or specialized solutions, which we cannot observe. It is also worth noting that our analysis excludes agentic AI tools like Antigravity where scientists can use Gemini and other LLMs to leverage specialized model capabilities (e.g., Applebaum et al., 2026)—incorporating agentic data into our analysis is an important next step for our research program.
Second, our specialized model inventory is an initial step towards capturing the wealth of domain-specific AI models in science, but at this point should not be considered representative. For example, we currently focus on notable models that are easier to detect through web searches, bibliometric platforms and coding repositories, but miss bespoke models designed during individual studies building on open source frameworks or fine-tuning open weight models, as well as commercial models. The specialized model inventory and the citation analysis building on it are also inherently lagging which means we are capturing a snapshot in a dynamic landscape. It could be that long-term citation patterns evolve differently as peer-review standards for AI-assisted papers change.
Our third data source—the scientist survey—could suffer from selection bias. In particular those scientists who are already enthusiastic about or frequently use AI might be more likely to respond to a survey about AI in science. This could potentially overestimate the real adoption rate and the average weekly time savings. Our sample size also limits our ability to explore disciplinary differences in key questions such as shifting bottlenecks in downstream validation (e.g., physical experimentation and clinical trials), which could vary by discipline.
[Footnote 26] While chemistry and medicine could face large capital/temporal barriers in physical validation, disciplines like computer science or mathematics can validate hypotheses entirely in-silico (computationally). Our aggregate findings, as a result, may over-represent the physical bottlenecks of “wet” sciences compared to “dry” computational sciences.
We rely on MIT FutureTech’s Scientific Task Taxonomy to integrate our analysis of LLM interaction log and specialized model inventory data. Deriving a task taxonomy from job postings has limitations. Defining the line between science research and the broader STEM economy is challenging, and there will be some false negatives and positives in the science research job postings used (and hence extracted tasks). The taxonomy will continue to be improved as we fine-tune the ability to characterize the tasks performed across scientific subfields.
The mapping procedure we use to classify logs and tasks extracted from specialized models into the MIT FutureTech Scientific Task Taxonomy also has limitations. Our classifier currently has the inherent issues associated with trying to disambiguate what tasks are science (and all downstream field/task classifications) using LLMs, without knowing user identity. For example, we are explicitly trying to assign a “science” label (through the LLM making this classification) only when the content of the conversation seems to suggest it. As a result we may systematically omit implicit or loosely framed scientific work differentially across task types (e.g. context may be less prevalent when scheduling meetings, and more prevalent when doing data analysis for a paper).
Additionally, and leaving aside classification error emerging from the need to assign noisy descriptions of scientific activity “in the wild” to ambiguous and in some case overlapping labels in a large taxonomy, our classification procedure maps each log/model task to a single MIT task, ignoring the fact that, for example, some specialized tasks might be relevant for multiple downstream scientific activities, or multiple tasks might be performed in a single Gemini conversation. It might also be the case that scientists “find their own use” for open source models, leveraging them for tasks that the developers had not anticipated or mentioned in the model descriptions. We plan to address these limitations and improve our classification and filtering procedure in future work.
We also note our paper captures a snapshot of AI adoption through the early and mid-period of 2026. This may matter for some of our findings. For example, we found a clear division of labor between LLMs and specialized models in this period. However, this may be subject to change. The emergence of new frontier LLMs (and specialized models) increasingly able to solve difficult problems in fields like mathematics, genomics or life sciences (e.g., Anthropic, 2026; Callaway, 2026; OpenAI, 2026; Romera-Paredes et al., 2023) creates uncertainty about the future evolution in the division of labor between different types of model. We plan to track this in future iterations of our work.
Finally, our results establish associations rather than causality. Future work is needed to isolate the causal impact of using AI tools on scientific productivity.
06
Discussion: the verification tax and move 37
AI tools are dividing the labor by task; about half are pushed toward safer problems; the verification tax is high; next may be an LLM orchestrating specialized models, with people choosing questions and setting the standard of evidence.
6.2 Discussion
Notwithstanding these limitations, we believe that our findings provide an important early look into how AI is reshaping science. Rather than a simple story of labor substitution or runaway productivity gains from automated discovery, our evidence points to a more nuanced picture.
AI usage is widespread across scientific disciplines, so it is perhaps unsurprising that scientists appear more over-represented in terms of AI usage relative to other occupations. Of course, the diffusion of AI in science is closely related to other distributions (and capacity to make investments, e.g. Brynjolfsson et al. (2021)) in computational or intangible capital. Gemini usage scales roughly in proportion to national researcher populations and the creation of specialized AI models is mostly concentrated in a small group of developed economies. Cutting-edge scientific AI often requires complements like compute availability and domain specific human capital. This suggests a risk that without policy action under-resourced or developing regions could fall behind. The frontier of AI-augmented science might require investment to mitigate global disparities in research capacity.
In terms of task analyses, specialized AI models and Gemini usage appear similar. Both have the primary use category in quantitative modeling and data analysis. But when we dig into specific types of analysis, we start seeing evidence of model specialization and a division of labor within the scientific production function. Specialized models function as a kind of domain-specific capital—generating synthetic data, predicting molecular properties, and executing simulations—tasks that seem better served by a specialized model instead of a general LLM (at least for now). Meanwhile, general-purpose models absorb a wider range of tasks like writing code, synthesizing literature and operations. Similar to the way in which researchers specialize across disciplines or workflow components, AI tools may be specializing as well.
Where adopted, AI seems to deliver meaningful productivity dividends. Researchers report substantial net time savings that they recycle back into other projects. However, these changes enabled by AI in science will not only be on the extensive margin. About half of surveyed researchers report that AI directs them toward safer, incremental questions where clean data and benchmarks exist. Only 28 percent instead pursue riskier projects–despite evidence that most important scientific progress relies on high-risk exploration (Azoulay et al., 2011) and novel hypotheses (Uzzi et al., 2013). It remains to be seen whether AI models will lead researchers toward more incremental work, enable “moonshot” projects, or possibly both; future work should examine the relationship between AI usage and the risk profile of scientists. It might be that AI delivers a “Streetlight Effect” (Hoelzemann et al., 2024; Nagaraj and Tranchero, 2023) where the costs of executing incremental work drops relative to higher-risk projects. Incremental work made easy by AI could grow as a proportion of research output overall, at least in the short-run. On the other hand, scientists report that AI enables more interdisciplinary work, potentially acting as a translator helping researchers recombine ideas more effectively (Fang and Evans, 2026) and helping overcome the burden of knowledge (Jones, 2009).
[Footnote 28] This is not necessarily a negative outcome—for example, incremental “normal science” that enhances the understanding of disease mechanisms contributes to evidence on drug development (McNamee et al., 2017).
Despite widespread task-level time savings, macro-level scientific discovery rates remain bound by workflow task dependencies. These hidden scientific “organizational” complementarities are similar to AI’s impact on work in a general sense (Demirer et al., 2026; Gans and Goldfarb, 2026; Kremer, 1993). For example, it may be that computational modeling and data analysis are accelerated by AI tooling or that specialized AI models provide more promising biological structures to test, but the residual tasks in the bundle of scientific work are harder to scale and absorb the time savings. This bundling and task interlock for scientific work means that broader productivity gains may require redrawing job boundaries, and for some types of work the task bundling is stronger than others (Garicano et al., 2026). Enhanced productivity in data analysis might, for example, push the queue of work into hard-to-automate, physical stages that become a bottleneck (Jones, 2025).
Furthermore, AI outputs are not necessarily “turn-key”. Almost 90 percent of researchers who report time savings spend a meaningful share of their saved time verifying AI outputs, and about 46 percent spend more than a quarter of their saved time doing so. As with software engineering, scientists who adopt AI tools may be pushed toward auditing, debugging, and quality screening. Especially in the case of scientific work, being correct is of paramount importance. Lowering the marginal cost of hypothesis generation, code generation, analysis, or testing will create a high complementary marginal value of verifying results (or, perhaps, human judgment e.g., Agrawal et al. (2019)). This high “verification tax” likely arises from the high value of reliable and correct output in science.
The division of labor we document also suggests what the scientist of the near future might look like. Today a researcher moves between tools by hand: a chatbot for code and drafting, AlphaFold or a materials model for prediction, a spreadsheet or a lab notebook in between. The next step may be an LLM that orchestrates the specialized models directly, planning the experiment, calling the right model for each step, checking outputs against each other, and returning a candidate answer for the scientist to judge. The human’s job shifts toward choosing the question, setting the standard of evidence, deciding what is worth the cost of a physical test. Our data say we are not there yet: verification still consumes a large share of the time AI saves and the backlog of untested hypotheses is growing. The hope, in the end, is larger than time savings. When AlphaGo (Silver et al., 2016) played its 37th move against Lee Sedol in 2016, it played a move that its own model estimated a human would make about one time in ten thousand, and it won. The promise of more advanced AI systems in science is that they will eventually make moves of that kind: discoveries beyond the limit of what human researchers would have tried on their own.
Our work here has provided some initial measurements of the adoption of AI tools in science, its potential impact and bottlenecks that could limit this impact. As these limitations arise, so do new opportunities for cross-disciplinary work and new directions of inquiry in computationally demanding domains. Scientific production shares many of the features of other kinds of production. The rapid integration and adoption of AI tools by scientists combined with the substantive importance of breakthrough discoveries makes this an ideal domain to understand opportunities posed by AI. By continuing this work over time, we can watch the experimenters experiment, learning from scientific successes and failures to unlock AI-driven productivity gains in scientific R&D and elsewhere in the economy.
Where Indigo landsFurther
Indigo's conclusion
The hardest first-hand evidence yet that verification is incompressible: AI crushes generation, the bottleneck moves to verification and physical experiment, and 89% of time-savers spend over a tenth of the saved time checking AI output. But it's Gemini grading itself plus a self-reported survey; read the numbers as an early Google snapshot.
What to remember
The bottleneck moves to verification: AI crushes upstream generation, constraints move to verification and physical experiments; 41% report a hypothesis backlog; 89% of time-savers spend over a tenth of it checking AI output.
LLMs and specialized models are economic complements with almost no overlap at the finest level, pointing to LLMs orchestrating specialized models: value lies in orchestration, specialized models and data.
49% are pushed toward safer, incremental problems versus 28% toward riskier ones: verifiable domains speed up, possibly at the cost of risky questions.
Nearly 7 hours saved a week, mostly reinvested in research, 8 in 10 report higher output; but it's self-reported perception, not measurement.
Three discounts: Gemini only, a self-reported non-probability survey, and a classifier only 28% accurate at the finest level. Trust the direction; treat the numbers as a snapshot.
Back on the long-running theses
confirms
Verification is incompressible The hardest first-hand pillar yet: Google measured, with 15 million interactions and 600+ scientists, that AI compresses the upstream, the bottleneck moves down, and verification eats the dividend.
conflicts
Can we scale beyond easy-to-verify fields? 49% moving to safer problems is an early warning: conquering verifiable domains isn't advancing the frontier.
confirms
You don't own the model LLMs orchestrating specialized models, plus data and benchmarks, place the value in orchestration, unique specialized models and data, not the general LLM.
adds to
The politics of open weights Lower-income countries citing specialized models without developing them is first-hand evidence of open models diffusing widely.
confirms
DeepMind Co-Scientist The same honesty: AI helping research is real, but final verification and the physical world are the bottlenecks.