Mind · In / Out · In · 文章

AI 在数学中的严重错位

A Severe Misalignment of AI in Mathematics

mathandai.org · 2026-09-11

约 25 位菲尔兹奖得主联署:解题只是手段,批量生产真假命题会毁掉孕育新想法的土壤。

Indigo 的结论

即使机器验证了命题为真、优先权也理清了,数学的价值仍不在那个真假答案里,而在理解、提问的能力和人与人的传承。它不争事实对错,争的是事实对错是不是重点。

怎么读这篇 数学界的集体声明,立场鲜明:反对把解数学题当成 AI 的跑分来冲。这也是一份行业立场文件,真诚的科学关切和对本行身份、劳动的自我保护交织在一起,两面都要看到。它发在 OpenAI 称解开 Navier–Stokes、Buckmaster 优先权争议之后几天。

需要记住的几件事

  1. 一个领域有便宜的机器终审,反而更容易被 AI 用成批的正确输出掏空。
  2. 看任何「AI 攻下了 X」的说法,先问:X 里被验证的那一层,是不是 X 真正的价值所在。
  3. 真诚的科学关切和行业自保可以同时成立,两面都记下,不偏向任何一面。

拆解 · 3 步

  1. 01

    AI 公司和数学界的目标严重错位

    大模型已经能解多个领域的重大未决问题,但把解题当跑分来冲,对这门科学有害。 读这一段原文 →

  2. 02

    解题只是手段,目标是理解和洞见

    难题历来是地标和灯塔;解开之后,要经过讲座、讨论、简化,才沉淀成教科书。 读这一段原文 →

  3. 03

    不是反 AI,是要人来做决定

    把数学的问题推广到所有智力工作:走向取决于掌控这项技术的人怎么决定。 读这一段原文 →

什么会让我改口

AI 产出的证明被数学家接手、吸收进数学的主干,人与人的传承没有断。

怎么读这篇

数学界的集体声明,立场鲜明:反对把解数学题当成 AI 的跑分来冲。这也是一份行业立场文件,真诚的科学关切和对本行身份、劳动的自我保护交织在一起,两面都要看到。它发在 OpenAI 称解开 Navier–Stokes、Buckmaster 优先权争议之后几天。

拆解 · 3 步
  1. AI 公司和数学界的目标严重错位
  2. 解题只是手段,目标是理解和洞见
  3. 不是反 AI,是要人来做决定
01

AI 公司和数学界的目标严重错位

大模型已经能解多个领域的重大未决问题,但把解题当跑分来冲,对这门科学有害。

发表于 2026 年 9 月 11 日 · DOI 10.5281/zenodo.22737750

过去几个月里,大语言模型的数学能力突飞猛进,已经到了能解出许多数学分支中重大未决问题的地步。然而,AI 公司把"解数学题"当作 benchmark 来冲的这种推进方式,对数学这门科学、对数学共同体都是有害的。AI 公司的目标和数学共同体的目标严重错位。我们认为,这是影响其他科学与创造性职业、乃至整个社会的更广泛对齐问题的一部分。

02

解题只是手段,目标是理解和洞见

难题历来是地标和灯塔;解开之后,要经过讲座、讨论、简化,才沉淀成教科书。

研究数学处理的是形状、数与自然现象的基本结构。经过一代又一代人的积累,它建起了一整套庞大而精致的想法、方法、抽象和工具,用来理解数学的地貌。反过来,现代技术与科学正是建立在数学工具之上。

著名难题常常充当地标和灯塔,人们借它们来衡量自己对这片地貌的理解是否有所长进。解开其中一道题,一向是新洞见与有趣方法的确证;随后,一个数学家群体会通过漫长而艰苦的讲座、讨论与简化过程,把它研究透。理想情况下,这个过程的终点是一种教科书式的呈现,任何研究生、甚至本科生都能拿来学。有些数学想法还会继续往前走,在几十年甚至几个世纪之后,变成被所有人理解和使用的工具。

数学共同体在许多方面像是人类的一个微缩版。它由采用各种不同路数的个体组成,被一组核心价值连在一起。我们这个行当最宝贵的资源是学生和想法,而我们极其用心地培育它们。我们觉得自己有责任让它们长到饱满,直到能在数学世界里过上属于自己的生活。给学生出题时,我们的核心用意往往是培养技能,让他们在研究上、以及在别的地方,都处在能有所推进的位置上。我们的想法则通过讲座、私下讨论和仔细的书面工作传播出去,并与他人先前的想法连接起来。这些过程无一例外都需要时间,而且建立在人与人的互动之上。

最近几个月,AI 解开重大数学问题的成功,连数学圈之外都上了头条。但解题只是一个工具和代理,真正的首要目标是概念性的理解与洞见。在 AI 的世界里忘掉这一点,可能会让工具反过来对付这个首要目标。事实上,以越来越快的速度批量生产"真/假"命题,可能会摧毁那片沃土,而不是给新想法注入生命。

这些解答常常是在匆忙中宣布的,来不及好好写下来,来不及把新方法和新思想分离出来,也来不及引用他人相关的前作。和所有创造性职业里一样,这引出了严重的归属与剽窃问题。更何况,如果没有愿意接手的数学家去照料它们的发展、把它们整合进数学的正典,AI 构想出来的想法永远不会真正活过来,数学家之间那条关键的人类传承链也会就此断掉。

03

不是反 AI,是要人来做决定

把数学的问题推广到所有智力工作:走向取决于掌控这项技术的人怎么决定。

我们正在见证一种对智力工作的普遍威胁:使用 AI 的结果,与使用它的初衷之间出现了错位。在许多领域和活动中,多年的训练传统上不只是为了产出一个最终答案或产品,也是为了养成理解力,以及提出新问题、新想法的能力。然而,AI 系统建立在人类此前庞大的工作之上,正越来越有能力直接产出这类工作的结果,于是这两个目标不再重合。数学共同体眼下面对的问题,与其他科学和创造性职业正在面对的问题类似,也预示着全人类可能都要面对的问题:当 AI 改变工作的方式时,我们要怎样确保自己没有忘掉这份工作原本是为了达成什么。

AI 确实有潜力增强并加速真正的数学研究与理解。数学作为一门职业,将需要在好几个方面适应这些变化。但这些变化最终是让这个领域受益,还是产生破坏性的后果,很大程度上取决于掌控这项新技术的那些人的决定。

这些问题必须被紧迫地处理——在数学共同体内部,由开发这些技术的公司,以及更广泛地,由一个将在许多其他形式的智力工作中遭遇同类问题的社会来处理。

Artur Avila(菲尔兹奖 2014)Manjul Bhargava(菲尔兹奖 2014)Caucher Birkar(菲尔兹奖 2018)Pierre Deligne(菲尔兹奖 1978)Yu Deng(菲尔兹奖 2026)Simon Donaldson(菲尔兹奖 1986)Vladimir Drinfeld(菲尔兹奖 1990)Hugo Duminil-Copin(菲尔兹奖 2022)Alessio Figalli(菲尔兹奖 2018)Martin Hairer(菲尔兹奖 2014)June Huh(菲尔兹奖 2022)Maxim Kontsevich(菲尔兹奖 1998)Elon Lindenstrauss(菲尔兹奖 2010)Pierre-Louis Lions(菲尔兹奖 1994)James Maynard(菲尔兹奖 2022)Curtis McMullen(菲尔兹奖 1998)Shigefumi Mori(菲尔兹奖 1990)Ngô Bảo Châu(菲尔兹奖 2010)Andrei Okounkov(菲尔兹奖 2006)Peter Scholze(菲尔兹奖 2018)Stanislav Smirnov(菲尔兹奖 2010)Terence Tao(菲尔兹奖 2006)Maryna Viazovska(菲尔兹奖 2022)Cédric Villani(菲尔兹奖 2010)Wendelin Werner(菲尔兹奖 2006)Efim Zelmanov(菲尔兹奖 1994)

判断收口延伸

Indigo 的结论

即使机器验证了命题为真、优先权也理清了,数学的价值仍不在那个真假答案里,而在理解、提问的能力和人与人的传承。它不争事实对错,争的是事实对错是不是重点。

需要记住的几件事

  1. 一个领域有便宜的机器终审,反而更容易被 AI 用成批的正确输出掏空。
  2. 看任何「AI 攻下了 X」的说法,先问:X 里被验证的那一层,是不是 X 真正的价值所在。
  3. 真诚的科学关切和行业自保可以同时成立,两面都记下,不偏向任何一面。

放回主线

冲突+补充

验证不可压缩:瓶颈·护城河·断点 补一个反转:在有便宜机器终审的领域,能被验证恰恰是被掏空的原因,价值可能不在验证这一层。

冲突

可验证域能否泛化 这条判断在问可验证的领域会不会扩大;声明是反向证据:领域被攻下,不等于人的价值被接住。

冲突+补充

OpenAI 称解决 Navier–Stokes 千禧难题 声明不争 Navier–Stokes 的数学对错,争的是这种解题方式在伤害数学。

补充

Buckmaster 声明:LLM 助推流体 PDE 新 blowup 从一位学者的一手操守指控,升级为整个数学界最高权威的集体立场。

冲突

Anthropic:Claude 把费马大定理形式化进 Lean 机器终审让「真」变便宜,也让「真」贬值;两篇对照读,判断机器可验证是护城河还是陷阱。

什么会让我改口

AI 产出的证明被数学家接手、吸收进数学的主干,人与人的传承没有断。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Essay

A Severe Misalignment of AI in Mathematics

mathandai.org · 2026-09-11

About 25 Fields medalists sign on: solving problems is only a means, and mass-producing true/false statements destroys the ground new ideas grow in.

Indigo's conclusion

Even with a statement machine-checked as true and priority settled, math's value isn't in the true/false answer; it is in understanding, the ability to ask questions, and what passes between people. It doesn't dispute the facts; it disputes whether the facts are the point.

How to read this A collective statement from the math community with a clear stance: against treating math problems as benchmarks for AI to chase. It is also a guild position paper, where real scientific concern is woven with protecting the profession's standing and work; see both. It came days after OpenAI's Navier–Stokes claim and the Buckmaster priority dispute.

What to remember

  1. A field with a cheap machine judge is easier for AI to hollow out with correct output in bulk.
  2. For any “AI has conquered X” story, ask first: is the layer of X that gets verified where X's real value lies?
  3. Real scientific concern and a profession protecting itself can both be true; note both and lean toward neither.

Breakdown · 3 steps

  1. 01

    AI companies and mathematicians want very different things

    Language models can now solve major open problems in several fields, but chasing solutions as benchmarks harms the science. Read this part →

  2. 02

    Solving is a means; the goal is understanding

    Hard problems have always been landmarks and beacons. After a solution come lectures, discussion and simplification before it settles into textbooks. Read this part →

  3. 03

    Not anti-AI: people have to decide

    They widen math's problem to all intellectual work: where it goes depends on the decisions of those who control the technology. Read this part →

What would change my mind

proofs produced by AI are taken up by mathematicians and absorbed into the core of the field, with the handing-on between people unbroken.

How to read this

A collective statement from the math community with a clear stance: against treating math problems as benchmarks for AI to chase. It is also a guild position paper, where real scientific concern is woven with protecting the profession's standing and work; see both. It came days after OpenAI's Navier–Stokes claim and the Buckmaster priority dispute.

Breakdown · 3 steps
  1. AI companies and mathematicians want very different things
  2. Solving is a means; the goal is understanding
  3. Not anti-AI: people have to decide
01

AI companies and mathematicians want very different things

Language models can now solve major open problems in several fields, but chasing solutions as benchmarks harms the science.

Published 11 September 2026 · DOI 10.5281/zenodo.22737750

Over the last few months, the mathematical capabilities of LLMs have improved dramatically, to the point that they can solve major outstanding problems in many fields of mathematics. However, the push by AI companies to solve mathematical problems as a benchmark is detrimental to the science of mathematics, and to the mathematical community. The goals of the AI companies and the goals of the mathematical community are severely misaligned. We see these as part of broader alignment issues impacting other scientific and creative professions, as well as the whole of society.

02

Solving is a means; the goal is understanding

Hard problems have always been landmarks and beacons. After a solution come lectures, discussion and simplification before it settles into textbooks.

Research mathematics deals with understanding basic structures of shapes, numbers, and natural phenomena. Over the course of generations, it has built a large corpus of sophisticated ideas, methods, abstractions, and other tools to comprehend the mathematical landscape. In turn, modern technologies and sciences are based on mathematical tools.

Famous problems have often served as landmarks and lighthouses against which one can measure an improved understanding of this landscape. Solving one of these problems has been a certain sign of new insights and interesting methods, which would then be studied by a community of mathematicians, through a long and arduous process of talks, discussions, simplifications. At the end of this process, one will ideally find a textbook presentation of the results suitable for any graduate or even undergraduate student to study. Some of the mathematical ideas pursue their journey even further to become, decades or centuries after, tools that are understood and used by the whole population.

The mathematical community functions, in many ways, as a miniature version of humanity. It consists of individuals using a wide variety of different approaches, joined by core values. The most precious resources of our profession are students and ideas, and these we nurture with great care. We feel responsible to let them grow to their full potential, until they can live a life of their own in the mathematical world. For students we often suggest problems with the core intention of developing skills making them well-positioned for advances in research and elsewhere. Our ideas we disseminate in talks, private discussions and careful writeups, connecting them to the previous ideas of others. These processes invariably take time and are based on human interaction.

In recent months, the success of AI in solving major mathematical problems has made headlines even outside mathematical circles. But solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight. Forgetting this in the world of AI may turn the tool against the primary goal. Indeed, the mass production at faster and faster pace of "true/false" statements could destroy fertile ground instead of breathing life into new ideas.

Often these solutions are announced in a rush, leaving no time for a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others. As in all creative professions, this raises severe attribution and plagiarism questions. Moreover, without the willing mathematicians who must take care of their development and integration into the mathematical canon, AI-conceived ideas would never become fully alive and the crucial human transmission chain between mathematicians would be lost.

03

Not anti-AI: people have to decide

They widen math's problem to all intellectual work: where it goes depends on the decisions of those who control the technology.

We are witnessing a general threat to intellectual work, with misalignment between the outcome of the use of AI and its initial purpose. In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas. However, building on a vast body of previous human work, AI systems are becoming increasingly capable of producing the results of such work directly, and these goals cease to align. The issues the mathematical community faces now are similar to issues that other scientific and creative professions are facing, and indicate issues that all of humanity might face: how to make sure that, as AI changes the way work is done, we do not lose sight of what that work was meant to achieve in the first place.

AI offers the potential of enhancing and accelerating genuine mathematical study and understanding. Mathematics as a profession will need to adapt to these changes in several ways. However, whether these changes ultimately benefit the field or have a destructive effect will in large part be determined by the decisions of the humans in control of this new technology.

These issues must be addressed urgently, in the mathematical community, by the companies developing these technologies and, more broadly, by a society that will confront similar problems in many other forms of intellectual work.

Artur Avila (Fields Medal 2014) Manjul Bhargava (Fields Medal 2014) Caucher Birkar (Fields Medal 2018) Pierre Deligne (Fields Medal 1978) Yu Deng (Fields Medal 2026) Simon Donaldson (Fields Medal 1986) Vladimir Drinfeld (Fields Medal 1990) Hugo Duminil-Copin (Fields Medal 2022) Alessio Figalli (Fields Medal 2018) Martin Hairer (Fields Medal 2014) June Huh (Fields Medal 2022) Maxim Kontsevich (Fields Medal 1998) Elon Lindenstrauss (Fields Medal 2010) Pierre-Louis Lions (Fields Medal 1994) James Maynard (Fields Medal 2022) Curtis McMullen (Fields Medal 1998) Shigefumi Mori (Fields Medal 1990) Ngô Bảo Châu (Fields Medal 2010) Andrei Okounkov (Fields Medal 2006) Peter Scholze (Fields Medal 2018) Stanislav Smirnov (Fields Medal 2010) Terence Tao (Fields Medal 2006) Maryna Viazovska (Fields Medal 2022) Cédric Villani (Fields Medal 2010) Wendelin Werner (Fields Medal 2006) Efim Zelmanov (Fields Medal 1994)

Where Indigo landsFurther

Indigo's conclusion

Even with a statement machine-checked as true and priority settled, math's value isn't in the true/false answer; it is in understanding, the ability to ask questions, and what passes between people. It doesn't dispute the facts; it disputes whether the facts are the point.

What to remember

  1. A field with a cheap machine judge is easier for AI to hollow out with correct output in bulk.
  2. For any “AI has conquered X” story, ask first: is the layer of X that gets verified where X's real value lies?
  3. Real scientific concern and a profession protecting itself can both be true; note both and lean toward neither.

Back on the long-running theses

conflicts + adds to

Verification can't be compressed: bottleneck, moat, breaking point A reversal: where a cheap machine judge exists, being checkable is why a field gets hollowed out, and value may not sit in the verified layer.

conflicts

Whether verifiable domains generalize This view asks whether checkable domains will spread; the declaration is counter-evidence: taking a domain doesn't mean taking over the human value in it.

conflicts + adds to

OpenAI says it solved the Navier–Stokes Millennium Problem The declaration doesn't dispute the math; it says this way of solving problems harms mathematics.

adds to

Buckmaster statement: LLM-assisted new blowups in fluid PDEs One scholar's first-hand complaint about conduct becomes the collective stance of the field's highest authorities.

conflicts

Anthropic: Claude formalizes Fermat's Last Theorem in Lean A machine judge makes truth cheap, and cheapens it. Read both to decide whether machine checking is a moat or a trap.

What would change my mind

proofs produced by AI are taken up by mathematicians and absorbed into the core of the field, with the handing-on between people unbroken.

Finished. Indigo's take on this piece is in two places: