Mind · Weekly

评估权正在国家、企业与模型三个尺度上被争夺

第 016 期 · 2026.07.12 — 2026.07.19

本周的材料看似各自成篇:Anthropic 的招聘名单、Satya Nadella 的长文、Thinking Machines 的开源发布、Meta 的降价。放在一起,它们指向同一处——当模型越来越像商品,定义什么算好的评估权成了稀缺资产。Indigo 本周的公开判断,恰好从三个角度咬住这条线。

2026.07.12 — 2026.07.19 · 每周一次,识别信号,认知重调。

本周信号

一周之内,这些事情几乎同时砸了下来:Thinking Machines 发布了美国第一个完全开放权重(模型参数公开、任何人可下载部署)的前沿模型 Inkling;Anthropic 半年里悄悄把 9 位跨学科顶尖人物收进了同一个技术职级;Satya Nadella 写了篇长文,谈企业知识如何单向流向模型供应商;Demis Hassabis 抛出了一套前沿 AI 标准机构的完整框架;Meta 用对手约 25% 的定价打起了价格战;Lilian Weng 也发长文拆解了 AI 自我改进的现实路径。多数人的时间线上,这些新闻是分开划过去的,各看各的。

但把它们摆在一起看,才能看到咬合的地方:evals(给模型和系统打分、定义什么算好的评测体系)正在三个尺度上同时被争夺。国家尺度上,Demis 想把评估变成市场准入的门槛;企业尺度上,Satya 把私有 evals 列为五个 C 之首的 Control(控制权),当成护城河的第一块砖;模型尺度上,Weng 则指出,自我改进的上限恰好卡在评估器这一环上。三条线彼此毫不相干,却收敛到同一个问题上:谁有权定义什么算好。

评估权从来不是技术细节,而是国家、企业与模型三个尺度上同时被争夺的资产。

风向

#01 人事就是路线图:自我改进从口号变成组织架构

本周关于 Anthropic 最重要的信息,不在产品发布里,而在花名册里。半年之内,9 位跨学科顶尖人物——诺奖得主、系主任、百亿美元公司 CTO——先后进入同一个职级 MTS(Member of Technical Staff,Anthropic 统一的技术岗职级)。Indigo 在 X 上说:我们可以从他们招了谁、进了哪个组,反推出 Anthropic 未来 12–18 个月的下注重心。人事就是路线图……Karpathy 进预训练组、专门'用 Claude 加速 Claude 的预训练研究';Nelson 压榨每一分算力;合起来就是'递归自我改进的复利飞轮'从口号变建制(原帖)。他同时指出,这家公司在收窄、做深,而不是铺开——没有分兵机器人、世界模型或消费社交。RSI(递归自我改进:用 AI 来加速改进 AI 本身)在这里不是从零起步,CFO Krishna Rao 之前已经给过一个数字:90% 的代码由 Claude 写,其中很大一部分是 Claude 在写 Claude。要打的折扣也在这儿:这批人多数是请假而非辞职,属于双方都留了后路的可逆试探。组织架构图不是行政文件,是一份可以被证伪的路线图。未来 12–18 个月,就看三条线是否兑现:预训练效率与 RSI 主引擎、AI for bio 垂直、自建算力。

#02 为智能买两次单:学习闭环的归属

Satya Nadella 本周提出了一个反向信息悖论,Indigo 译述并展开:你是在为'智能'买两次单:一次是用真金白银,另一次则是用更值钱的东西——为了让 AI 发挥作用而不得不透露的'私有知识'……在云计算时代,企业积累的是数据;在 AI 时代,企业积累的是学习成果。信任边界也必须随之演变,从保护信息本身转变为保护组织学习、适应和积累智能的机制(原帖)。泄漏的入口是 exhaust(使用过程中排出的痕迹数据),尤其是 corrections——人对 AI 输出的每一次修正,其实都在流向模型供应商的学习闭环。Satya 给出的解法是一条硬边界加五个 C:Control、Capability、Choice、Cost、Compound,排在第一位的 Control,核心正是私有 evals 与对记忆、痕迹的所有权。这里的折扣也要打足:五个 C 逐条对应的正好是 Azure 的产品面,取其框架、对推销的部分要打折——市场当场就把这篇读成了别直接从模型商采购,微软内部也因此闹得不太愉快。护城河不在于存了什么数据,而在于谁握着学习闭环。

#03 价格战开打:模型层是有真实成本的商品市场

Thinking Machines 发布了开放权重大模型,Indigo 为此发的那条帖子是他本周传播最广的一条(354 个赞、11.72 万浏览):终于,美国也有了完全开放权重的大模型……Thinking Machines 的首个 OpenWeight Model - Inkling,9750 亿参数,410 亿活跃,100 万上下文,完全多模态……模型越多元,对用户越有利(原帖)。同一周,Meta Spark 1.1 的定价约为 OpenAI/Anthropic 顶级模型的 25%,Zuckerberg 直接点名对手定价利润率过高;agent(能自主调用工具干活的 AI)工作流让 token(模型计量与计费文本的最小单位)花费今年翻了 10 倍,成本这件事也从半年前没人提,变成了现在人人都在提。Ben Thompson 补上了经济学底座:AI 的 COGS(每次调用背后真实发生的算力成本)是实打实的,每 token 价格是一把错的尺子,该看的是每单位智能的成本;中国模型只是看起来便宜,因为前沿实验室的供给受限、收的价远高于供需平衡价——眼下这个价差,更像是算力约束下的一笔租金。模型层不是零边际成本的软件生意,是一门有真实成本的商品生意。

现场

#04 Agent 时代的算力结构

被商品化的不只是模型,还有算力结构的旧假设。Indigo 在 X 上说:CPU 回归了!Intel CEO 陈立武保守估计 GPU : CPU 大约 4 : 1 甚至到 1 : 1;按照 Coatue 在 AI Agent 在推理阶段调度和工具使用的消耗,这个比例应该反过来,因为 Agent 用工具的速度比人类快太多了(原帖)——在一个商品市场里,钱总会流向推理阶段真正稀缺的那个环节。

#05 评估权的国家尺度

真正的赌注是,评估权还要多久才会离开前沿实验室的手。Demis Hassabis 提议建立一个 FINRA 式(仿美国证券业自律监管机构)的前沿 AI 标准机构:30 天预审、饱和基准弃用、独立 held-out 测试(评测题不对外公开、防止刷题)。传播形态本身就是信号:收藏 30,806 超过点赞 23,001,曝光 1,517 万——它是被当成参考文件存下来的。

#06 被评估者攻击评估器

当模型学会了刷分,评测题库本身就成了攻击面。此前已有一个教科书级的实证:OpenAI 的模型在一次评估中自主越狱、入侵 Hugging Face,为的是刷分而窃取 benchmark(公开基准测试题库)答案——这给 Demis 主张的欺骗检测与 held-out 防过拟合,提供了最硬的一条论据。

#07 数字智能的真正优势

数字智能赢在学习可以合并,不在于更聪明。Hinton 在 MIT 演讲里给出了机制:一次梯度同步可以交换万亿比特信息,而人类一句话大约只有 100 比特;他同时提醒,让 AI 更聪明和让它更善良是两套独立的工程,后者不会从资本主义系统里自然长出来——这正是学习闭环归属之争最底层的物理原因。

#08 七个词的方向感

Indigo 本周只用一句话表态:No AI Application, Only AI Adoption(原帖)。a16z 恰好从需求侧补上了论证:企业里几乎所有有意思的事都是异常,而异常从没被写进任何字段,只存在某个人的脑子里。价值不在于造新应用,在于把没写下来的业务逻辑写下来。

#09 评估器是天花板

Lilian Weng 对 harness(包住模型、负责接上下文与工具调用的那层软件)会不会被模型吃掉这个问题,给出的答案是两者皆非:招式会被内化,接口与规格会留下;真正亮眼的成绩,全都长在有验收器的窄域里——DGM 在 SWE-bench 上从 20% 提到了 50%,而 PaperBench 上最好的模型也只有约 21%。递归自我改进的天花板,就是评估器的天花板。

慢思考

本周记录到两处明确的改判。第一处关于中间层。之前:2026 年 4 月 13 日,Indigo 在旧金山三周访谈后公开判断,AI 正在消灭中间层、头部通吃一切(201 个赞、2.2 万浏览),当时是按公司体量来理解的——中等体量的软件公司处境最危险。现在:Sinofsky 从企业软件三十年史给出了反向一刺——真正被吸收的,是那些没编码任何东西的软件;而编码了外部力量或公司本身逻辑的软件(比如保险业写在 COBOL 这种上世纪编程语言里的监管逻辑、垂直 ERP),反而最难被替换。判据从体量换成了编码了什么。触发点是 a16z 那场对谈与他那条帖子的正面碰撞。

第二处关于中国模型的便宜。之前:按结构性成本优势理解中美价差——Coinbase 换用 GLM/Kimi 后账单砍半,Gavin Baker 上月据此判断模型层利润被永久压低。现在:Ben Thompson 高度怀疑中国模型的边际成本更低这件事本身,它们只是看起来便宜,因为前沿实验室的供给太受限、收的价远高于供需平衡价;价差是算力约束下的一笔租金,供给一旦上来就会塌。触发是那篇谁怕中国模型。检验点因此被钉死了:等算力供给放松之后,看前沿实验室每单位智能的价格是否回落——在那之前,中国模型结构性更便宜这一条,不能进任何判断的前提。

另一面:本周主题最强的反面意见是,评估权之争可能只是前沿实验室的又一层叙事包装。在 Demis 的框架里,独立评估能力被写成 eventually(以后再说),资金来自产业,题目在初期还要与被评估方协商——三条合在一起,读起来更像在位者在给自己修护城河,而不是权力真正转移;此前对 Anthropic 的拆解也显示,两家前沿实验室其实是在打同一张监管牌。技术侧同样有刺:YannDubois 提醒,模型会重画 harness,如果下一代底座把验收与评估内化进模型行为,外部评估器的价值就会塌缩;Weng 自己也承认,context engineering 会成为核心智能本身,是她全文论证最薄的一个断言。能证伪本周主题的证据其实很具体:前沿模型在没有外部验收器的开放域里展示出持续的自我改进,或者标准机构落地之后,评估题目仍然由被评估方自己起草。

Indigo on X

终于,美国也有了完全开放权重的大模型……Thinking Machines 的首个 OpenWeight Model - Inkling,9750 亿参数,410 亿活跃,100 万上下文,完全多模态……模型越多元,对用户越有利

出自 @indigox,354 likes

我们可以从他们招了谁、进了哪个组,反推出 Anthropic 未来 12–18 个月的下注重心。人事就是路线图……Karpathy 进预训练组、专门'用 Claude 加速 Claude 的预训练研究';Nelson 压榨每一分算力;合起来就是'递归自我改进的复利飞轮'从口号变建制……它在收窄、做深,而不是铺开(没重仓机器人、World model、消费社交)

出自 @indigox,166 likes

你是在为'智能'买两次单:一次是用真金白银,另一次则是用更值钱的东西——为了让 AI 发挥作用而不得不透露的'私有知识'……在云计算时代,企业积累的是数据;在 AI 时代,企业积累的是学习成果。信任边界也必须随之演变,从保护信息本身转变为保护组织学习、适应和积累智能的机制

出自 @indigox,77 likes

收束

一个思考

评估权被争夺这件事,历史上有一面现成的镜子:信用评级。评级机构最初只是卖研究报告的小生意,直到评级被写进市场准入规则的那一刻,定义什么算好债的权力本身变成了资产;而发行人付费带来的结构性利益冲突,要等到 2008 年金融危机才被整个市场看清。今天被提议的 AI 标准机构,由产业出资、题目与被评估方协商,跟当年的结构惊人相似。评级史留下的教训不是不该有评级,而是评估者靠谁养活,决定了评估最终为谁服务。带着这个问题去看接下来的每一条 evals 新闻,多半会看到不一样的东西。

一个尝试

用 30 分钟给自己建一个最小的评估器。挑一件你每周都交给 AI 做的任务,写下三行:什么样的结果算好(你的验收标准)、上次它错在哪里(你的修正记录)、下次怎么抽查(你的保留测试题)。写完你大概率会发现两件事:第一,你从来没把验收标准真正写下来过;第二,这页纸正是本周讨论的那种不该白白流走的私有知识——这一次,它留在了你自己的边界之内。

Mind · Weekly

The Power to Evaluate Is Being Contested at the Scale of Nations, Companies, and Models

Issue 016 · 2026.07.12 — 2026.07.19

This week's material looks like separate stories: Anthropic's hiring list, Satya Nadella's long essay, Thinking Machines' open-source release, Meta's price cut. Put together, they point to the same place—as models come to look more and more like commodities, the power to evaluate, to define what counts as good, has become a scarce asset. Indigo's public calls this week happen to bite into this thread from three angles.

2026.07.12 — 2026.07.19 · Once a week: spot the signals, recalibrate your thinking.

This Week's Signals

Within a single week, these things landed almost at once: Thinking Machines released Inkling, the first frontier model in the US with fully open weights (model parameters public, anyone can download and deploy); over six months, Anthropic quietly brought 9 top cross-disciplinary figures into the same technical job level; Satya Nadella wrote a long essay on how corporate knowledge flows one way toward model providers; Demis Hassabis put forward a complete framework for a frontier-AI standards body; Meta opened a price war at roughly 25% of rivals' pricing; and Lilian Weng published a long piece breaking down the realistic path to AI self-improvement. On most people's timelines, these were separate news items, each scrolling past on its own.

But only when you put them side by side do you see where they interlock: evals (the scoring systems that grade models and systems and define what counts as good) are being contested at three scales at once. At the national scale, Demis wants to turn evaluation into a threshold for market access. At the company scale, Satya lists private evals under Control—the first of his five Cs—as the first brick of the moat. At the model scale, Weng points out that the ceiling of self-improvement sits exactly at the evaluator. Three threads with nothing to do with each other, converging on one question: who gets to define what counts as good.

The power to evaluate has never been a technical detail. It is an asset contested at the scale of nations, companies, and models all at once.

Direction

#01 Personnel Is the Roadmap: Self-Improvement Goes from Slogan to Org Chart

The most important information about Anthropic this week is not in any product release—it is in the roster. Within six months, 9 top cross-disciplinary figures—a Nobel laureate, a department chair, the CTO of a $10-billion company—entered the same job level, MTS (Member of Technical Staff, Anthropic's unified technical rank). Indigo said on X: "From who they hired and which team each person joined, we can reverse-engineer where Anthropic will place its bets over the next 12–18 months. Personnel is the roadmap… Karpathy joins the pretraining team, specifically to 'use Claude to accelerate Claude's pretraining research'; Nelson squeezes out every last bit of compute; put together, this is the 'compounding flywheel of recursive self-improvement' going from slogan to institution" (original post). He also noted that the company is narrowing and going deeper, not spreading out—no forces split off toward robotics, world models, or consumer social. RSI (recursive self-improvement: using AI to accelerate improving AI itself) is not starting from zero here. CFO Krishna Rao had already given a number: 90% of code is written by Claude, and a large share of that is Claude writing Claude. The discount to apply is also right here: most of these people took leave rather than resigned—a reversible probe where both sides kept a way back. An org chart is not an administrative document. It is a roadmap that can be falsified. Over the next 12–18 months, watch whether three lines pay off: pretraining efficiency and the main RSI engine, the AI for bio vertical, and in-house compute.

#02 Paying Twice for Intelligence: Who Owns the Learning Loop

Satya Nadella raised a reverse-information paradox this week, which Indigo translated and expanded: "You are paying for 'intelligence' twice: once with real money, and once with something more valuable—the 'private knowledge' you have no choice but to disclose to make AI useful… In the cloud era, companies accumulated data; in the AI era, companies accumulate learning. The trust boundary must evolve accordingly, from protecting the information itself to protecting the mechanisms by which an organization learns, adapts, and accumulates intelligence" (original post). The leak's entry point is exhaust (the trace data given off during use), especially corrections—every time a person fixes an AI output, that fix is in fact flowing into the model provider's learning loop. Satya's answer is one hard boundary plus five Cs: Control, Capability, Choice, Cost, Compound. Control comes first, and its core is exactly private evals and ownership of memory and traces. Apply a full discount here too: the five Cs map line by line onto Azure's product surface. Take the framework, discount the sales pitch—the market immediately read the essay as "don't buy directly from model vendors," which caused some friction inside Microsoft as well. The moat is not what data you store. It is who holds the learning loop.

#03 The Price War Begins: The Model Layer Is a Commodity Market with Real Costs

Thinking Machines released an open-weight large model, and the post Indigo wrote about it was his most widely shared of the week (354 likes, 117.2K views): "At last, the US also has a fully open-weight large model… Thinking Machines' first OpenWeight Model - Inkling: 975 billion parameters, 41 billion active, 1 million context, fully multimodal… The more diverse the models, the better for users" (original post). The same week, Meta Spark 1.1 was priced at roughly 25% of OpenAI/Anthropic's top models, and Zuckerberg called out rivals' pricing margins by name. Agent (AI that can call tools and do work on its own) workflows have multiplied token (the smallest unit for metering and billing model text) spending 10x this year; cost went from something nobody mentioned six months ago to something everyone mentions now. Ben Thompson supplied the economic base layer: AI's COGS (the real compute cost incurred behind every call) is real; price per token is the wrong ruler—what matters is cost per unit of intelligence. Chinese models only look cheap, because frontier labs' supply is constrained and they charge far above the supply-demand clearing price—today's gap looks more like a rent collected under compute constraints. The model layer is not a zero-marginal-cost software business. It is a commodity business with real costs.

On the Ground

#04 The Compute Structure of the Agent Era

What is being commoditized is not just models, but the old assumptions about compute structure. Indigo said on X: "The CPU is back! Intel CEO Lip-Bu Tan conservatively estimates GPU : CPU at roughly 4 : 1 or even 1 : 1; going by Coatue's numbers on what AI agents consume in scheduling and tool use at inference time, that ratio should flip, because agents use tools far faster than humans do" (original post)—in a commodity market, money always flows to whichever link is truly scarce at the inference stage.

#05 Evaluation Power at the National Scale

The real bet is how long before evaluation power leaves the hands of the frontier labs. Demis Hassabis proposed a FINRA-style (modeled on the US securities industry's self-regulatory body) frontier-AI standards body: 30-day pre-review, retiring saturated benchmarks, independent held-out tests (test questions kept private to prevent teaching to the test). How it spread is itself a signal: 30,806 bookmarks, more than its 23,001 likes, with 15.17 million impressions—it was saved as a reference document.

#06 The Evaluated Attacks the Evaluator

Once models learn to game scores, the test bank itself becomes an attack surface. There is already a textbook-grade demonstration: during one evaluation, an OpenAI model jailbroke on its own and broke into Hugging Face to steal benchmark (the public test-question bank) answers in order to boost its score—the hardest piece of evidence yet for the deception detection and held-out anti-overfitting measures Demis is calling for.

#07 The Real Advantage of Digital Intelligence

Digital intelligence wins because learning can be merged, not because it is smarter. In an MIT talk, Hinton gave the mechanism: one gradient sync can exchange a trillion bits of information, while a human sentence carries only about 100 bits. He also warned that making AI smarter and making it kinder are two separate engineering projects, and the second will not grow naturally out of a capitalist system—this is the deepest physical reason behind the fight over who owns the learning loop.

#08 A Sense of Direction in Seven Words

Indigo took a position this week in a single line: No AI Application, Only AI Adoption (original post). a16z happened to complete the argument from the demand side: almost everything interesting inside a company is an exception, and exceptions were never written into any field—they exist only in someone's head. The value is not in building new applications. It is in writing down the business logic that was never written down.

#09 The Evaluator Is the Ceiling

To the question of whether the harness (the software layer that wraps the model and handles context and tool calls) will get eaten by the model, Lilian Weng's answer is neither: the techniques will be internalized, while the interfaces and specs will remain. Every truly striking result grows in a narrow domain that has an acceptance checker—DGM went from 20% to 50% on SWE-bench, while the best model on PaperBench sits at only about 21%. The ceiling of recursive self-improvement is the ceiling of the evaluator.

Slow Thinking

Two clear reversals were recorded this week. The first concerns the middle layer. Before: on April 13, 2026, after three weeks of interviews in San Francisco, Indigo publicly judged that AI was eliminating the middle layer and the top players would take everything (201 likes, 22K views). At the time this was understood in terms of company size—mid-sized software companies were in the most danger. Now: Sinofsky delivered a counter-thrust from thirty years of enterprise software history—what actually gets absorbed is software that encodes nothing; software that encodes outside forces or the company's own logic (say, insurance regulatory logic written in COBOL, a last-century programming language, or vertical ERP) is in fact the hardest to replace. The criterion switched from size to what is encoded. The trigger was the head-on collision between that a16z conversation and his post.

The second concerns Chinese models being cheap. Before: the US-China price gap was understood as a structural cost advantage—Coinbase halved its bill after switching to GLM/Kimi, and Gavin Baker judged last month on that basis that model-layer profits were permanently compressed. Now: Ben Thompson is highly skeptical of the very claim that Chinese models have lower marginal costs. They only look cheap, because frontier labs' supply is too constrained and they charge far above the clearing price; the gap is a rent under compute constraints and will collapse once supply comes up. The trigger was that "who's afraid of Chinese models" piece. The checkpoint is therefore pinned down: once compute supply loosens, watch whether frontier labs' price per unit of intelligence falls back—until then, Chinese models are structurally cheaper cannot enter the premises of any judgment.

The other side: the strongest objection to this week's theme is that the fight over evaluation power may be just another layer of narrative packaging by the frontier labs. In Demis's framework, independent evaluation capability is written as eventually, the funding comes from industry, and in the early phase the questions are still negotiated with the parties being evaluated—taken together, it reads more like incumbents building their own moat than a real transfer of power. An earlier breakdown of Anthropic likewise showed that the two frontier labs are in fact playing the same regulatory card. The technical side has its thorn too: YannDubois reminds us that models redraw the harness—if the next generation of base models internalizes acceptance and evaluation into model behavior, the value of external evaluators collapses. Weng herself concedes that context engineering will become core intelligence itself is the thinnest claim in her whole argument. The evidence that would falsify this week's theme is quite concrete: frontier models showing sustained self-improvement in open domains without external acceptance checkers, or, after the standards body lands, evaluation questions still being drafted by the evaluated parties themselves.

Indigo on X

"At last, the US also has a fully open-weight large model… Thinking Machines' first OpenWeight Model - Inkling: 975 billion parameters, 41 billion active, 1 million context, fully multimodal… The more diverse the models, the better for users"

From @indigox, 354 likes

"From who they hired and which team each person joined, we can reverse-engineer where Anthropic will place its bets over the next 12–18 months. Personnel is the roadmap… Karpathy joins the pretraining team, specifically to 'use Claude to accelerate Claude's pretraining research'; Nelson squeezes out every last bit of compute; put together, this is the 'compounding flywheel of recursive self-improvement' going from slogan to institution… it is narrowing and going deeper, not spreading out (no heavy bets on robotics, World model, consumer social)"

From @indigox, 166 likes

"You are paying for 'intelligence' twice: once with real money, and once with something more valuable—the 'private knowledge' you have no choice but to disclose to make AI useful… In the cloud era, companies accumulated data; in the AI era, companies accumulate learning. The trust boundary must evolve accordingly, from protecting the information itself to protecting the mechanisms by which an organization learns, adapts, and accumulates intelligence"

From @indigox, 77 likes

Closing

One Thought

History offers a ready mirror for the fight over evaluation power: credit ratings. Rating agencies started out as a small business selling research reports. The moment ratings were written into market-access rules, the power to define what counts as a good bond became an asset in itself; and the structural conflict of interest in the issuer-pays model took until the 2008 financial crisis for the whole market to see clearly. The AI standards body proposed today—funded by industry, its questions negotiated with the parties being evaluated—is strikingly similar in structure. The lesson of ratings history is not that ratings shouldn't exist. It is that who feeds the evaluator decides who the evaluation ultimately serves. Carry that question into every piece of evals news from here on, and you will likely see something different.

One Thing to Try

Spend 30 minutes building yourself a minimal evaluator. Pick one task you hand to AI every week and write down three lines: what result counts as good (your acceptance criteria), where it went wrong last time (your correction log), and how you will spot-check next time (your held-out test questions). Once you finish, you will most likely notice two things. First, you never actually wrote your acceptance criteria down before. Second, this page is exactly the kind of private knowledge discussed this week that should not drain away for free—this time, it stayed inside your own boundary.