Mind · Weekly

护城河不在模型里,在私有判断数据里

第 014 期 · 2026.06.28 — 2026.07.05

本周的材料排出一场罕见的正面对撞:带着私有数据的专用小模型在窄任务上碾压最强前沿,通用阵营则断言历史终将重演。Indigo 没有选边,他把原本的一条付费铁律拆成一张二维地图——这个拆法本身,比任何单一结论都更接近可复用的判断力。

2026.06.28 — 2026.07.05 · 每周一次,识别信号,认知重调。

本周信号

一周之内,两份第一手材料撞到了一起。桥水(Bridgewater,全球最大对冲基金之一)和 Thinking Machines 联手做了一项研究:用开源模型 Qwen3-235B 做 RL 微调(强化学习微调:拿私有专家数据在特定任务上继续训练模型),跑了六个金融判断类任务,平均正确率 84.7%,比最强的前沿模型 GPT 5.5 还高——GPT 5.5 是 78.2%,错误率低 29.8%,推理成本还便宜 13.8 倍。同一周,Naval 在一场圆桌上重申了苦涩教训(bitter lesson:拉长时间看,吃算力的通用方法总会打败人工雕琢的专用方法),并给出了通用这一侧的量级:推理需求预计会涨到约 90,000 倍,能靠模型本身直接赚钱的玩家,正从五家收敛到 OpenAI 和 Anthropic 两家。

大多数人会把这读成一场谁对谁错的争论,但真正值得注意的信息在别处:两边测的根本不是同一件事。桥水赢在高频、窄域、只能靠人类专家验证的地方——正因为这些领域没法自动打分,私有专家数据才成了稀缺资产,报告里把它称作 differentiated intelligence(差异化智能:护城河在私有数据和判断力上,不在模型本身)。Naval 的逻辑赢在长链条、深递归、公开域的地方,那里误差会一路滚雪球。两边还共享同一块背景板:整条前沿线集体撞在约 78% 的天花板上(Claude Opus 4.8 是 78.0%,还是单价最贵的那个),够不到 80% 的信任门槛。

护城河的判断,正在从选哪个模型下沉到谁握有私有判断数据——通用与专用不是二选一,而是同一张矩阵的两条轴。

风向

#01 误差的复利与 78% 的天花板

这场对撞的一侧,是 Indigo。他本周把圆桌上那条核心算术,亲手写成了一条纪律:一个 99.9% 正确的 AI vs 一个 90% 正确的 AI,单次差距看着小,但递归循环跑 100 次,90% 那个掉正确率会掉到 13%,99.9% 那个还有 80-90%。『智能的误差同样会有复利』—— 所以在判断类、高杠杆任务上,永远为最强的智能付费(原帖)。这条算术背后还藏着一个前提:成本正在坍缩。圆桌举的例子里,同样的产出,每人月成本从 $100 降到了 $2.84——智能越便宜,任务链条就被拉得越长,误差的复利效应也就越致命。但对撞的另一侧立刻补上了限制条件:在桥水的金融任务上,前沿模型集体够不到 80% 的信任门槛,GPT 5.4 比 5.2 贵 43%,换来的只是边际提升;而下定义、提猜想这类最高阶的创造,连奖励函数都写不出来。为最强付费不是一条铁律,而是一条有边界的杠杆纪律。

#02 存储进入了另一种周期

本周传播最广的一条帖子(68 个赞、12.7k 浏览)来自 Indigo 对 AI 半导体的原创推演:HBM 因结构性需求变化和技术快速迭代,进入成长型周期已是市场共识;DRAM 因 Agentic CPU 获得新的结构性指数需求 …… CPU 服务器 TAM 暴涨、每 CPU core 的 DRAM 配比涨 3–4 倍 …… 我对未来 12–24 个月的预测:HBM 上行延续;DRAM 现货价 2026 内大概率续强,但 2027 进入『兑现 vs 透支』的检验 …… NAND 跟涨但波动更大(原帖)。把名词换成人话:HBM 是堆在 AI 加速芯片旁边的高带宽内存,DRAM 是通用服务器内存,NAND 是闪存,Agentic CPU 指 AI 代理(能自主执行多步任务的程序)把海量任务压回 CPU 服务器,TAM 就是整体市场规模。一位与 Indigo 素不相识的外部分析师,在一套 14 层产业链框架里给出了同构的判断:HBM 没有消灭存储周期,而是让它被 AI 带宽瓶颈重新定义了。存储的周期性不是被消灭,是被重新定义——只是部分摆脱,不是彻底摆脱。 检验点钉死在 2027:届时要看 Agentic CPU 对 DRAM 的真实数据拉动;NAND 的结构性动力最弱,也会是链条上最先出问题的一环。

#03 实现变便宜之后,品味成了瓶颈

本周,Indigo 转发并背书了一条关于 agentic coding(让 AI 代理替人写代码的工作方式)的方法论:产出瓶颈从『模型能力』转移到了『你能不能把未知讲清楚』。你给的 prompt/context 是地图,代码/真实约束是领土,差值就是『未知』。减少并预留未知,是 agentic coding 的真功夫,产出 = f(你澄清未知的能力)(原帖)。佐证的数字来自 OpenAI 内部:接近 100% 的员工每周都在用 Codex(OpenAI 的编程代理),用量涨了 6 倍——实现已经不是贵的部分了,贵的是品味。GrantSanderson 给这个存活下来的人类角色起了个名字:策展人,负责决定什么值得说、什么值得放上去。三篇互不相干的材料,收敛到了同一条主轴上:当实现被商品化到近乎免费,价值就会往上移到租不到的东西——品味、判断力,以及把未知讲清楚的能力。品味不是装饰,而是能力商品化之后冒出来的新瓶颈。

现场

#04 自动化的排序判据

自动化的时间表不是按行业排的,是按能不能被碾磨排的。 Dwarkesh 给出的判据是:一个领域被攻克的速度,取决于能不能在确定性、可重放的模拟器里并行跑成千上万条试错轨迹——编程最快,computer use(让 AI 直接操作电脑界面)偏慢,建公司、打官司、交易垫底。而碾磨不到的地方,恰恰是私有专家数据最值钱的地方。

#05 伽罗瓦的百年验证回路

可碾磨性,比可验证性更关键。 GrantSanderson 用伽罗瓦的故事压测了 RLVR(只用可自动验证的奖励信号做强化学习)这条主线:当年的验证器——学院——拒了他的稿,奖励几十年后才结算。最高阶的创造,本就活在任何自动打分体系之外,这给通用阵营的乐观画了一条硬边界。

#06 记忆要蒸回权重

持续学习比的不是谁存的上下文多,是谁能把经验蒸回权重。 蒸回权重(蒸馏:把在岗经验压进模型参数本身)正是桥水这条路线的实证核心;Dwarkesh 补了个刺眼的数字:实验室 30–50% 的算力花在推理上,却对改进模型毫无贡献。Indigo 的追问踩在同一个点上:AI 真的需要有人教吗?好奇心是最好的老师 Curiosity is everything(原帖)。

#07 卖铲人的卖铲人

产业链的盲区,往往藏在卖铲人的卖铲人那一层。 那份 14 层产业链框架补的正是这块盲区,价值在上游:材料(信越、SUMCO、林德)、EDA/IP(Synopsys、Cadence、Arm)、载板(欣兴、Ibiden、Shinko)——模型层打得再热闹,真正收钱的位置始终在瓶颈处。分析师同时也明说:预期已经大多计入价格了。

#08 电力簇的两面账

运营在加速是真的,估值把这份加速提前收走了,也是真的。 Bloom Energy 在手订单约 200 亿,产能从 1GW 冲向年底超过 2GW,和 Oracle 签约后 55 天就点亮了 50+MW;但市值约 930 亿,约 46x P/S(市销率:市值除以年收入),一年涨了超 1500%,CEO 自己都点名瓶颈是天然气供应。卖铲人的逻辑站得住,纪律在于分清运营和预期,别把两者混成一件事。

#09 Meta 的隐性表态

给超额资本开支找盈利点,本身就是一种能力焦虑的信号。 本周 Indigo 另一条高互动原帖(35 个赞、29 条回复)谈的是 Meta 做 NeoCloud(对外出租算力的新型云生意),他的判断是:这更像是给超额 Capex(资本开支)找一个盈利出口,反过来映衬出自家模型能力的不济——是一次站在通用能力才是硬通货这一侧的隐性表态。

慢思考

这一周,Indigo 有一处明确的改判。他原来的纪律是一条铁律:判断类、高杠杆任务,永远为最强的通用智能付费。现在,这条铁律被拆成了两个维度——任务链条的长度,乘以数据的私有程度。链条长、递归深、又落在公开域的任务,仍然为最强的通用模型付费,因为误差会复利;高频、窄域、握有私有数据的任务,反而是微调过的小模型更准也更便宜。而且最强本身也有硬边界:前沿模型集体撞在约 78% 的信任门槛下,最高阶的创造根本写不成奖励函数。触发这次改判的,是桥水和 Naval 两份一手报告在同一周撞在了一起——他自己那条讲误差复利的帖子,恰好站在通用这一侧,被桥水从专用这一侧补上了另外一半。落到方法上:对 AI 应用层公司的判断,从是否用了最强模型这一个维度,升级成了一张二维矩阵;护城河的检查点,也从模型选型下沉到了私有判断数据。

另一面是:本周最有力的反驳,恰恰来自 Naval 那条老判断——历史上每一代通用模型,都吞掉过上一代的专用优化。84.7% 的胜利可能只是时间差,因为它建立在当代前沿集体撞 78% 天花板这个前提之上。能证明这套框架错了的证据其实很具体:下一代通用前沿模型,在同类窄域任务上跨过 80% 的信任门槛,同时推理价格继续坍缩,把专用微调的正确率优势和 13.8 倍的成本优势一起抹平。到那一天,护城河会从私有数据退回算力和研究本身——该修正的就不是某个细节,而是整张矩阵。

Indigo on X

HBM 因结构性需求变化和技术快速迭代,进入成长型周期已是市场共识;DRAM 因 Agentic CPU 获得新的结构性指数需求 …… CPU 服务器 TAM 暴涨、每 CPU core 的 DRAM 配比涨 3–4 倍 …… 我对未来 12–24 个月的预测:HBM 上行延续;DRAM 现货价 2026 内大概率续强,但 2027 进入『兑现 vs 透支』的检验 …… NAND 跟涨但波动更大

出自 @indigox,68 likes

一个 99.9% 正确的 AI vs 一个 90% 正确的 AI,单次差距看着小,但递归循环跑 100 次,90% 那个掉正确率会掉到 13%,99.9% 那个还有 80-90%。『智能的误差同样会有复利』—— 所以在判断类、高杠杆任务上,永远为最强的智能付费

出自 @indigox,34 likes

产出瓶颈从『模型能力』转移到了『你能不能把未知讲清楚』。你给的 prompt/context 是地图,代码/真实约束是领土,差值就是『未知』。减少并预留未知,是 agentic coding 的真功夫,产出 = f(你澄清未知的能力)

出自 @indigox,24 likes

收束

一个思考

上一次通用能力被商品化,是电。电网铺开之后,发电本身很快就不再稀缺了,真正拉开差距的是那些围绕电机重新排布生产线的工厂——同样的电价,走出了不一样的流程。通用模型像电网,专用微调像那条私有产线:电越便宜,懂得怎么用电的私有工艺就越值钱。本周桥水和 Naval 的这场对撞,不过是这个百年剧本在智能上的重演——商品化的部分负责把价格压下去,没被商品化的部分负责把价值攒起来。

一个尝试

拿一张纸,写下你最近反复冒出来的十件工作任务,给每一件打两个分:链条有多长(一步搞定,还是要递归几十步)、数据有多私有(网上查得到,还是只有你的团队才知道)。把这十件事摆进一张二维矩阵:长链条、又落在公开域的那些,交给你能拿到的最强通用模型,并且盯住每一步的错误率——误差是会复利的;窄域、握着私有数据的那些,问自己一个问题:这里的专家判断,能不能写成一份清单、被反复调用?三十分钟后,你会得到一张属于自己的通用与专用地图。

Mind · Weekly

The Moat Isn't in the Model — It's in Private Judgment Data

Issue 014 · 2026.06.28 — 2026.07.05

This week's material lines up a rare head-on collision: a specialized small model armed with private data crushes the strongest frontier model on a narrow task, while the general-purpose camp insists history will repeat itself. Indigo does not pick a side. He splits what used to be a single iron rule about paying for AI into a two-dimensional map — and that split, in itself, comes closer to reusable judgment than any single conclusion.

2026.06.28 — 2026.07.05 · Once a week: spot the signals, recalibrate your thinking.

This Week's Signal

Within one week, two first-hand pieces of material collided. Bridgewater (one of the world's largest hedge funds) and Thinking Machines published a joint study: they took the open-source model Qwen3-235B and did RL fine-tuning (reinforcement-learning fine-tuning: continuing to train a model on a specific task using private expert data), ran six financial-judgment tasks, and got an average accuracy of 84.7% — higher than the strongest frontier model, GPT 5.5, which scored 78.2%, with a 29.8% lower error rate and inference costs 13.8 times cheaper. The same week, at a roundtable, Naval restated the bitter lesson (over a long enough horizon, general methods that consume compute always beat hand-crafted specialized methods) and gave the general side's order of magnitude: inference demand is expected to grow roughly 90,000x, and the players who can make money directly from models themselves are narrowing from five down to two — OpenAI and Anthropic.

Most people will read this as an argument about who is right. But the real information sits elsewhere: the two sides were not measuring the same thing at all. Bridgewater wins in places that are high-frequency, narrow-domain, and verifiable only by human experts — precisely because those domains cannot be auto-scored, private expert data becomes the scarce asset. The report calls it differentiated intelligence (the moat sits in private data and judgment, not in the model itself). Naval's logic wins in places with long chains, deep recursion, and public domains, where errors snowball all the way down. The two sides also share the same backdrop: the entire frontier line collectively hit a ceiling around 78% (Claude Opus 4.8 scored 78.0%, and was still the most expensive per unit), short of the 80% trust threshold.

The moat question is shifting from which model do you pick down to who holds private judgment data — general vs. specialized is not an either/or, but two axes of the same matrix.

Which Way the Wind Blows

#01 Compounding Errors and the 78% Ceiling

On one side of this collision stands Indigo. This week he wrote the roundtable's core arithmetic into a discipline in his own hand: "A 99.9%-correct AI vs. a 90%-correct AI — the gap looks small on a single pass, but run a recursive loop 100 times and the 90% one's accuracy drops to 13%, while the 99.9% one still holds 80-90%. 'Errors in intelligence compound too' — so on judgment-type, high-leverage tasks, always pay for the strongest intelligence" (original post). Behind this arithmetic hides a premise: cost is collapsing. In the roundtable's example, the same output fell from $100 to $2.84 per person-month — the cheaper intelligence gets, the longer task chains stretch, and the deadlier the compounding of errors becomes. But the other side of the collision immediately added the constraint: on Bridgewater's financial tasks, frontier models collectively fell short of the 80% trust threshold — GPT 5.4 costs 43% more than 5.2 for only marginal gains — and for the highest-order creative work, like writing definitions or proposing conjectures, no reward function can even be written. Paying for the strongest is not an iron rule; it is a leverage discipline with boundaries.

#02 Storage Has Entered a Different Kind of Cycle

The week's most widely shared post (68 likes, 12.7k views) was Indigo's original analysis of AI semiconductors: "It is already market consensus that HBM has entered a growth-type cycle, driven by structural demand shifts and rapid technology iteration; DRAM is gaining new structural, exponential demand from Agentic CPUs … CPU server TAM is surging, and DRAM per CPU core is rising 3–4x … My forecast for the next 12–24 months: HBM's upcycle continues; DRAM spot prices most likely stay strong through 2026, but 2027 brings the 'delivery vs. overshoot' test … NAND rises along with them but with more volatility" (original post). In plain terms: HBM is the high-bandwidth memory stacked next to AI accelerator chips, DRAM is general-purpose server memory, NAND is flash storage, Agentic CPU means AI agents (programs that carry out multi-step tasks on their own) pushing huge workloads back onto CPU servers, and TAM is total market size. An outside analyst with no connection to Indigo reached a structurally identical conclusion within a 14-layer supply-chain framework: HBM has not killed the storage cycle — it has let the cycle be redefined by the AI bandwidth bottleneck. Storage's cyclicality is not being eliminated; it is being redefined — a partial escape, not a full one. The test point is nailed to 2027: that is when Agentic CPUs' real data pull on DRAM must show up; NAND has the weakest structural driver and will be the first link in the chain to break.

#03 Once Implementation Gets Cheap, Taste Becomes the Bottleneck

This week Indigo reposted and endorsed a piece of methodology on agentic coding (having AI agents write code in a person's place): "The output bottleneck has shifted from 'model capability' to 'whether you can spell out the unknown.' The prompt/context you give is the map, the code/real constraints are the territory, and the gap between them is the 'unknown.' Reducing unknowns, and budgeting for them, is the real craft of agentic coding. Output = f(your ability to clarify the unknown)" (original post). The supporting numbers come from inside OpenAI: nearly 100% of employees use Codex (OpenAI's coding agent) every week, and usage has grown 6x — implementation is no longer the expensive part; taste is. GrantSanderson gave the surviving human role a name: the curator, who decides what is worth saying and what goes up. Three unrelated pieces of material converged on the same axis: once implementation is commoditized to near-free, value moves up to what cannot be rented — taste, judgment, and the ability to spell out the unknown. Taste is not decoration; it is the new bottleneck that emerges after capability gets commoditized.

On the Ground

#04 The Sorting Rule for Automation

The automation timetable is not sorted by industry; it is sorted by what can be milled. Dwarkesh's criterion: how fast a field gets cracked depends on whether you can run thousands of parallel trial-and-error trajectories in a deterministic, replayable simulator — coding is fastest, computer use (letting AI operate a computer interface directly) is slower, and building companies, litigating, and trading sit at the bottom. And the places the mill cannot reach are exactly where private expert data is worth the most.

#05 Galois and the Hundred-Year Verification Loop

Millability matters more than verifiability. GrantSanderson stress-tested the RLVR mainline (reinforcement learning using only auto-verifiable reward signals) with the story of Galois: the verifier of his day — the academy — rejected his paper, and the reward settled decades later. The highest-order creation lives outside any automatic scoring system. That draws a hard boundary on the general camp's optimism.

#06 Memory Must Be Distilled Back Into Weights

Continuous learning is not about who stores more context; it is about who can distill experience back into weights. Distilling into weights (distillation: compressing on-the-job experience into the model's own parameters) is exactly the empirical core of Bridgewater's route; Dwarkesh added a glaring number: labs spend 30–50% of their compute on inference, which contributes nothing to improving the model. Indigo's follow-up question lands on the same point: Does AI really need to be taught? Curiosity is the best teacher. Curiosity is everything (original post).

#07 The Shovel-Sellers' Shovel-Sellers

The supply chain's blind spot often hides one layer up — with the shovel-sellers' shovel-sellers. That 14-layer supply-chain framework fills exactly this blind spot; the value sits upstream: materials (Shin-Etsu, SUMCO, Linde), EDA/IP (Synopsys, Cadence, Arm), substrates (Unimicron, Ibiden, Shinko) — however loud the model-layer fight gets, the positions that actually collect money stay at the bottlenecks. The analyst also said plainly: most of the expectation is already priced in.

#08 Two Ledgers for the Power Cluster

Operations really are accelerating — and valuation really has collected that acceleration in advance. Both are true. Bloom Energy has roughly 20 billion in backlog, capacity climbing from 1GW toward over 2GW by year-end, and lit up 50+MW just 55 days after signing with Oracle; but the market cap is about 93 billion, roughly 46x P/S (price-to-sales: market cap divided by annual revenue), up more than 1500% in a year, and the CEO himself named natural-gas supply as the bottleneck. The shovel-seller logic holds; the discipline is to keep operations and expectations separate, and not blend the two into one thing.

#09 Meta's Implicit Confession

Hunting for a profit outlet for excess capital spending is itself a signal of capability anxiety. Indigo's other high-engagement original post of the week (35 likes, 29 replies) was about Meta doing NeoCloud (a new cloud business renting out compute). His read: this looks more like finding a profit outlet for excess Capex (capital expenditure), which in turn reflects the inadequacy of Meta's own model capability — an implicit vote for the side that says general capability is the hard currency.

Slow Thinking

This week, Indigo made one clear revision. His old discipline was an iron rule: on judgment-type, high-leverage tasks, always pay for the strongest general intelligence. Now that rule has been split into two dimensions — the length of the task chain, times the privacy of the data. Tasks with long chains, deep recursion, and public-domain footing still deserve the strongest general model, because errors compound; high-frequency, narrow-domain tasks backed by private data are actually better served by fine-tuned small models — more accurate and cheaper. And strongest itself has a hard edge: frontier models collectively hit a wall just under the roughly 78% trust threshold, and the highest-order creation simply cannot be written as a reward function. What triggered the revision was the collision of the Bridgewater and Naval first-hand reports in the same week — his own post on compounding errors happened to stand on the general side, and Bridgewater supplied the other half from the specialized side. In practical terms: judging AI application-layer companies upgraded from one dimension — do they use the strongest model — to a two-dimensional matrix; and the moat checkpoint dropped from model selection down to private judgment data.

The other side of it: the week's strongest rebuttal comes precisely from Naval's old call — every generation of general models in history has swallowed the previous generation's specialized optimizations. The 84.7% victory may just be a timing gap, because it rests on the premise that today's frontier is collectively stuck at the 78% ceiling. The evidence that would prove this framework wrong is actually concrete: the next generation of general frontier models crosses the 80% trust threshold on the same narrow-domain tasks, while inference prices keep collapsing, wiping out both the specialized fine-tune's accuracy edge and its 13.8x cost advantage at once. On that day, the moat retreats from private data back to compute and research itself — and what needs correcting would not be a detail, but the whole matrix.

Indigo on X

"It is already market consensus that HBM has entered a growth-type cycle, driven by structural demand shifts and rapid technology iteration; DRAM is gaining new structural, exponential demand from Agentic CPUs … CPU server TAM is surging, and DRAM per CPU core is rising 3–4x … My forecast for the next 12–24 months: HBM's upcycle continues; DRAM spot prices most likely stay strong through 2026, but 2027 brings the 'delivery vs. overshoot' test … NAND rises along with them but with more volatility"

From @indigox, 68 likes

"A 99.9%-correct AI vs. a 90%-correct AI — the gap looks small on a single pass, but run a recursive loop 100 times and the 90% one's accuracy drops to 13%, while the 99.9% one still holds 80-90%. 'Errors in intelligence compound too' — so on judgment-type, high-leverage tasks, always pay for the strongest intelligence"

From @indigox, 34 likes

"The output bottleneck has shifted from 'model capability' to 'whether you can spell out the unknown.' The prompt/context you give is the map, the code/real constraints are the territory, and the gap between them is the 'unknown.' Reducing unknowns, and budgeting for them, is the real craft of agentic coding. Output = f(your ability to clarify the unknown)"

From @indigox, 24 likes

Closing

One Thought

The last time a general capability got commoditized, it was electricity. Once the grid spread, generating power itself quickly stopped being scarce; what opened the real gap were the factories that rearranged their production lines around the electric motor — same price of power, different processes. General models are the grid; specialized fine-tuning is that private production line: the cheaper power gets, the more valuable the private craft of knowing how to use it becomes. This week's Bridgewater-vs-Naval collision is just that century-old script replaying on intelligence — the commoditized part drives the price down, and the uncommoditized part stores the value up.

One Exercise

Take a sheet of paper and write down the ten work tasks that keep coming back to you lately. Score each one on two axes: how long is the chain (done in one step, or dozens of recursive steps), and how private is the data (findable online, or known only to your team). Place the ten tasks on a two-dimensional matrix: for the long-chain, public-domain ones, hand them to the strongest general model you can get, and watch the error rate at every step — errors compound; for the narrow-domain ones sitting on private data, ask yourself one question: can the expert judgment here be written into a checklist and called on repeatedly? Thirty minutes later, you will have your own map of general vs. specialized.