This week's material lines up a rare head-on collision: a specialized small model armed with private data crushes the strongest frontier model on a narrow task, while the general-purpose camp insists history will repeat itself. Indigo does not pick a side. He splits what used to be a single iron rule about paying for AI into a two-dimensional map — and that split, in itself, comes closer to reusable judgment than any single conclusion.
2026.06.28 — 2026.07.05 · Once a week: spot the signals, recalibrate your thinking.
This Week's Signal
Within one week, two first-hand pieces of material collided. Bridgewater (one of the world's largest hedge funds) and Thinking Machines published a joint study: they took the open-source model Qwen3-235B and did RL fine-tuning (reinforcement-learning fine-tuning: continuing to train a model on a specific task using private expert data), ran six financial-judgment tasks, and got an average accuracy of 84.7% — higher than the strongest frontier model, GPT 5.5, which scored 78.2%, with a 29.8% lower error rate and inference costs 13.8 times cheaper. The same week, at a roundtable, Naval restated the bitter lesson (over a long enough horizon, general methods that consume compute always beat hand-crafted specialized methods) and gave the general side's order of magnitude: inference demand is expected to grow roughly 90,000x, and the players who can make money directly from models themselves are narrowing from five down to two — OpenAI and Anthropic.
Most people will read this as an argument about who is right. But the real information sits elsewhere: the two sides were not measuring the same thing at all. Bridgewater wins in places that are high-frequency, narrow-domain, and verifiable only by human experts — precisely because those domains cannot be auto-scored, private expert data becomes the scarce asset. The report calls it differentiated intelligence (the moat sits in private data and judgment, not in the model itself). Naval's logic wins in places with long chains, deep recursion, and public domains, where errors snowball all the way down. The two sides also share the same backdrop: the entire frontier line collectively hit a ceiling around 78% (Claude Opus 4.8 scored 78.0%, and was still the most expensive per unit), short of the 80% trust threshold.
The moat question is shifting from which model do you pick
down to who holds private judgment data — general vs. specialized is not an either/or, but two axes of the same matrix.
Which Way the Wind Blows
#01 Compounding Errors and the 78% Ceiling
On one side of this collision stands Indigo. This week he wrote the roundtable's core arithmetic into a discipline in his own hand: "A 99.9%-correct AI vs. a 90%-correct AI — the gap looks small on a single pass, but run a recursive loop 100 times and the 90% one's accuracy drops to 13%, while the 99.9% one still holds 80-90%. 'Errors in intelligence compound too' — so on judgment-type, high-leverage tasks, always pay for the strongest intelligence" (original post). Behind this arithmetic hides a premise: cost is collapsing. In the roundtable's example, the same output fell from $100 to $2.84 per person-month — the cheaper intelligence gets, the longer task chains stretch, and the deadlier the compounding of errors becomes. But the other side of the collision immediately added the constraint: on Bridgewater's financial tasks, frontier models collectively fell short of the 80% trust threshold — GPT 5.4 costs 43% more than 5.2 for only marginal gains — and for the highest-order creative work, like writing definitions or proposing conjectures, no reward function can even be written. Paying for the strongest is not an iron rule; it is a leverage discipline with boundaries.
#02 Storage Has Entered a Different Kind of Cycle
The week's most widely shared post (68 likes, 12.7k views) was Indigo's original analysis of AI semiconductors: "It is already market consensus that HBM has entered a growth-type cycle, driven by structural demand shifts and rapid technology iteration; DRAM is gaining new structural, exponential demand from Agentic CPUs … CPU server TAM is surging, and DRAM per CPU core is rising 3–4x … My forecast for the next 12–24 months: HBM's upcycle continues; DRAM spot prices most likely stay strong through 2026, but 2027 brings the 'delivery vs. overshoot' test … NAND rises along with them but with more volatility" (original post). In plain terms: HBM is the high-bandwidth memory stacked next to AI accelerator chips, DRAM is general-purpose server memory, NAND is flash storage, Agentic CPU means AI agents (programs that carry out multi-step tasks on their own) pushing huge workloads back onto CPU servers, and TAM is total market size. An outside analyst with no connection to Indigo reached a structurally identical conclusion within a 14-layer supply-chain framework: HBM has not killed the storage cycle — it has let the cycle be redefined by the AI bandwidth bottleneck. Storage's cyclicality is not being eliminated; it is being redefined — a partial escape, not a full one. The test point is nailed to 2027: that is when Agentic CPUs' real data pull on DRAM must show up; NAND has the weakest structural driver and will be the first link in the chain to break.
#03 Once Implementation Gets Cheap, Taste Becomes the Bottleneck
This week Indigo reposted and endorsed a piece of methodology on agentic coding (having AI agents write code in a person's place): "The output bottleneck has shifted from 'model capability' to 'whether you can spell out the unknown.' The prompt/context you give is the map, the code/real constraints are the territory, and the gap between them is the 'unknown.' Reducing unknowns, and budgeting for them, is the real craft of agentic coding. Output = f(your ability to clarify the unknown)" (original post). The supporting numbers come from inside OpenAI: nearly 100% of employees use Codex (OpenAI's coding agent) every week, and usage has grown 6x — implementation is no longer the expensive part; taste is. GrantSanderson gave the surviving human role a name: the curator, who decides what is worth saying and what goes up. Three unrelated pieces of material converged on the same axis: once implementation is commoditized to near-free, value moves up to what cannot be rented — taste, judgment, and the ability to spell out the unknown. Taste is not decoration; it is the new bottleneck that emerges after capability gets commoditized.
On the Ground
#04 The Sorting Rule for Automation
The automation timetable is not sorted by industry; it is sorted by what can be milled.
Dwarkesh's criterion: how fast a field gets cracked depends on whether you can run thousands of parallel trial-and-error trajectories in a deterministic, replayable simulator — coding is fastest, computer use (letting AI operate a computer interface directly) is slower, and building companies, litigating, and trading sit at the bottom. And the places the mill cannot reach are exactly where private expert data is worth the most.
#05 Galois and the Hundred-Year Verification Loop
Millability matters more than verifiability. GrantSanderson stress-tested the RLVR mainline (reinforcement learning using only auto-verifiable reward signals) with the story of Galois: the verifier of his day — the academy — rejected his paper, and the reward settled decades later. The highest-order creation lives outside any automatic scoring system. That draws a hard boundary on the general camp's optimism.
#06 Memory Must Be Distilled Back Into Weights
Continuous learning is not about who stores more context; it is about who can distill experience back into weights. Distilling into weights (distillation: compressing on-the-job experience into the model's own parameters) is exactly the empirical core of Bridgewater's route; Dwarkesh added a glaring number: labs spend 30–50% of their compute on inference, which contributes nothing to improving the model. Indigo's follow-up question lands on the same point: Does AI really need to be taught? Curiosity is the best teacher. Curiosity is everything
(original post).
#07 The Shovel-Sellers' Shovel-Sellers
The supply chain's blind spot often hides one layer up — with the shovel-sellers' shovel-sellers. That 14-layer supply-chain framework fills exactly this blind spot; the value sits upstream: materials (Shin-Etsu, SUMCO, Linde), EDA/IP (Synopsys, Cadence, Arm), substrates (Unimicron, Ibiden, Shinko) — however loud the model-layer fight gets, the positions that actually collect money stay at the bottlenecks. The analyst also said plainly: most of the expectation is already priced in.
#08 Two Ledgers for the Power Cluster
Operations really are accelerating — and valuation really has collected that acceleration in advance. Both are true. Bloom Energy has roughly 20 billion in backlog, capacity climbing from 1GW toward over 2GW by year-end, and lit up 50+MW just 55 days after signing with Oracle; but the market cap is about 93 billion, roughly 46x P/S (price-to-sales: market cap divided by annual revenue), up more than 1500% in a year, and the CEO himself named natural-gas supply as the bottleneck. The shovel-seller logic holds; the discipline is to keep operations and expectations separate, and not blend the two into one thing.
#09 Meta's Implicit Confession
Hunting for a profit outlet for excess capital spending is itself a signal of capability anxiety. Indigo's other high-engagement original post of the week (35 likes, 29 replies) was about Meta doing NeoCloud (a new cloud business renting out compute). His read: this looks more like finding a profit outlet for excess Capex (capital expenditure), which in turn reflects the inadequacy of Meta's own model capability — an implicit vote for the side that says general capability is the hard currency.
Slow Thinking
This week, Indigo made one clear revision. His old discipline was an iron rule: on judgment-type, high-leverage tasks, always pay for the strongest general intelligence. Now that rule has been split into two dimensions — the length of the task chain, times the privacy of the data. Tasks with long chains, deep recursion, and public-domain footing still deserve the strongest general model, because errors compound; high-frequency, narrow-domain tasks backed by private data are actually better served by fine-tuned small models — more accurate and cheaper. And strongest
itself has a hard edge: frontier models collectively hit a wall just under the roughly 78% trust threshold, and the highest-order creation simply cannot be written as a reward function. What triggered the revision was the collision of the Bridgewater and Naval first-hand reports in the same week — his own post on compounding errors happened to stand on the general side, and Bridgewater supplied the other half from the specialized side. In practical terms: judging AI application-layer companies upgraded from one dimension — do they use the strongest model — to a two-dimensional matrix; and the moat checkpoint dropped from model selection down to private judgment data.
The other side of it: the week's strongest rebuttal comes precisely from Naval's old call — every generation of general models in history has swallowed the previous generation's specialized optimizations. The 84.7% victory may just be a timing gap, because it rests on the premise that today's frontier is collectively stuck at the 78% ceiling. The evidence that would prove this framework wrong is actually concrete: the next generation of general frontier models crosses the 80% trust threshold on the same narrow-domain tasks, while inference prices keep collapsing, wiping out both the specialized fine-tune's accuracy edge and its 13.8x cost advantage at once. On that day, the moat retreats from private data back to compute and research itself — and what needs correcting would not be a detail, but the whole matrix.
Indigo on X
"It is already market consensus that HBM has entered a growth-type cycle, driven by structural demand shifts and rapid technology iteration; DRAM is gaining new structural, exponential demand from Agentic CPUs … CPU server TAM is surging, and DRAM per CPU core is rising 3–4x … My forecast for the next 12–24 months: HBM's upcycle continues; DRAM spot prices most likely stay strong through 2026, but 2027 brings the 'delivery vs. overshoot' test … NAND rises along with them but with more volatility"
From @indigox, 68 likes
"A 99.9%-correct AI vs. a 90%-correct AI — the gap looks small on a single pass, but run a recursive loop 100 times and the 90% one's accuracy drops to 13%, while the 99.9% one still holds 80-90%. 'Errors in intelligence compound too' — so on judgment-type, high-leverage tasks, always pay for the strongest intelligence"
From @indigox, 34 likes
"The output bottleneck has shifted from 'model capability' to 'whether you can spell out the unknown.' The prompt/context you give is the map, the code/real constraints are the territory, and the gap between them is the 'unknown.' Reducing unknowns, and budgeting for them, is the real craft of agentic coding. Output = f(your ability to clarify the unknown)"
From @indigox, 24 likes
Closing
One Thought
The last time a general capability got commoditized, it was electricity. Once the grid spread, generating power itself quickly stopped being scarce; what opened the real gap were the factories that rearranged their production lines around the electric motor — same price of power, different processes. General models are the grid; specialized fine-tuning is that private production line: the cheaper power gets, the more valuable the private craft of knowing how to use it becomes. This week's Bridgewater-vs-Naval collision is just that century-old script replaying on intelligence — the commoditized part drives the price down, and the uncommoditized part stores the value up.
One Exercise
Take a sheet of paper and write down the ten work tasks that keep coming back to you lately. Score each one on two axes: how long is the chain (done in one step, or dozens of recursive steps), and how private is the data (findable online, or known only to your team). Place the ten tasks on a two-dimensional matrix: for the long-chain, public-domain ones, hand them to the strongest general model you can get, and watch the error rate at every step — errors compound; for the narrow-domain ones sitting on private data, ask yourself one question: can the expert judgment here be written into a checklist and called on repeatedly? Thirty minutes later, you will have your own map of general vs. specialized.