Mind · In / Out · In · 文章

没有人在认真讨论 AI 需求

Nobody is talking seriously about AI demand

Giovanni Cattani · X article · 2026-08-31

AI 投资只算供给不算需求,这篇补上需求这一半:前沿模型约一半收入在自我循环。

Indigo 的结论

前沿模型约一半收入来自自我强化的需求,AI 研发排第一。同一组事实,可以读成上涨的引擎,也可以读成泡沫的裂缝。Giovanni 画出了裂缝在哪,却站在看多一边;他的框架其实更支持他一笔带过的看空理由。

怎么读这篇 作者不卖东西,但立场偏多,他自己说对未来需求「极度乐观」。三个框架可以直接拿来用:token 即时间、有界与无界任务、自我强化的需求。具体比例(AI 研发 50% 以上、软件工程 15%、交易 15%)他自己标了是估计和传闻,当猜测看,不当事实。

需要记住的几件事

  1. 需求这一半一直没人认真建模:之前讨论 AI 资本开支,几乎都在算供给。
  2. token 即时间,加上有界和无界的切法很好用:有上限的任务比便宜,没上限的任务比前沿。
  3. 自我强化的需求只属于前沿模型,领先 6 个月就能通吃;三类依次是 AI 研发、软件工程(主要来自初创和 AI 公司)、交易。
  4. 自我强化、彼此相关、极度顺周期:上涨时是引擎,下跌时是裂缝。多空双方争的不是事实,而是它什么时候反转。

拆解 · 6 步

  1. 01

    需求决定钱能投多少、谁赢、价值落在哪

    分析师默认需求无限,却没人拆需求从哪来。开头引 Dwarkesh 的问题:前沿智能每 GW 预期 $100B 收入,实验室为什么不多投算力? 读这一段原文 →

  2. 02

    token 即时间:用 METR 的图切出两根轴

    每个模型的 token 对应它能独立做完的任务时长:o3 约 30 分钟,Mythos 约 3 小时。再按两根轴切:模型做不做得到,任务有没有上限。 读这一段原文 →

  3. 03

    有界任务比便宜,无界任务比前沿

    有界任务的回报是省成本,非前沿和开源模型主导,自动化一次就够;无界任务的回报是增收,先到者通吃,前沿模型主导。 读这一段原文 →

  4. 04

    一半的前沿收入在自我强化

    AI 研发、软件工程、交易三类约占前沿实验室推理收入的 50%,共用一个循环:多花 token,换来更多收入和融资,再花更多 token。 读这一段原文 →

  5. 05

    自我强化是双刃剑,他却站在看多一边

    三类需求高度相关、又极度顺周期,监管、加息或外部冲击都可能触发向下的连锁收缩;他承认这种需求脆弱,但判断短期不会反转。 读这一段原文 →

  6. 06

    价值流向无界任务的团队,资本配置被重新定价

    十年内,有界任务的价值从人力转到数据中心,很少留在公司手里;市场会给可复现的现金流打折,给善于配置长期投入的团队更高估值。 读这一段原文 →

对 Rewired Index 意味着什么

需求模型是供给模型缺的另一半,直接帮助判断价值落在哪一层、哪类任务。相关标的:SPCX、TSLA(资本配置被重新定价的现成例子,Giovanni 点了名);前沿实验室自我强化的收入结构(未上市,只观察)。

什么会让我改口

监管、加息或外部冲击出现后,前沿模型的 token 需求没有跟着收缩;或者三类自我强化的任务被证实远不到前沿推理收入的 50%。

怎么读这篇

作者不卖东西,但立场偏多,他自己说对未来需求「极度乐观」。三个框架可以直接拿来用:token 即时间、有界与无界任务、自我强化的需求。具体比例(AI 研发 50% 以上、软件工程 15%、交易 15%)他自己标了是估计和传闻,当猜测看,不当事实。

拆解 · 6 步
  1. 需求决定钱能投多少、谁赢、价值落在哪
  2. token 即时间:用 METR 的图切出两根轴
  3. 有界任务比便宜,无界任务比前沿
  4. 一半的前沿收入在自我强化
  5. 自我强化是双刃剑,他却站在看多一边
  6. 价值流向无界任务的团队,资本配置被重新定价
01

需求决定钱能投多少、谁赢、价值落在哪

分析师默认需求无限,却没人拆需求从哪来。开头引 Dwarkesh 的问题:前沿智能每 GW 预期 $100B 收入,实验室为什么不多投算力?

2028 年的 AI 资本开支预计将超过法国的国家预算。前沿 AI 实验室的收入爬坡曲线几乎可以为任何数字辩护。AI 需求中的反身性是一把双刃剑,现在正是开始讨论它的好时机。

在最新一期播客中,@dwarkesh_sp 提出疑问:既然前沿智能的预期收入高达 $100B/GW,前沿实验室为什么不在算力上投入更多?按这个费率,如果 Anthropic 把预期产能全部变现,到 2026 年底它的 ARR 可以达到 $500B——足以成为全球营收第三大的公司。

没有人在认真讨论 AI 需求。今天,分析师们在任何供给水平下都假设 AI 需求是无限的,却并不真正拆解这些需求从何而来。但需求极其重要:它决定了前沿实验室能投入多少资本、哪些模型会被使用,以及价值将在哪里沉淀。

下面是我思考 AI 需求的框架、看多前沿 token 的理由,以及几个相关话题。

我对 AI 的未来需求极度乐观。但我的框架显示,前沿 token 的需求可能带有部分反身性,由 AI 热潮本身推动。这在上行时很棒,会加速增长,但它也可能朝相反方向螺旋。

所有数字均为估算。

02

token 即时间:用 METR 的图切出两根轴

每个模型的 token 对应它能独立做完的任务时长:o3 约 30 分钟,Mythos 约 3 小时。再按两根轴切:模型做不做得到,任务有没有上限。

Token 是时间的单位

METR 的长时程任务图,可能是当今世界上最重要的一张图。

基于这张图,我常说可以把 token 看作时间的单位。每个模型的 token 对应一定的任务时程,而前沿模型的时程最长。图上,o3 约为 30 分钟,而 Mythos 约为 3 小时。使用 o3 的软件工程师可以委派最长 30 分钟的任务,而使用 Mythos 的工程师则被显著提速——可以委派长达 3 小时的任务。

这张图还提供了一个梳理人类活动的框架。我们来定义:

长时程任务 vs. 短时程任务。「长时程」任务是目前还没有任何模型能胜任的任务——即位于那条从 GPT-2 延伸到 Mythos 的曲线之上的部分。曲线之下的其余一切都是「短时程」——实际上已经被 AI 解决了。

有界任务 vs. 无界任务。「有界」任务是指曲线在某个点之后不再继续扩展的任务:报税就是有界任务,它的复杂度存在上限。「无界」任务是指你永远可以做得更多的任务:AI 研究、探索太空、长寿科学。

03

有界任务比便宜,无界任务比前沿

有界任务的回报是省成本,非前沿和开源模型主导,自动化一次就够;无界任务的回报是增收,先到者通吃,前沿模型主导。

劳动分类学

对 AI 的需求归根结底是对劳动的需求。要思考 AI 需求侧的动态,我们需要一套新的劳动分类学。想象一个 2x2 矩阵,把有界—无界与长时程—短时程两组维度组合起来:

有界任务

对这类任务,AI 很快会把所有基准测试打满:你只需相信 METR 的趋势线。今天用任何 AI 模型都能轻松一次性搞定一个简单前端(有界、短时程),而税务筹划(有界、长时程)可能还需要一两年,但我们终会到达。

有界任务也是人们通常希望花时间越少越好的任务(上行空间不大,主要是别出错)。在这里,对 AI 的需求会极其旺盛,但会收敛到最便宜的选项:有界任务的 ROI 主要来自成本节约而非增量收入(对会计服务的需求终归有限)。

因此,非前沿模型将在这里占主导:没必要为实验室的利润率买单,优秀的推理服务商随处可得,对小模型做后训练也很有效。开源模型会继续追赶前沿——而对这类任务来说,这个时间差基本无关紧要,因为你只需要把它们自动化一次。

无界任务

无界任务截然不同:按定义,你永远可以做得更多。探索太空,你总能探索得更远。缩小芯片制程节点,你总能再缩小一点。改进 AI,你永远能把它做得更好。软件永远可以再多写一些,交易策略永远需要更新。

换句话说,无界任务是你希望投入尽可能多时间的任务——这就是为什么前沿 token 在这里基本没有商量余地。这类任务往往一辈子都不够用。由于这些任务极强的凸性、幂律特征,其 ROI 方程主要围绕增量收入而非成本节约。更快严格意义上就是更好,而竞争动态意味着抢先是唯一最重要的事。你想要最好的 AI 模型、品类里最好的软件、alpha 最多的交易策略。

前沿模型将在这里占主导。就像 F1 一样,赛车只能撑 9 个月并不重要——你永远需要开最好的那一辆。

04

一半的前沿收入在自我强化

AI 研发、软件工程、交易三类约占前沿实验室推理收入的 50%,共用一个循环:多花 token,换来更多收入和融资,再花更多 token。

反身性需求

上面描述的框架今天已经在实时上演。

我的看法是:前沿模型史无前例的收入爬坡,是由一小组无界、长时程任务驱动的。我的最佳猜测是:(1) AI 研发,(2) 软件工程,(3) 交易。

按顺序来说:

AI 研发:传闻各实验室如今把 60% 的算力预算花在训练上、只有 40% 用于推理,也可以合理假设全球 50% 的 AI 算力被用于训练。仅此一项就足以让 AI 研发成为毫无疑问的第一大用例。但我还预计推理这部分里也有相当比例本质上属于某种 AI 研发。比如,第二梯队的 AI 实验室用前沿模型做研究和合成数据生成,或是应用型 AI 公司在做后训练。姑且算前沿 AI 实验室推理收入的约 20% 来自 AI 研发。

软件工程:软件工程是 AI 需求最显而易见的任务类别,但微妙之处在于,这里对前沿 token 的需求有很大一块来自初创公司和 AI 公司——而不是传统的 F500 企业。初创公司不受企业官僚体系束缚,能更快地交付更多代码。交付得越多,从客户和投资人那里拿到的钱就越多。此外,初创公司之间为市场份额、人才和 VC 资金激烈竞争——竞争迫使它们使用前沿 token。成熟的 AI 公司(NVIDIA、Amazon 等)也是同理。姑且算这又占前沿 AI 实验室推理收入的约 15%。

交易:如果「量化交易公司是前沿 AI token 最大买家之一」的传闻属实,且「客户在前沿 AI token 上的支出呈幂律分布」的传闻也属实,那么交易占到前沿 token 收入两位数百分比,我不会感到意外。其他投资机构同样如此。姑且也算这占前沿 AI 实验室推理收入的 15%。

如果这个归因大致正确,那么可以说仅这三类任务就贡献了前沿 AI 实验室约 50% 的推理收入。

特别有意思的是,这三类任务在 token 支出与收入之间都存在一个紧密的反馈回路,使得收入增长具有反身性:

AI 研发:AI 实验室持续把研发投入转化为更强的模型能力,能力提升几乎即刻带来收入增长和融资能力增强,反过来又让实验室能在研发上投入更多,如此往复;

软件工程:一家用 AI 的初创公司可以快速构建大量软件,借此扩大收入并完成股权融资,再迅速把这些资源投向更多的 token,如此往复;

交易:交易公司可以在市场上即时检验 token 支出的效果,利润越高,就越能把资金再投入到 AI 驱动的策略中,如此往复。

我把这些任务的需求描述为反身性的,因为它们共享同一个模式:更大的 token 支出带来更多收入和更强的融资能力,进而提供更多可投入 token 的资源,如此循环——看起来没有上限。

对这些任务而言,唯一的约束是你能往这个问题上分配多少 token。最好的公司是那些能募到足够资本并把它配置到正确赌注上的公司——正如 Anthropic 率先把重心收窄到编程所做的那样。而这里妙的地方在于,总可及市场大致是无限的——token 支出会拓展我们所能成就之事的边界,ROI 由更高的收入驱动。

而且对这些任务来说,只有前沿模型才重要。这正是前沿模型在今天成为如此优秀生意的原因。反身性需求只属于前沿模型,哪怕只领先开源六个月,也足以吃下整个市场。

相比之下,有界任务正在被成百上千的初创公司蚕食。这很好,也不可避免。然而这类收入没有反身性。你或许能削减成本,但更低的成本与更高的收入之间没有即时反馈回路。而且节约下来的成本还得分一部分给 RLaaS 服务商,降本本身也存在天花板。

关于开源与闭源之争,或者消费级与企业级之辩,其中的含义就留给读者自己体会了。

05

自我强化是双刃剑,他却站在看多一边

三类需求高度相关、又极度顺周期,监管、加息或外部冲击都可能触发向下的连锁收缩;他承认这种需求脆弱,但判断短期不会反转。

反身性是双刃剑

前沿 token 的反身性需求是一把双刃剑。反身性在上行时妙不可言,在下行时则糟糕透顶。

其中涉及的数字如此之大,以至于必须考虑这样一种情形:至少在某段时间内,前沿 token 的需求出现收缩。届时的问题在于,关键的需求驱动因素彼此相关,且极度顺周期。

相关性

三类关键任务中任何一类的需求收缩,都可能直接冲击前沿 token 收入的 15–20%,而由于相关性,冲击面可能高达 50%。

眼下的循环是:(i) 前沿 AI 收入走高推升实验室的股权价值,(ii) 这鼓励更多资金流入 VC,其中很大一部分被花在 token 上,(iii) 这又抬高了实验室和初创公司的估值,进而 (iv) 推动公开市场上涨,量化公司从中获利,并 (v) 加大在前沿 token 上的支出。这只是一个例子,但传染可以从这个回路的任何一环开始。

极度顺周期

这些任务极度顺周期:周期上行时其需求会加速,而周期下行时可能朝反方向加速。

市场状态直接影响三大关键驱动因素的需求。市场上涨让量化交易公司能在 token 上花更多钱,也给了 AI 实验室和初创公司更多可投的资本。

在这个框架下,也许只需要一个简单的触发因素,就能启动一场向下的反身性修正。信手举几个例子:监管拖慢 AI 能力的进展;利率上升拖慢基建扩张;任何外生冲击。

一个 AI 需求模型

显然,我们需要一个前沿 token 需求的模型。

极度顺周期、高度相关的需求是脆弱的。虽然一路上有些颠簸,但第一版 ChatGPT 的发布几乎与 NASDAQ 最近一次相对底部重合,此后一级和二级市场都一路向右上方走。我们甚至还没探讨过 AI 实验室收入若放缓可能如何冲击市场。短期内这不太可能发生(下一代模型或许还会成为加速的催化剂),但恰恰因为如此,现在正是思考这个话题的好时机。

考虑到当前每年 AI 资本开支和收入的量级,一个 AI 需求模型将与我们现有的供给模型互为补充,并为关键的投资决策——从融资到投资——提供依据,尤其是对暴露于这轮基建扩张的公司。上面的框架或许可以作为一个起点。

06

价值流向无界任务的团队,资本配置被重新定价

十年内,有界任务的价值从人力转到数据中心,很少留在公司手里;市场会给可复现的现金流打折,给善于配置长期投入的团队更高估值。

超越基建的价值

需求模型还应指导基建之外的资本配置。价值会继续沉淀到物理 AI 供应链。但一个较少被讨论的方向是:价值也会沉淀到那些追逐无界、长时程任务的团队身上。

在大约十年内,今天由有界任务创造的价值中,将有一大块从人类劳动手中被拿走、转移到数据中心。从财务上看,可以把这些任务对应的劳动 GDP 打个折扣,再把它挪到 AI 供应链上。非前沿模型将在量上占主导,人也许仍能赚钱(做销售、设计以及一些长时程规划),但沉淀到公司层面的价值会很少。

我预计大部分价值将沉淀到无界任务上。尤其是那些能让世界相信自己有能力把资本开支(即 token)配置到这些长时程、无界任务上的团队。这是模型能力之外的领域:你永远可以用更长的时间视野去思考,我们将看到创业者去做那些以今天的方式需要好几辈子才能建成的公司。其中一部分可能是 AI 实验室自己,一部分将是新公司。

今天,市场奖励经常性、可预测的现金流——并厌恶看不到短期实际成果的研发与资本开支。未来我们可能看到相反的景象:市场将重度折价有界任务带来的可重复现金流,同时重估那些能为长期明智配置资本开支的团队。

Elon Musk 不是异类——他是这一模式的第一个样本。Tesla 和 SpaceX 的交易价格是老派金融分析师给其现金流定价的 10X。但市场一直在为 Elon 把研发投入配置到任何理性人都会认为不可能之事上的能力定价。有了超级智能,还会出现更多个 Elon。

这里还有更多可说的——但那是另一天的故事了。

判断收口延伸

Indigo 的结论

前沿模型约一半收入来自自我强化的需求,AI 研发排第一。同一组事实,可以读成上涨的引擎,也可以读成泡沫的裂缝。Giovanni 画出了裂缝在哪,却站在看多一边;他的框架其实更支持他一笔带过的看空理由。

需要记住的几件事

  1. 需求这一半一直没人认真建模:之前讨论 AI 资本开支,几乎都在算供给。
  2. token 即时间,加上有界和无界的切法很好用:有上限的任务比便宜,没上限的任务比前沿。
  3. 自我强化的需求只属于前沿模型,领先 6 个月就能通吃;三类依次是 AI 研发、软件工程(主要来自初创和 AI 公司)、交易。
  4. 自我强化、彼此相关、极度顺周期:上涨时是引擎,下跌时是裂缝。多空双方争的不是事实,而是它什么时候反转。

可回查的判断

判断谁说的何时见分晓证据多硬
三类自我强化的任务(AI 研发、软件工程、交易)约占前沿实验室推理收入的 50%Giovanni现在一手估计,他自己标了「猜测、传闻」
有界任务由非前沿和开源模型主导,无界任务由前沿模型主导Giovanni进行中一手,框架判断
前沿 token 需求极度顺周期、彼此高度相关,一次外部冲击就可能触发向下的连锁收缩Giovanni无期一手,机制判断
十年内有界任务的价值从人力转到数据中心,很少留在公司手里Giovanni约 2036一手,方向判断
市场会给可复现的现金流打折,给善于配置长期投入的团队更高估值(会出现更多 Elon)Giovanni无期一手,愿景

放回主线

补充

价值上移到不可租用的东西 Giovanni 说价值会流向做无界任务、敢为长期投入的团队,这是这条判断在资本市场上的版本。

补充

可验证域能否泛化 有界与无界、能验证与难验证,两把尺子一起用,能更细地判断价值落在哪。

冲突+补充

NeilMovva:把 token 做到过剩,让智能便宜 1000 倍 两者不矛盾:Neil 说最便宜的 token 赢,Giovanni 补上适用范围,只对有上限的任务成立。

补充

fin AI 半导体终局 III:把泡沫定位到时间错配 fin 从供给侧算账,Giovanni 补上需求侧;两边都指向同一个真风险:投入和回报在时间上错开。

补充

a16z Machine Age Fund:需求真实,但绕开了融资脆弱 a16z 说需求是真的;Giovanni 说清了真在哪:三类没有上限的长任务。

对 Rewired Index 意味着什么

需求模型是供给模型缺的另一半,直接帮助判断价值落在哪一层、哪类任务。相关标的:SPCX、TSLA(资本配置被重新定价的现成例子,Giovanni 点了名);前沿实验室自我强化的收入结构(未上市,只观察)。

什么会让我改口

监管、加息或外部冲击出现后,前沿模型的 token 需求没有跟着收缩;或者三类自我强化的任务被证实远不到前沿推理收入的 50%。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Essay

Nobody is talking seriously about AI demand

Giovanni Cattani · X article · 2026-08-31

Everyone models AI supply and nobody models demand. This fills in the demand half: about half of frontier revenue feeds itself.

Indigo's conclusion

About half of frontier revenue comes from self-reinforcing demand, AI R&D first of all. The same facts read as an engine on the way up and a crack on the way down. Giovanni shows where it would crack, then sides with the bulls; his own framework supports the bear case he waves away.

How to read this The author sells nothing, but he leans bullish; by his own account he is “extremely optimistic” about future demand. The three frameworks are worth keeping: tokens as time, bounded vs. unbounded tasks, self-reinforcing demand. The shares (AI R&D 50%+, software engineering 15%, trading 15%) he labels himself as estimates and rumors. Treat them as guesses, not facts.

What to remember

  1. Nobody had seriously modeled the demand half; the debate over AI spending has been almost all supply.
  2. Tokens as time, plus bounded vs. unbounded, is a useful cut: tasks with a ceiling compete on price, tasks without one on the frontier.
  3. Self-reinforcing demand belongs only to the frontier, where a 6-month lead takes the market. The three sources: AI R&D, software engineering (mostly startups and AI companies), trading.
  4. Self-reinforcing, correlated and swinging with the cycle: an engine going up, a crack going down. Bulls and bears agree on the facts and disagree on when it turns.

Breakdown · 6 steps

  1. 01

    Demand decides how much gets spent, who wins, and where value lands

    Analysts assume unlimited demand and never ask where it comes from. The hook is Dwarkesh's question: if frontier intelligence is expected to earn $100B per GW, why don't labs buy more compute? Read this part →

  2. 02

    Tokens are time: two axes from the METR chart

    Each model's tokens stand for how long a task it can finish on its own: o3 about 30 minutes, Mythos about 3 hours. Then two axes: can a model do it yet, and does the task have a ceiling. Read this part →

  3. 03

    Bounded tasks go to the cheapest, unbounded to the frontier

    Bounded tasks pay off in saved cost, so non-frontier and open models win and you automate once. Unbounded tasks pay off in revenue, first to arrive takes all, so the frontier wins. Read this part →

  4. 04

    Half of frontier revenue feeds itself

    AI R&D, software engineering and trading make up about 50% of frontier labs' inference revenue. All three run the same loop: spend more on tokens, earn and raise more, spend more on tokens. Read this part →

  5. 05

    It cuts both ways, and he still sides with the bulls

    The three sources of demand move together and swing hard with the cycle. Regulation, rate hikes or an outside shock could set off a downward spiral. He admits the demand is fragile but expects no reversal soon. Read this part →

  6. 06

    Value goes to teams on unbounded work; capital allocation gets repriced

    Within a decade the value of bounded tasks moves from human labor to data centers, and little of it stays with companies. Markets will discount repeatable cash flows and pay up for teams that invest well for the long run. Read this part →

What it means for Rewired Index

A demand model is the missing half of the supply model, and it helps judge which layer and which kind of task value lands in. Related names: SPCX and TSLA (live cases of capital allocation being repriced; Giovanni names both); the self-reinforcing revenue mix of frontier labs (private; watching only).

What would change my mind

regulation, rate hikes or an outside shock arrives and frontier token demand does not shrink; or the three self-reinforcing categories turn out to be far below 50% of frontier inference revenue.

How to read this

The author sells nothing, but he leans bullish; by his own account he is “extremely optimistic” about future demand. The three frameworks are worth keeping: tokens as time, bounded vs. unbounded tasks, self-reinforcing demand. The shares (AI R&D 50%+, software engineering 15%, trading 15%) he labels himself as estimates and rumors. Treat them as guesses, not facts.

Breakdown · 6 steps
  1. Demand decides how much gets spent, who wins, and where value lands
  2. Tokens are time: two axes from the METR chart
  3. Bounded tasks go to the cheapest, unbounded to the frontier
  4. Half of frontier revenue feeds itself
  5. It cuts both ways, and he still sides with the bulls
  6. Value goes to teams on unbounded work; capital allocation gets repriced
01

Demand decides how much gets spent, who wins, and where value lands

Analysts assume unlimited demand and never ask where it comes from. The hook is Dwarkesh's question: if frontier intelligence is expected to earn $100B per GW, why don't labs buy more compute?

AI capex for 2028 is forecast to be larger than the budget of France. Frontier AI labs’ revenue ramp justifies almost any number. Reflexivity in AI demand is a double-edged sword, and it's now a good time to start talking about it.

On his latest podcast, @dwarkesh_sp wonders why the frontier labs are not spending even more on compute, given expectations of $100B/GW in revenue for frontier intelligence. At these rates, if Anthropic were to monetize all of its expected capacity, it could be at $500B ARR by the end of 2026 - enough to be the third-largest company in the world by revenue.

Nobody is talking seriously about AI demand. Today, analysts model infinite demand for AI for any level of supply, without really breaking down where that demand is coming from. But demand is extremely important, it dictates how much capital frontier labs can invest, which models will be used, and where the value will accrue.

Below is my framework for thinking about AI demand, the case for frontier tokens, and a few adjacent topics.

I am extremely optimistic about future demand for AI. But my framework suggests that demand for frontier tokens may be partly reflexive, fueled by the AI boom itself. This is great and accelerates growth, but it may also spiral in the opposite direction.

All numbers are estimates.

02

Tokens are time: two axes from the METR chart

Each model's tokens stand for how long a task it can finish on its own: o3 about 30 minutes, Mythos about 3 hours. Then two axes: can a model do it yet, and does the task have a ceiling.

Tokens as units of time

METR’s long-horizon chart is possibly the most important chart in the world today.

Based on the chart, I often say we can think about tokens as units of time. Tokens from each model represent a certain task-horizon, and frontier models have the longest duration. On the chart, o3 is ~ 30 min, while Mythos is ~ 3h. A software engineer using o3 can delegate tasks of up to 30 min, while one using Mythos is significantly sped up - delegating up to 3h.

The chart also provides a framework for structuring human activities. Let’s define:

Long-horizon vs. short-horizon tasks. “Long-horizon” tasks are the ones that no model succeeds at yet - i.e., above the curve which runs from GPT-2 to Mythos. Everything else, below the curve, is “short-horizon” - effectively, it’s already been solved by AI.

Bounded vs. unbounded tasks. “Bounded” tasks are those for which, at some point, the curve stops scaling: doing taxes is a bounded task, there’s a limit to its complexity. “Unbounded” tasks are those for which you can always do more: AI research, exploring space, longevity.

03

Bounded tasks go to the cheapest, unbounded to the frontier

Bounded tasks pay off in saved cost, so non-frontier and open models win and you automate once. Unbounded tasks pay off in revenue, first to arrive takes all, so the frontier wins.

Labor taxonomy

Demand for AI is ultimately demand for labor. And to think about AI demand-side dynamics, we need a new labor taxonomy. Imagine a 2-by-2 matrix, combining the bounded-unbounded categories with long-horizon and short-horizon ones:

Bounded Tasks

For these tasks, AI will soon saturate all benchmarks: one can just trust METR’s trendline. You can easily one-shot a simple frontend (bounded, short-horizon) with any AI model today, while tax planning (bounded, long-horizon) may take another year or two, but we’ll get there.

Bounded tasks are also those for which one would generally like to spend the least amount of time possible (little upside, mostly just a matter of not making mistakes). Here, demand for AI will be extremely high, but converge on the cheapest possible option: the ROI for bounded tasks is mostly a function of cost savings rather than additional revenue (there’s only so much demand for accounting).

Therefore, non-frontier models will dominate: no need to pay for the labs’ margins, great inference providers are available, and post-training of smaller models is effective. And open source models will continue to catch up with the frontier - with a lag that is mostly irrelevant for these tasks, since you just have to automate them once.

Unbounded Tasks

Unbounded tasks are very different: by definition, you can always do more. Exploring space, you can always explore further. Shrinking the node on a chip, you can always shrink it more. Improving AI, you will always be able to make it better. And you can always build more software, and you always need to update trading strategies.

In other words, unbounded tasks are those for which you’d want to spend as much time as possible - that’s why frontier tokens are mostly non-negotiable. These are the kinds of tasks for which one life is often not enough. Due to the very convex, power-law nature of these tasks, the ROI equation is mostly focused on additional revenue rather than cost savings. Going faster is strictly better, and competitive dynamics mean that being first is the single most important thing. You want the best AI model, the best software in the category, the trading strategy with the most alpha.

Frontier models will dominate here. Just like in F1, it doesn't matter if the car only lasts 9 months - you always need to drive the best possible one.

04

Half of frontier revenue feeds itself

AI R&D, software engineering and trading make up about 50% of frontier labs' inference revenue. All three run the same loop: spend more on tokens, earn and raise more, spend more on tokens.

Reflexive demand

The framework described above is already playing out today, in real time.

My view is that the unprecedented revenue ramp for frontier models is driven by a small set of unbounded, long-horizon tasks. My best guess is: (1) AI R&D, (2) software engineering, and (3) trading.

In that order:

AI R&D: Rumor has it that labs today spend 60% of their compute budget on training and only 40% on inference, and it’s fair to assume that 50% of the world’s AI compute capacity is used for training. That would already make AI R&D the clear #1 use case. But I’d also expect a significant portion of the inference bucket to be some form of AI R&D. For instance, runner-up AI labs using frontier models for research and synthetic data generation, or the applied AI companies doing post-training. Let’s say ~ 20% of inference revenue for the frontier AI labs is coming from AI R&D.

Software engineering: Software engineering is the obvious task category for AI demand, but the nuance is that a large chunk of the demand for frontier tokens here is coming from startups and AI companies - not traditional F500 companies. Startups are unconstrained by corporate bureaucracy, and can just ship more code, faster. The more they ship, the more money from clients and investors. Additionally, startups compete fiercely with each other for market share, talent, and VC money - and competition forces them to use frontier tokens. The same goes for the mature AI companies (NVIDIA, Amazon, etc.). Let’s say this is another ~ 15% of inference revenue for the frontier AI labs.

Trading: If rumors that quant trading firms are some of the largest spenders on frontier AI tokens are true, and if rumors that spend on frontier AI tokens by customers is power-law distributed are true, then it wouldn’t surprise me to see trading as a double-digit percentage of frontier token revenue. This holds for other investment firms as well. Let’s say this too is 15% of inference revenue for the frontier AI labs.

If this attribution is roughly right, then we could say that these three categories of tasks alone account for ~ 50% of frontier AI lab inference revenue.

What is especially interesting is that these three sets share a tight feedback loop between token spend and revenue, such that revenue growth is reflexive:

AI R&D: AI labs have consistently translated AI R&D spend into stronger model capabilities, which increase revenue and the ability to raise capital almost instantaneously, and in turn allow the labs to spend more on R&D, and so on;

Software engineering: A startup using AI can build a lot of software fast, use that to scale revenue and raise equity, and quickly deploy those resources on even more tokens, and so on;

Trading: A trading firm can test the impact of token spend instantaneously on the market, and the higher the profits, the more it can reinvest in AI-powered strategies, and so on.

I describe demand for these tasks as reflexive because they all share the same pattern: larger spend on tokens yields more revenue and a stronger ability to raise capital, which in turn gives them more resources to invest in tokens, and so on - seemingly with no upper bound.

For these tasks, the only constraint is how many tokens you can allocate to the problem. The best companies are the ones that can raise enough capital and allocate it to the right bets - as Anthropic did by being the first to narrow its focus to coding. And the cool thing here is that the total addressable market is roughly infinite - token spend expands the horizon of what we can accomplish, and ROI is driven by higher revenue.

And for these tasks, only frontier models matter. This is what makes frontier models such a great business today. Reflexive demand is only for frontier models, even a six-month lead over open source is more than enough to capture the whole market.

By contrast, bounded tasks are being addressed by dozens and dozens of startups. This is great and inevitable. However, there is no reflexivity for this kind of revenue. You can perhaps cut costs, but there is no immediate feedback loop between lower costs and increased revenue. And a percentage of the cost savings has to be shared with the RLaaS provider, and there’s a ceiling on cost cuts.

On the debate between open and closed source, or on consumer versus enterprise, I will leave the implications to the reader.

05

It cuts both ways, and he still sides with the bulls

The three sources of demand move together and swing hard with the cycle. Regulation, rate hikes or an outside shock could set off a downward spiral. He admits the demand is fragile but expects no reversal soon.

Reflexivity is double-edged

Reflexive demand for frontier tokens is a double-edged sword. Reflexivity is great on the way up, awful on the way down.

The numbers at stake are so large that one has to consider the scenario where, at least temporarily, demand for frontier tokens contracts. In such a case, the issue is that the key demand drivers are correlated and super procyclical.

Correlation

Demand contraction for any of the three key tasks could hit up to 15–20% of frontier token revenue directly and, due to correlation, up to 50%.

Right now: (i) higher revenue for frontier AI drives the equity value of the labs up, which (ii) encourages more investment in VC, a large share of which is spent on tokens, which (iii) increases the value of both labs and startups, and in turn (iv) makes the public markets go up, with the quant firms profiting from it and (v) increasing their spend on frontier tokens. This is just one example, but contagion could start from any of the steps in this loop.

Super procyclicality

These tasks are super procyclical because their demand accelerates as the cycle goes up, and may accelerate in the opposite direction when the cycle goes down.

The state of the market directly impacts demand for all three key drivers. Market going up allows quant trading firms to spend more on tokens, but also grants AI labs and startups more capital to invest.

Under this framework, you may only need one simple trigger to start a reflexive correction downwards. As examples, off the top of my head: regulation slowing down progress in AI capabilities; higher interest rates slowing down the buildout; any exogenous shock.

A model for AI demand

Obviously, we need a model for frontier token demand.

Super procyclical, heavily correlated demand is fragile. While there have been some bumps along the way, the first ChatGPT release almost coincided with the most recent NASDAQ relative bottom, and both private and public markets have gone up and to the right. We haven’t even explored how a potential slowdown in revenue for the AI labs could impact the markets. It’s unlikely that this will happen anytime soon (next-generation models may be a catalyst for acceleration), but precisely for this reason it is now a great time to think about the topic.

Given current levels of annual AI capex and revenue, a model for AI demand would complement the one we currently have for supply, and inform critical investment decisions - from financing to investing - especially for the companies exposed to the buildout. The framework above may be a starting point.

06

Value goes to teams on unbounded work; capital allocation gets repriced

Within a decade the value of bounded tasks moves from human labor to data centers, and little of it stays with companies. Markets will discount repeatable cash flows and pay up for teams that invest well for the long run.

Value beyond

A model for demand should also inform capital allocation beyond the buildout. Value will keep accruing to the physical AI supply chain. But while less discussed, value will also accrue to teams going after unbounded, long-horizon tasks.

Within a decade or so, a large chunk of the value generated by bounded tasks today will be taken from human labor and moved to data centers. Financially, one could take labor GDP for those tasks, apply a percentage cut, and move it to the AI supply chain. Non-frontier models will dominate volumes, people may still make money (doing sales, design, and some long-horizon planning), but little value will accrue to the company.

I expect most of the value to accrue to unbounded tasks. And in particular, to teams that can convince the world of their ability to allocate capex (i.e., tokens) to go after these long-horizon, unbounded tasks. This is the domain outside of model capabilities: you can always think with a longer horizon, and we will see founders going after companies that would take several lifetimes to build today. Some of it may be the AI labs themselves, some of it will be new companies.

Today, the market rewards recurring, predictable cash flows - and hates R&D and capex spent with no short-term tangible results in sight. In the future, we may see the inverse: the market will heavily discount repeatable cash flow from bounded tasks, while repricing teams that can wisely allocate capex for the long term.

Elon Musk is not an anomaly - he’s the first example of this. Tesla and SpaceX trade at 10X what an old-fashioned financial analyst would price their cash flows at. But the market routinely prices Elon’s ability to allocate R&D spend to what any reasonable person would consider impossible. With superintelligence, there will be several more Elons.

More to say here - but this is a story for another day.

Where Indigo landsFurther

Indigo's conclusion

About half of frontier revenue comes from self-reinforcing demand, AI R&D first of all. The same facts read as an engine on the way up and a crack on the way down. Giovanni shows where it would crack, then sides with the bulls; his own framework supports the bear case he waves away.

What to remember

  1. Nobody had seriously modeled the demand half; the debate over AI spending has been almost all supply.
  2. Tokens as time, plus bounded vs. unbounded, is a useful cut: tasks with a ceiling compete on price, tasks without one on the frontier.
  3. Self-reinforcing demand belongs only to the frontier, where a 6-month lead takes the market. The three sources: AI R&D, software engineering (mostly startups and AI companies), trading.
  4. Self-reinforcing, correlated and swinging with the cycle: an engine going up, a crack going down. Bulls and bears agree on the facts and disagree on when it turns.

Claims you can check later

ClaimWhoWhen we will knowHow firm
Three self-reinforcing categories (AI R&D, software engineering, trading) make up about 50% of frontier labs' inference revenueGiovanniNowFirst-hand estimate he labels guess and rumor
Bounded tasks go to non-frontier and open models; unbounded tasks go to the frontierGiovanniUnder wayFirst-hand; a framework call
Frontier token demand swings hard with the cycle and moves as one; a single outside shock could set off a downward spiralGiovanniNo dateFirst-hand; a call on mechanism
Within a decade the value of bounded tasks moves from human labor to data centers, and little of it stays with companiesGiovanniAround 2036First-hand; a call on direction
Markets will discount repeatable cash flows and pay up for teams that invest well for the long run (more Elons)GiovanniNo dateFirst-hand; a vision

Back on the long-running theses

adds to

Value moves up to what can't be rented Giovanni says value flows to teams on unbounded work that dare to invest for the long run: the capital-markets version of this view.

adds to

Whether verifiable domains generalize Use both rulers, bounded vs. unbounded and easy vs. hard to verify, to judge more finely where value lands.

conflicts + adds to

Neil Movva: make tokens abundant, make intelligence 1000x cheaper No contradiction: Neil says the cheapest token wins; Giovanni adds that this holds only for tasks with a ceiling.

adds to

fin, AI semiconductor endgame III: the bubble is a timing mismatch fin does the supply-side math, Giovanni the demand side. Both point to the same real risk: spending and returns arriving at different times.

adds to

a16z Machine Age Fund: demand is real, financing fragility skipped a16z says demand is real; Giovanni says where: three kinds of long tasks with no ceiling.

What it means for Rewired Index

A demand model is the missing half of the supply model, and it helps judge which layer and which kind of task value lands in. Related names: SPCX and TSLA (live cases of capital allocation being repriced; Giovanni names both); the self-reinforcing revenue mix of frontier labs (private; watching only).

What would change my mind

regulation, rate hikes or an outside shock arrives and frontier token demand does not shrink; or the three self-reinforcing categories turn out to be far below 50% of frontier inference revenue.

Finished. Indigo's take on this piece is in two places: