Mind · In / Out · In · 视频

硅谷淘金热:AI 如何推动新一代芯片

The Silicon Gold Rush: How AI Is Driving the Development of New Chips

Bill Dally、Norm Jouppi · YouTube · 2026-09-10

两位最该知道的人说:能源是第一瓶颈,DRAM 已经见顶,CUDA 的护城河正被 AI 填平。

第 1 段 / 共 7 段 · 2:10
这次淘金热是真需求

真实经济需求、应用简单到只需加速几个基本操作、没有老代码包袱,和 90 年代的超算热潮不同;两家都在给内存厂付大钱。

拆解 · 7 步

  1. 01

    2:10 – 10:35

    这次淘金热是真需求

    真实经济需求、应用简单到只需加速几个基本操作、没有老代码包袱,和 90 年代的超算热潮不同;两家都在给内存厂付大钱。 读这一段视频稿 →

  2. 02

    10:35 – 20:51

    GPU 和 TPU 没有趋同,系统才是护城河

    NVIDIA 靠数值格式和稀疏性领先,却背着「多客户税」;TPU 只服务 Google,从白纸起步做超算;万卡即开即用的系统经验,创业公司没有。 读这一段视频稿 →

  3. 03

    20:51 – 29:31

    软件护城河被填平,精度快压到头

    有规格说明,让 Claude Code 重建调优过的库不难;每年 2 倍的提升只有 3 倍来自制程,精度还剩一两轮,好点子还能撑八到十年。 读这一段视频稿 →

  4. 04

    29:31 – 38:47

    芯片要三年,模型三个月一变

    只能瞄在鸭子前面,把基本运算做好;混合专家对互连延迟要求更高;InferenceMAX 取代 MLPerf,成了大家真正看的基准。 读这一段视频稿 →

  5. 05

    38:47 – 49:35

    能源是第一瓶颈

    输电线难批,转向现场发电;燃气轮机约五年内卖光,锂电太贵,热电池有希望;按平准化成本,天然气仍难被核能打败。 读这一段视频稿 →

  6. 06

    49:35 – 51:56

    DRAM 微缩停了,需求没停

    内存厂同样的零件能卖好几倍价钱,扩产动力反而变小;它们过去几十年总是过度建厂、价格崩盘,这次会不会不同,没有答案。 读这一段视频稿 →

  7. 07

    51:56 – 1:13:03

    量子没用、烧进芯片太险,AI 会设计芯片

    量子计算是大计算小数据,AI 正相反;前沿模型每月一版,把权重烧进 ROM 风险太大;模拟计算会输;验证占了 75% 的人力。 读这一段视频稿 →

Indigo 的结论

两位架构师一手确认:能源排第一,DRAM 超级周期是结构性的而不是周期性的;Dally 自己承认 CUDA 护城河在变浅、Google TPU 有结构上的单一客户优势。在泡沫之争上,他们给出了「这次是真需求」最硬的卖方论证。

怎么读这篇 两位顶级芯片架构师的公开对谈。两家都在卖铲子,「淘金热是真的、我们稳赚」天然对自己有利。能源、内存、护城河、基准这些技术判断当硬料收,「淘金热是真的」当卖铲人的乐观打折。公开场合,没有对抗性追问,但两人有几处罕见的坦白。

需要记住的几件事

  1. 能源是第一瓶颈:输电难批就现场发电,燃气轮机约五年卖光,热电池有望,核能在平准化成本上仍难赢天然气。
  2. DRAM 微缩停了、需求没停,内存超级周期是结构性的;但内存厂会不会长期维持高价,Jouppi 留了问题没答。
  3. TPU 的单一客户优势对 NVIDIA 的多客户税,Dally 说「有点羡慕」:这是商业模式的差别,不是技术强弱。
  4. SemiAnalysis 的 InferenceMAX 取代 MLPerf,成了胜出的基准:它衡量的正是每瓦多少 token。

什么会让我改口

内存厂重回「过度建厂、价格崩盘」的老路,DRAM 结构性供给受限的判断就不成立了。

怎么读这篇

两位顶级芯片架构师的公开对谈。两家都在卖铲子,「淘金热是真的、我们稳赚」天然对自己有利。能源、内存、护城河、基准这些技术判断当硬料收,「淘金热是真的」当卖铲人的乐观打折。公开场合,没有对抗性追问,但两人有几处罕见的坦白。

拆解 · 7 步
  1. 这次淘金热是真需求
  2. GPU 和 TPU 没有趋同,系统才是护城河
  3. 软件护城河被填平,精度快压到头
  4. 芯片要三年,模型三个月一变
  5. 能源是第一瓶颈
  6. DRAM 微缩停了,需求没停
  7. 量子没用、烧进芯片太险,AI 会设计芯片

据视频字幕整理,按说话人分段。

01

这次淘金热是真需求

真实经济需求、应用简单到只需加速几个基本操作、没有老代码包袱,和 90 年代的超算热潮不同;两家都在给内存厂付大钱。

01:24 · 两位生涯平行的架构师

2:10Dave Patterson: 我来介绍两位从 1980 年代就认识的朋友。那时我在伯克利当助理教授,Bill 在加州理工读博,Norm 在斯坦福。Bill Dally 从加州理工开始就一直在造联网的计算机:先在 MIT,后回到斯坦福当上计算机系主任,2009 年起任 NVIDIA 首席科学家,现在也是高级副总裁。他最有名的贡献包括虫洞路由,以及和 Brian Towles 合写的互连网络那本书;他在斯坦福的流处理项目,是现代 GPU 计算的重要前身。Norm 在斯坦福读研时参与了 MIPS 项目,毕业后在 DEC 西部研究实验室待了大约十年,之后去了惠普实验室。2013 年,他的朋友 Jeff Dean 请他到 Google 为深度学习造硬件。Norm 见过太多 AI 的炒作,很怀疑,但 Jeff 说服了他:深度学习用在什么上都管用。他现在是 Google 的院士兼副总裁,以 TPU 和受害者缓存、预取缓冲这类存储层级的工作闻名。两人的生涯惊人地平行:都在斯坦福师从 John Hennessy,都是美国国家工程院院士,都拿过计算机体系结构的最高奖 Eckert-Mauchly 奖。

06:10 · 这次淘金热有什么不同

6:35Dave Patterson: 今晚的题目是「硅谷淘金热」:密集的创新、投资和竞争。站在 NVIDIA 和 Google 的位置,你们觉得这个 AI 硬件时代的特征是什么?和以往的计算机体系结构时代有什么根本不同?

6:56Bill Dally: 有三个特征让它非常不同。第一,强烈的经济需求。就像 Jeff Dean 说的,AI 用在什么上都管用,于是对更多 token、更多算力、更多 AI 的需求永远填不满。第二,应用变化很快,但和那些动辄几百万行代码的大型超算应用相比,它相对简单。Transformer 是个相对简单的东西,很容易看出要造什么才能让它跑快。第三,大家愿意很快地改,没有积满灰尘的老代码。对比 90 年代初,那时围绕超算也有一轮体系结构淘金热。但没有强烈的经济需求,它是靠 DARPA 的战略计算计划撑起来的,这笔钱让人以为下游有个大市场,其实没有。很多人创了业,比如 Thinking Machines,大多数最后都倒闭了。那时的应用是积满灰尘的老代码,有的接近一百万行。你可以加速看起来最重要的那个内核,但 Amdahl 定律会反咬你,因为另外 99% 的代码没被加速。所以,一个简单的应用、加速几个基本操作就能得到很好的结果,再加上真实的经济需求,让这一轮站得住。这不是 80 年代的 Lisp 机热潮,也不是 90 年代的超算热潮,它在创造真实的价值。

8:49Norm Jouppi: 我一直关注全球经济,现在投进来的钱多得惊人。纽约有一段百年老隧道通往新泽西,一直漏水,差十亿美元修不了;而 Google 宣布明年资本开支 $1050 亿,其他公司也投类似的数目。会有赢家也会有输家,每个人都想跑得越快越好。

9:34Dave Patterson: 放到真正的淘金热里,谁是淘金的人,谁是卖镐头的人?

9:42Bill Dally: 淘金的是那些找到垂直领域的人。找到适合用 AI 的垂直领域,就能像当年的淘金者一样发大财。NVIDIA 和 Google 这样做基础设施的公司,是当年的 Leland Stanford,卖镐头和铲子给淘金的人。他们发不发财,我们都会赚钱。

10:12Norm Jouppi: 不过,我们俩都还在给内存公司付大把的钱。

10:19Bill Dally: 我记得 Micron 上个季度赚的钱,和它之前 19 年加起来一样多。有意思的是,它们的营收在涨,卖出去的零件数量却持平。

10:31Norm Jouppi: 没错,利润率高到天上去了。

02

GPU 和 TPU 没有趋同,系统才是护城河

NVIDIA 靠数值格式和稀疏性领先,却背着「多客户税」;TPU 只服务 Google,从白纸起步做超算;万卡即开即用的系统经验,创业公司没有。

10:29 · 殊途同归?GPU 和 TPU

10:35Dave Patterson: 我有位同事说训练加速器是「趋同演化」:一块尽可能大的计算芯片,一个大的脉动阵列矩阵单元,周边塞满 HBM,再用最快的 SerDes 做定制互连。他错在哪?GPU 和 TPU 走的路是不是其实不同?你们各自欣赏对方设计的哪一点?

11:22Norm Jouppi: 别忘了,寒武纪大爆发之后还有大灭绝。那时有各种奇奇怪怪的生物,像长着怪异附肢的巨型鼠妇,都没挺过灭绝事件。现在也一样:很多创业公司、很多不同的想法,而 NVIDIA 和 Google 都成功了,我认为是因为很多别的想法没那么能打,它们可能会灭绝。

12:21Bill Dally: 我不觉得我们趋同了,TPU 和 GPU 看起来很不一样,除非从两万英尺的高空看。应用决定了一些需求:有矩阵乘法,就要矩阵乘法单元;有 softmax 和归一化,就要能算超越函数的向量单元;还要一定的内存容量、内存带宽和通信带宽。谁都得有这些。但怎么组合,细节很多。多年来 NVIDIA 在数值格式上一直领先:几年前我们发了一篇讲向量缩放的论文,由此有了 NVFP4,后来各种 MX 格式也跟进了。这个细节能让你在同样精度下,比只用 FP8 的对手快 2 倍。稀疏性也一样:2015 年我和 Song Han 写过一篇论文,讲神经网络天然非常稀疏,从 Ampere 这一代开始我们在硬件里支持稀疏。不是每个怪虫子都有这个。

13:41Bill Dally: Google 设计 TPU 有个很大的优势:直到最近,它们只有一个客户,就是 Google 自己,所以能想做什么就做什么。NVIDIA 很幸运,在推理和训练两个市场都占 68% 的份额,但随之而来的是我们有很多客户,要让他们都满意。他们会提功能需求,如果是够大的客户,我们至少得认真对待,在硬件里加进东西,让这些不同的人都满意。如果能只为自己造最想要的东西,我们也许能做得更好。所以我有点羡慕。我也很喜欢他们训练 TPU 用的 3D 环面互连。我在 MIT 和 90 年代在 Cray 造过很多 3D 环面网络的超算,能从里面看到我那本书的合著者 Brian Towles 的手笔。

14:58Bill Dally: 说到寒武纪大爆发:90 年代末到 2000 年代初,硅谷有不下一百家图形芯片创业公司。适者生存之后只剩两家,NVIDIA 和 ATI,后者被 AMD 收购了。我认为这一轮的结局很可能也差不多。

16:03Norm Jouppi: Bill 说得对,有些操作我们都得支持,矩阵乘法、向量运算。TPU 不同的地方在于,第一代是块只做推理的 PCIe 卡,但从第二代开始,它就是按超算来设计的,用的就是 Bill 和 Brian 书里的环面网络。所以我们没经历过那种痛苦:加一个功能,又和别的功能打架。我们从一张白纸、一个超算设计起步,液冷也已经用了八年。

17:40 · 系统才是护城河

17:36Dave Patterson: NVIDIA 和 Google 都是世界级的系统公司。系统工程,也就是互连、内存、供电、散热、封装,是不是已经成了比芯片本身更大的护城河?你们在互连上的做法,怎样解决扩展时的关键瓶颈?

17:59Bill Dally: 产品是整个系统:不只是 GPU,不只是装着 GPU、CPU 和网络设备的主板,而是所有硬件、所有软件,以及让它们配合得好的配置。大约从 2015 年的 Pascal 这一代开始,我们推出了 DGX SuperPOD,那时我们还没有大规模网络,也还没收购 Mellanox。我们告诉你:用这个交换机,照这样配置,你就能搭起大概一万块 GPU,一开机就能用。相比之下,我们给美国能源部造大型超算的时候,比如 2011、2012 年的 Titan,还有后来的 Summit 和 Sierra,硬件都好了之后,通常还要花六个月调试,因为网络配置之类的小问题。你交付的是一整个系统,要在很长的训练任务里可靠运行、可用性非常高。这需要大量系统经验;而我们卖进的数据中心各有各的做法,所以一套标准化配置非常关键。不同在哪?主要在纵向扩展的网络。Google 用 3D 环面加光路交换机,可以绕开坏掉的部分。我们用更传统的做法:纵向扩展用 Clos 网络,加一条专有的超低延迟链路,避开以太网的很多开销;横向扩展用常规以太网。

19:57Norm Jouppi: 很多创业公司没有这种系统经验。Luiz Barroso 他们写过「数据中心就是一台计算机」,很多教训只有亲手造过、吃过苦头才明白。这让我们俩都比创业公司有优势。

03

软件护城河被填平,精度快压到头

有规格说明,让 Claude Code 重建调优过的库不难;每年 2 倍的提升只有 3 倍来自制程,精度还剩一两轮,好点子还能撑八到十年。

20:34 · 软件护城河正在被沙子填平

20:51Dave Patterson: 所以远不止芯片。NVIDIA 和 Google 的竞争优势,有多少在硬件,有多少在软件、编译器和库?换个说法:如果把网表甚至版图都给一家创业公司,但不给软件栈,它能竞争吗?

21:39Norm Jouppi: 设计 TPU 时,我们遵循我博士导师的建议:能在编译时做的,别拖到运行时。我们在一个十个人的房间里设计体系结构,其中两个人来自编译器团队,就是要造一台容易编译的机器。我们一直保持同样的总体架构,内存大小可以增减,就像给电脑加内存条,Word 照样能跑。框架在不断演进,所以要跟上、要支持好;也有人喜欢自己写内核,这也得支持。

23:00Bill Dally: 产品是整个系统,软件是其中不可分割的一部分。但从很多方面看,深度学习的软件问题,比之前通用 GPU 计算的问题容易。2006 年我们随 G80 推出 CUDA 时,有成千上万的应用能从并行里受益,但代码是串行的,很多是 Fortran 写的,要移植上百万行代码。我们在 CUDA 生态上投入巨大,包括语言本身,以及做 FFT、做矩阵运算的各种库。2010 年我们开始第一个深度学习项目,和斯坦福的 Andrew Ng 合作,后来成了 cuDNN。当时的感受是:哇,这个应用真小,真正要紧的只有几个内核,把它们做快,整个就快了。看过上百万行的气象代码之后,这简直是一股清风。它也变得非常关键:我们把一个模型跑起来,接下来六个月能把性能翻一倍,因为很容易在里面丢性能,所以我们做了各种工具来分析性能、融合内核、调优。软件非常关键。但它对今天的创业公司还算多大的门槛,我不知道,因为现在你只要跟 Claude 说「我要一套软件栈」,然后出去过个周末就行了。

24:55Dave Patterson: 这正是我的下一个问题。家里有程序员的人,大概都在半夜听过一声惊呼:「快看它做了什么!」写代码正在交给机器。如果曾经有软件护城河,它是不是要消失了?还是这么想太天真?

25:35Bill Dally: 多年打磨、调优过、别人用起来很顺手的库,仍然有优势。但我认为任何软件护城河都已经被大大削弱了,因为只要你有那个库的规格说明,放 Claude Code 出去把它重建一遍,并不难。

25:55Norm Jouppi: 我同意。护城河越来越浅,正在被沙子填上,人都快能走过去了。

25:36 · 精度快压到头了

26:10Dave Patterson: 早些年我们靠降低精度榨出了大量速度,但只能降到那么低。当再也削不掉比特的时候,一代又一代的提升从哪来?

26:38Bill Dally: 过去 14 年,从 2012 年的 Kepler 开始,也就是我们第一代认真把它当应用来做,我们基本做到了每年 2 倍,其中只有 3 倍来自制程。这和 90 年代微处理器的黄金时代正相反,那时全部来自制程。NVIDIA 在数值格式上领先,从 Kepler 的 FP32 一路到出货 NVFP4,Vera Rubin 里还有些新东西。还能再转一两轮,但数值精度已经快到头了。还有别的方向可以继续创新:稀疏性、电路、局部性,甚至更好的模型,都能让每瓦产出更多 token。低垂的果子已经摘完了,要爬到树的更高处,才能找到那些 2 倍的果子。不过我们有不少好点子,我认为至少能撑过接下来四五代,然后我就可以退休了。

27:52Dave Patterson: 是四五年,还是八到十年?

28:07Bill Dally: 八到十年。

28:22Norm Jouppi: 我们在数值上也有创新。Jeff Dean 最早做 AI 的时候,用 FP32 计算,但存储时把低 16 位直接截掉。对搞数值分析的人来说,「截断」这个词就像指甲刮黑板,但它管用。所以我们做 TPU 时,能用 BF16 跑那些原本在 CPU 上跑的程序,得到同样的结果。这很有用,我们不必花时间在系统层面排查模型为什么不对,能先确认运行正确,之后再采用 FP8、FP4 这些更小的格式。

04

芯片要三年,模型三个月一变

只能瞄在鸭子前面,把基本运算做好;混合专家对互连延迟要求更高;InferenceMAX 取代 MLPerf,成了大家真正看的基准。

29:46 · 芯片要三年,模型三个月一变

29:31Dave Patterson: Google 有个理论上的优势:推进 AI 前沿的人和造硬件的人在同一家公司;我也读到 NVIDIA 要在内部建立这方面的能力。能接触到推进前沿的人,比起只用开源模型,优势有多大?

30:16Norm Jouppi: 排行榜显示,有些开源模型表现相当好,所以专有模型要保持领先是一场竞赛;而因为能针对某个系统,GPU 也好 TPU 也好,把模型调得更好,专有模型总会有一些优势。至于做硬件的人去问做机器学习的人下一步是什么:有一点用。问题是,从最初的想法到能大规模量产的系统,要两年半到三年,而机器学习的人大概每三个月就有一个新点子。你没法为某个特定模型设计一台超级专用的机器,因为到时候它已经变了。你只能把基本操作,矩阵运算、向量运算,做好。

31:54Bill Dally: 我们自己做模型已经有一段时间了,内部有很多专业能力。但我们和 OpenAI 这样的公司有区别。OpenAI 在 Hot Chips 上发了一篇很棒的论文,讲他们的 Jalapeño 处理器,基本上是为跑一个模型而设计的,所以他们清楚矩阵运算、向量运算和内存带宽的比例。如果你要支持外面所有的模型,差异非常大,会改变资源配置,尤其是内存带宽和算力之间的比例,特别是注意力机制。这从 DeepSeek 推出 MLA 注意力开始,现在很多模型用混合方案:三层状态空间模型配一层完整的平方复杂度注意力,交替排列;很多人做稀疏注意力,先快速过滤,再挑出前 k 个去关注。注意力上的这些变化,对硬件提出了很不一样的要求。作为给所有模型供货的芯片商,你得通盘看这些模型,决定造什么样的硬件,让大家尽量满意、又不让谁特别不满,还得决定重点支持哪些。就像 Norm 说的,要瞄在鸭子前面,因为产品要过两年才出来,其间大家会想出很多我们还没见过的聪明点子。好的计算机体系结构就是干这个的。

33:50Norm Jouppi: 说到模型,最近更具颠覆性的一件事是混合专家,虽然大约七年前就提出了。它需要多得多的互连带宽,而要让用户快速得到回应,就必须把延迟压下去。训练更多是带宽问题;在计算机系统里,低延迟比高带宽更难做到。这就是我们最新一代 TPU 那套互连配置的出发点。

33:44 · 又一场基准之战

34:41Dave Patterson: 在 Hot Chips 上,感觉像重演 1980 年代的战斗。那时 RISC 公司拿 MIPS 当指标,跑什么由你自己定,每家都说别人在撒谎。现在是每秒多少 token:跑的什么,模型多大?当年的解决办法是 SPEC,大家围绕它团结起来。MLPerf 本想提前解决这个问题,但在 Hot Chips 上没人提它。它是不是不如 SPEC 成功?为什么?

36:14Bill Dally: 我喜欢 MLPerf,它做到了基准该做的事,提供一个公平的场地,戳穿各种吹嘘。但它有两个问题。第一,几乎没人提交,因为工作量很大。我们每一代都把所有基准跑一遍。但很多创业公司说:如果守全部规矩,我们看起来就没那么好,所以挑一个结果放到幻灯片上。第二,今天大家真正关心的是每瓦多少 token、每美元多少 token,大多数 MLPerf 基准不衡量这个,最接近的那个也没衡量对。现在大家展示的,Hot Chips 上也确实有这样的幻灯片,是 SemiAnalysis 的 InferenceMAX 基准,因为它恰好给出大家想要的数据,而且可复现:要在那张图上占一个点,你得有一个 GitHub 或 Hugging Face 上的代码仓库,任何人都能加载,在那台硬件上跑那个模型。由 SemiAnalysis 来跑。我认为今天大家用 InferenceMAX 多过用 MLPerf。

38:00Norm Jouppi: 跑 MLPerf 还需要一个全职团队,而且不是小团队。它既贵,又给不出你想要的数字。

05

能源是第一瓶颈

输电线难批,转向现场发电;燃气轮机约五年内卖光,锂电太贵,热电池有希望;按平准化成本,天然气仍难被核能打败。

38:34 · 影响与责任

38:47Dave Patterson: 这场淘金热的影响超出了技术本身,有经济的、社会的,甚至地缘政治的。最深远的影响是什么?硬件架构师和这些公司的领导者,要承担什么责任?

39:11Bill Dally: 我不叫它硅谷淘金热,而叫 AI 淘金热,它几乎惠及生活的方方面面。医疗上,AI 最早用于医学影像分析,现在帮医生做更准确的诊断,我们还能有私人健康教练。教育上,每个学生都能有一位懂得什么能激励他、他怎么学习的个性化导师。工程上,AI 已经在自动化很多工作。我最近要设计一个硬件,写好规格说明,把 Claude 放出去,修正了我规格里的几个错误之后,它给出了一个相当好的设计。它做的正是我要求的东西,但那不是对的东西;你得学会写好规格。但它让各个领域的工程师都往上走了一层:他们不再做初级的计算,而是决定要造什么,管理一队 agent 小兵去干活。很多业务流程也在这样变。娱乐上,它帮我们做出电影、游戏、音乐里很棒的作品。

40:58Bill Dally: 另一面,任何伟大的技术都可以用来行善,也可以用来作恶。AI 最明显的是深度伪造,我们要尽快建立来源认证:除非一张图经过了适当的认证,否则就默认它是假的。网络安全最近很受关注,但超过 80% 的成功攻击其实是社会工程,是钓鱼,AI 当然会让钓鱼更厉害。AI 还可能被用来设计病原体。最大的风险是,当每个人都变得更高效,就业的性质会改变:有些岗位要更多人,有些要更少人。我们得想办法让从一个岗位转到另一个岗位的人,过渡得轻松一些。

42:11Norm Jouppi: 最大的影响之一会在科学上。有很多非常惊人的成果,大众媒体通常不报道。Google DeepMind 把人类已知的所有蛋白质结构都预测了出来,还拿了诺贝尔奖。Google 几乎在所有应用里都用 AI,但很多地方很隐蔽,你不会注意到。比如最早的应用之一是 Google 地图:原来它会说「开 1000 英尺,在 Tennyson 街右转」,结合街景数据库之后,现在会说「在壳牌加油站右转」。这对人友好多了,尤其是晚上路牌不亮的时候。很多这样的细节,大家不会注意,用久了就习以为常。

43:41 · 最大的瓶颈:能源

44:03Dave Patterson: 投票结果来了。问题是:AI 芯片继续大规模部署,最大的瓶颈是能源供应、半导体产能、芯片散热、内存容量和带宽,还是芯片成本?还在变,但目前观众选的是能源供应。你们俩怎么看?

45:02Norm Jouppi: 能源供应。大家在谈 1 吉瓦、5 吉瓦的数据中心,而输电线路很难拿到许可,社区不喜欢大电线从头顶经过。所以超大规模云厂商和其他公司在转向现场发电。有的用天然气发电,但如果选址好,大部分电可以来自风和太阳,这是我们在尝试的路子:不用碳的能源,离网运行,这样建数据中心就不会挤垮别的东西。

46:09Bill Dally: 数据中心需要稳定的电,太阳只在白天出,风只在刮的时候有,但总体上,市场力量让各种电源大体平衡。建数据中心要三样东西:土地、电力、厂房,需求就由它们推动。很多人在数据中心旁边放天然气发电机,由于这轮需求,燃气轮机未来大约五年都卖光了。有公司把飞机发动机拆下来改成发电机。

47:01Dave Patterson: 可如果转向天然气,碳足迹会非常大。

47:06Bill Dally: 它常常是给可再生能源做后备,因为太阳能和风能靠不住,我们目前也还没有足够的储能。锂电池按每千瓦时算太贵,撑不过 20 到 100 小时的缺口,也就是一段时间没有晴天时要补上的那种。热电池看起来很有希望,已经有人开始用它建数据中心,按每千瓦时算足够便宜,能填上那个缺口。另外,数据中心是一项资本资产:你花了 $100 亿建它,就希望它全天候忙着,所以这边的消费者睡觉时,就把 GPU 算力卖给地球另一边的消费者。

48:18Norm Jouppi: 要看用途。我们希望给用户很快的回应。

48:25Dave Patterson: 核能也不排碳,还有小型模块化反应堆,Google 在这方面也发过公告。你们怎么看核能给数据中心供电?

48:41Bill Dally: 我们在看所有技术,有专人负责跟踪能源技术、给数据中心提建议。核能看起来有希望,但按平准化度电成本算,一直比天然气贵。天然气很难打败,即使考虑碳排放、给天然气加上碳封存,我认为它在平准化度电成本上仍然胜过核能。我很喜欢抽水蓄能,但不太现实。

06

DRAM 微缩停了,需求没停

内存厂同样的零件能卖好几倍价钱,扩产动力反而变小;它们过去几十年总是过度建厂、价格崩盘,这次会不会不同,没有答案。

48:53 · 内存:DRAM 微缩停了

49:35Dave Patterson: 现在进入观众提问。票数最高的问题:AI 硬件的供应链瓶颈,也就是需求远超供给,解决起来最大的挑战是什么?

49:55Bill Dally: 最大的挑战是建一座晶圆厂要很长时间。眼下最痛的是内存,内存厂商乐坏了,同样的零件能卖供应充足时的好几倍价钱。所以它们建厂的动力,可能比本来要小。总是这样:你预测需求,需求超出预测,而扩产要两三年。

50:30Norm Jouppi: 我要是更早留意,也许能预见到一些,因为 DRAM 的密度微缩基本已经停了。就算不算 AI,计算机也需要越来越多的内存,所以我认为天平终于倾斜了,相当长一段时间里,DRAM 会是个更好的生意。它过去是最差的生意之一。

51:05Dave Patterson: 你教过我,SRAM 早就到平台期了。

51:10Norm Jouppi: DRAM 基本也到平台期了,但需求没有。这对内存公司会很有意思,它们的历史就是过度建厂、然后价格下跌,这样循环了几十年。现在它们会不会说:我们利润很高,现在这样有什么不好?

51:38Bill Dally: 还好,竞争有望解决这个问题:有人利润很高,想多要一点份额,就会扩建工厂,然后别人也跟着扩。但晶圆厂很贵,而且不只建厂慢,把良率提上来也慢。

07

量子没用、烧进芯片太险,AI 会设计芯片

量子计算是大计算小数据,AI 正相反;前沿模型每月一版,把权重烧进 ROM 风险太大;模拟计算会输;验证占了 75% 的人力。

51:52 · 量子、把模型烧进芯片、边缘与模拟计算

51:56Dave Patterson: 等量子计算成熟了,它对今晚讨论的话题影响有多大?

52:04Bill Dally: 非常非常小。量子计算是很棒的技术,但真正好的量子算法只有两个:一个是模拟量子化学,一个是 Shor 算法,分解两个大质数的乘积,破解很多现代密码。还有第三个,用于优化的 Grover 算法,但它只带来平方级加速,不是指数级。所有这些算法都利用了一点:量子计算机是计算量大、数据量小的机器。最乐观的设想,是做到几千个纠错后的量子比特,靠叠加的指数加速,在几千个比特上做海量计算。AI 正好相反,是数据量大、计算相对少的问题。所以量子计算不会对 AI 的训练或推理产生可测量的影响。

53:35Norm Jouppi: 它会产生影响的地方不在计算,而在通信:量子传输和量子密码。

54:46Dave Patterson: 你们怎么看 Etched 或 Taalas 这类架构,基本上是把大模型「刻」进芯片?

55:57Bill Dally: Taalas 是最极端的:他们打算把一组权重烧进 ROM,密度大概是 SRAM 的四倍,但远不如 DRAM。据我了解,Etched 一开始就是这个思路,后来越来越可编程。走 Taalas 的路,你得相信有某个模型变得不快,可前沿模型大约每月出一个新版本,就连开源模型也每隔几周就有一个聪明的新点子,任何和某个模型贴得太紧的设计都会被打破,即使权重是可编程的。

56:33Norm Jouppi: 我们还在早期,可编程性极其重要。

56:39Dave Patterson: 这和 CPU 完全是两个世界。CPU 有上百万行遗留代码、几百万程序员;这里代码量很小,创新压力巨大,硬件的人和算法的人都更容易创新和部署。挑一个模型、及时做出芯片,是个有意思的赌注。那集中计算和分布式计算怎么平衡?边缘硬件呢?

57:35Bill Dally: 我们有大量产品进了自动驾驶汽车,某个级别以上的每辆奔驰,自动驾驶功能里都有 NVIDIA 的处理器和软件。能放在数据中心,就放在数据中心,因为在那里提供算力经济得多。但如果因为延迟,或者断网时也得可靠运行,或者采集的数据太多没法全部传回去,你就得在边缘计算。一般规则是:除非这些因素卡住你,否则就放在数据中心。

58:15Norm Jouppi: Google Pixel 手机里有边缘 TPU,处理相对简单的任务,比如语音识别。太大的任务,在有网络时就交给云端。

59:11Dave Patterson: 既然低精度也能得到差不多的 AI 结果,模拟计算会不会有一席之地?在模拟电路里做乘法那么优雅。

59:35Norm Jouppi: 以造真实系统为生,就得测试它。如果一个东西是模拟的,每次结果不一样,就很难测试这个器件到底对不对。测晶圆时,你要用极高的速度送测试序列、读回比特,必须一位一位对上。

1:00:27Bill Dally: 就算能测试,我也很少见到模拟计算真有优势,我们一直在重新评估,因为这个说法很诱人:把激活值当电流输入,把电阻看成电导,电压等于电流乘以电导,乘法就白来了,求和也是白来的。但没有白来的乘法。更糟的是,模拟电路没法可靠地把一个值存很久,放在电容上会漏掉,在现代工艺里尤其如此。把模拟值搬到远处也很贵。所以大家通常做一个小的存内计算阵列,做激活乘权重的矩阵乘法,把结果加起来,然后必须做模数转换。按模数转换的基本能耗算,如果要 8 位精度,大约只能做到每瓦 5 万亿次运算,而我们用常规数字技术做的加速器能到每瓦 100 万亿次。如果能一直待在模拟域、不用转换,也许能做出有吸引力的器件;一旦要转换,你就输了。

1:02:57 · 由 AI 来设计芯片

1:02:41Dave Patterson: 我那些年轻同事对用 AI 设计硬件很兴奋。下一代芯片设计软件,会和过去 30 年用的完全不同吗?

1:02:57Bill Dally: 我认为芯片设计软件会是一个你可以对它说话的大模型:我要一块这样的芯片。EDA 公司有很多专业能力,也有很多需要接起来的零件,大模型得懂得怎么调用那些工具去做布局布线和验证。我看设计一块芯片时人的时间花在哪,是验证。大模型非常擅长写测试:确保这个能用,覆盖所有边界情况,它就给你写出一整套测试。我们 75% 的人力都花在那里。

1:03:47Norm Jouppi: 我也想说这个。Hot Chips 上发表的结果显示,设计本身,也就是版图和电路性能,大约能提升 10%;但团队里最大的部分是设计验证,他们最需要帮忙。

1:04:08Bill Dally: 说人的生产率,我们看到的不止于此:普通工程师翻倍,特别好的工程师 10 倍。有一项针对特定任务的基准研究,平均下来大约 3 倍。再说一次,AI 里真正的淘金者是攻垂直领域的人,现在有一大批创业公司在攻 EDA 这个垂直领域。

1:04:39 · CPU 该离 GPU 多近

1:05:15Dave Patterson: 为了照顾受内存限制的工作负载,要不要把 CPU 放得离 GPU 更近?

1:05:21Bill Dally: 我们的已经挺近了,通过芯片间的 NVLink 相连,那是一条单端的超高速链路。离得近能得到什么?第一,我们给这条链路配足了带宽,让 Vera CPU 上全部的 LPDDR5 带宽,大概是每秒 1.8 TB,都能被导到连着的两块 GPU 中的任何一块,所以链路永远不是瓶颈。第二是延迟:从那块内存取数本身延迟就比较高,链路增加的那一点不要紧。至于启动内核这类控制交互,别处的延迟已经够大,链路再快也不关键。我们的 GPU 上还有几十个小 RISC-V CPU,做用户看不到的杂务。对用户编程用的 CPU 来说,芯片间 NVLink 的距离差不多刚好。

1:07:00Norm Jouppi: 长期看,挂在加速器旁边的 CPU 主要做杂务。现在有意思的是 agent 式计算。如果你说「给我写个程序」,它总得在某处编译,而你不会在脉动阵列矩阵乘法器上编译。你需要真正的 CPU,配上像样的内存系统和网络,放在同一个数据中心里,做编译、工具调用这类事,不过不必紧挨着加速器。

1:08:36 · 给下一代的建议

1:08:28Dave Patterson: 观众里有个 10 岁的孩子问:为了将来,你们有什么建议?

1:08:50Bill Dally: 学数学和科学。数学和基础科学的扎实功底,没有什么可以替代。

1:09:06Norm Jouppi: 沟通能力也非常重要,学会把文章写好。我知道大模型能替你写,但总会有手边没有它、你照样得把话说清楚的时候。

1:09:33Dave Patterson: 那对今天刚入行的年轻计算机架构师和工程师呢?

1:09:57Bill Dally: 今天的世界很不一样。我职业生涯早期在贝尔实验室设计过一台机器,用铅笔在描图纸上画出所有电路,由技术员做版图,第一次流片就成功了,没做过一次仿真。现在,手工逻辑设计的本事已经毫无价值了。我有三条建议。第一,精通一个垂直领域,因为今天很多价值在于懂得该设计什么,而不是手工实现它的本事,AI 会帮你实现。第二,对计算机技术要有很宽的理解。很多架构师把自己限制住了:去问做内存的人能不能造出某种内存,对方说不行。如果你对电路设计有宽泛的理解,就会发现对方的思路不够开阔。贝尔实验室那台机器上有一种 3T DRAM,所有做电路的人都说行不通,我自己做了原型,说服自己它能行。第三,让 AI 成为你的搭档,想清楚在人和 AI 的合作里,哪部分是人独有的,并且练好它。去做 AI 能做的事没有意义,它会做得更好。但至少到现在,你还得能看出它什么时候搞砸了。我肯定看得出来,也许是因为我会手工做。或者,你可以让一个 AI 去检查另一个。

1:12:23Norm Jouppi: 计算机体系结构有 70 多年的历史,一路积累了很多教训。最好的办法之一,就是去学这些历史教训,有的今天还适用,有的不适用了。前面七八次都没做成的东西,没必要再发明一遍。

判断收口延伸

Indigo 的结论

两位架构师一手确认:能源排第一,DRAM 超级周期是结构性的而不是周期性的;Dally 自己承认 CUDA 护城河在变浅、Google TPU 有结构上的单一客户优势。在泡沫之争上,他们给出了「这次是真需求」最硬的卖方论证。

需要记住的几件事

  1. 能源是第一瓶颈:输电难批就现场发电,燃气轮机约五年卖光,热电池有望,核能在平准化成本上仍难赢天然气。
  2. DRAM 微缩停了、需求没停,内存超级周期是结构性的;但内存厂会不会长期维持高价,Jouppi 留了问题没答。
  3. TPU 的单一客户优势对 NVIDIA 的多客户税,Dally 说「有点羡慕」:这是商业模式的差别,不是技术强弱。
  4. SemiAnalysis 的 InferenceMAX 取代 MLPerf,成了胜出的基准:它衡量的正是每瓦多少 token。

可回查的判断

判断谁说的何时见分晓证据多硬
能源供应是 AI 芯片部署的第一瓶颈Dally、Jouppi 和现场投票当下一手,可信度高
DRAM 微缩已停、需求没停,内存会长期供给受限Jouppi持续一手架构判断
CUDA 这类软件护城河被 AI 显著削弱Dally、Jouppi进行中一手,自己承认
每年 2 倍的性能提升还能撑四五代(八到十年),数值精度快到头Dally到约 2034 年一手
量子计算对 AI 训练和推理几乎没有可测的影响Dally长期一手
燃气轮机约五年内已经卖光Dally现在一手转述

放回主线

证实

约束正从算法层迁到物理层 两位顶级架构师把这条判断顶格坐实:能源是第一瓶颈,DRAM 微缩已停。

补充

你拥有的不是模型 CUDA 护城河被 AI 填平,加上「系统就是护城河」:价值从软件库往系统、规模和生态迁移。

补充

AI capex 单引擎 「淘金热是真的、不同于 90 年代泡沫」,是「不得不烧、需求是真的」一侧最硬的卖方技术论证。

证实

Perplexity Lily:真信号是撞到内存带宽墙 Perplexity 在使用一侧量到「卡住的是内存带宽」,Jouppi 在供给一侧说「DRAM 微缩停了」。

补充

SemiAnalysis:商品化在能力层价值在 harness 层 Dally 亲口说 InferenceMAX 已取代 MLPerf,SemiAnalysis 的影响力被两位架构师确认。

什么会让我改口

内存厂重回「过度建厂、价格崩盘」的老路,DRAM 结构性供给受限的判断就不成立了。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Video

The Silicon Gold Rush: How AI Is Driving the Development of New Chips

Bill Dally, Norm Jouppi · YouTube · 2026-09-10

The two people best placed to know: energy is the first bottleneck, DRAM has plateaued, and AI is filling in the CUDA moat.

Part 1 of 7 · 2:10
This gold rush rests on real demand

Real economic demand, an application simple enough that accelerating a few primitives works, no legacy code: unlike the '90s supercomputing boom. Both companies are paying the memory makers a lot.

Breakdown · 7 steps

  1. 01

    2:10 – 10:35

    This gold rush rests on real demand

    Real economic demand, an application simple enough that accelerating a few primitives works, no legacy code: unlike the '90s supercomputing boom. Both companies are paying the memory makers a lot. Read this part →

  2. 02

    10:35 – 20:51

    GPUs and TPUs haven't converged; the system is the moat

    NVIDIA leads on number formats and sparsity but carries a multi-customer tax; TPUs serve only Google and started as supercomputers from a clean sheet; 10,000-GPU systems that just work are experience startups lack. Read this part →

  3. 03

    20:51 – 29:31

    The software moat is filling in; precision is nearly used up

    With a spec, having Claude Code rebuild a tuned library isn't hard. Of the 2x-a-year gains only 3x came from process; precision has a turn or two left, and good ideas can last eight to ten years. Read this part →

  4. 04

    29:31 – 38:47

    Chips take three years; models change every three months

    Aim ahead of the duck and do the basic operations well; mixture of experts pushes interconnect latency; InferenceMAX has replaced MLPerf as the benchmark people actually watch. Read this part →

  5. 05

    38:47 – 49:35

    Energy is the first bottleneck

    Transmission lines are hard to permit, so data centers generate on site. Gas turbines are sold out for about five years, lithium is too expensive, thermal batteries look promising, and gas still beats nuclear on levelized cost. Read this part →

  6. 06

    49:35 – 51:56

    DRAM scaling has stopped; demand hasn't

    Memory makers can charge several times more for the same part, so they have less reason to build. For decades they overbuilt and prices crashed; whether this time differs is left open. Read this part →

  7. 07

    51:56 – 1:13:03

    Quantum won't help, burning models into chips is risky, AI will design chips

    Quantum is big computation on small data, AI the reverse; with frontier models updated monthly, weights in ROM are too risky; analog loses; verification takes 75% of the labor. Read this part →

Indigo's conclusion

First-hand confirmation from two architects: energy comes first, and the DRAM super-cycle is structural, not cyclical. Dally admits the CUDA moat is getting shallower and Google's TPUs have a structural single-customer edge. On the bubble question they give the strongest sell-side case that this time the demand is real.

How to read this A public conversation between two top chip architects. Both companies sell shovels, so “the gold rush is real and we win either way” suits them. Take the technical calls on energy, memory, moats and benchmarks as hard material, and discount “the gold rush is real” as shovel-seller optimism. A public stage with no adversarial follow-ups, but both make a few rare admissions.

What to remember

  1. Energy is the first bottleneck: transmission is hard to permit, so generate on site; gas turbines sold out for about five years; thermal batteries promising; nuclear still loses to gas on levelized cost.
  2. DRAM scaling stopped while demand didn't, so the memory super-cycle is structural; whether memory makers keep prices high for long, Jouppi leaves open.
  3. The TPU's single-customer edge against NVIDIA's multi-customer tax, which Dally says he's a little envious of: a business-model difference, not technical strength.
  4. SemiAnalysis's InferenceMAX has replaced MLPerf as the winning benchmark: it measures exactly tokens per watt.

What would change my mind

memory makers return to overbuilding and crashing prices, and the case for structurally constrained DRAM supply falls apart.

How to read this

A public conversation between two top chip architects. Both companies sell shovels, so “the gold rush is real and we win either way” suits them. Take the technical calls on energy, memory, moats and benchmarks as hard material, and discount “the gold rush is real” as shovel-seller optimism. A public stage with no adversarial follow-ups, but both make a few rare admissions.

Breakdown · 7 steps
  1. This gold rush rests on real demand
  2. GPUs and TPUs haven't converged; the system is the moat
  3. The software moat is filling in; precision is nearly used up
  4. Chips take three years; models change every three months
  5. Energy is the first bottleneck
  6. DRAM scaling has stopped; demand hasn't
  7. Quantum won't help, burning models into chips is risky, AI will design chips

Compiled from the video's captions, by speaker.

01

This gold rush rests on real demand

Real economic demand, an application simple enough that accelerating a few primitives works, no legacy code: unlike the '90s supercomputing boom. Both companies are paying the memory makers a lot.

01:24 · Two architects with parallel lives

2:10Dave Patterson: I get to introduce two friends I've known since the 1980s, when I was an assistant professor at Berkeley, Bill was a PhD student at Caltech and Norm was at Stanford. Bill Dally has been building network-connected computers ever since Caltech: at MIT, then back at Stanford, where he became chair of computer science, and since 2009 as chief scientist at NVIDIA, now also a senior vice president. Among his best-known contributions are wormhole routing and the book on interconnection networks he wrote with Brian Towles, and his stream-processing projects at Stanford were an important precursor to modern GPU computing. Norm worked on the MIPS project as a Stanford graduate student, then spent about a decade at Digital's Western Research Lab and went on to HP Labs. In 2013 his friend Jeff Dean asked him to come to Google and build hardware for deep learning. Norm was skeptical given all the past hype about AI, but Jeff convinced him: deep learning worked on everything they tried. He is now a fellow and vice president at Google, known for the TPUs and for memory-hierarchy work like victim caches and prefetch buffers. Their careers run strikingly parallel: both studied under John Hennessy at Stanford, both are in the National Academy of Engineering, both won the Eckert-Mauchly Award, the highest award in computer architecture.

06:10 · What makes this gold rush different

6:35Dave Patterson: The title is "The Silicon Gold Rush": intense innovation, investment and competition. From your seats at NVIDIA and Google, what defines this era for AI hardware, and what makes it fundamentally different from earlier eras of computer architecture?

6:56Bill Dally: Three characteristics make it very different. First, intense economic demand. AI works on everything it's applied to, as Jeff Dean said, and the demand for more tokens, more flops, more AI is insatiable. Second, the application evolves rapidly but is relatively simple compared with the big supercomputing applications that had millions of lines of code. A transformer is a relatively simple thing, so it's easy to see what you have to build to make it go fast. Third, people are willing to evolve very quickly; there are no dusty decks. Contrast the early '90s, when there was a computer architecture gold rush around supercomputing. There was no intense economic demand; it was fueled by DARPA's strategic computing program, whose funding made people think there was a big downstream market, which there wasn't. Lots of people founded companies, like Thinking Machines, and most went bust. The applications were dusty decks with up to close to a million lines of code. You could accelerate what looked like the important kernel, but Amdahl's law would bite you, because the other 99% of the code wasn't accelerated. So the ease of accelerating a few primitives in a simple application, plus real economic demand, makes this one stick. This is not the Lisp machine frenzy of the '80s or the supercomputing frenzy of the '90s. It delivers real value.

8:49Norm Jouppi: I keep an eye on the global economy, and the amount of money being invested is incredible. For a while New York had a hundred-year-old leaky tunnel to New Jersey and was a billion dollars short, and now Google has announced $105 billion of capex next year, with other companies investing similar amounts. There will be winners and losers, and everyone wants to race as fast as possible.

9:34Dave Patterson: In the actual gold rush, who are the prospectors and who sells the pickaxes?

9:42Bill Dally: The prospectors are the people finding verticals. Find the right vertical to apply AI to and you can strike it rich, just like the prospectors. NVIDIA and companies like Google that make the infrastructure are the Leland Stanfords, selling picks and shovels to the prospectors. Whether they strike it rich or not, we're going to make money.

10:12Norm Jouppi: Well, we're both still paying the memory companies a lot of money.

10:19Bill Dally: I think Micron made as much money last quarter as it had in the previous 19 years. It's interesting to see their revenue go up while the number of parts sold stays level.

10:31Norm Jouppi: Exactly. The margins are through the roof.

02

GPUs and TPUs haven't converged; the system is the moat

NVIDIA leads on number formats and sparsity but carries a multi-customer tax; TPUs serve only Google and started as supercomputers from a clean sheet; 10,000-GPU systems that just work are experience startups lack.

10:29 · Convergent evolution? GPUs and TPUs

10:35Dave Patterson: A colleague accuses training accelerators of convergent evolution: a max-size compute die with a large systolic matrix unit, as many HBM stacks as fit around the perimeter, and the fastest SerDes for a custom link. Why is he wrong? Do GPUs and TPUs take divergent approaches, and is there anything you like about the other's design?

11:22Norm Jouppi: Remember that after the Cambrian explosion there was also an implosion. There were all these weird and wonderful creatures, like giant pill bugs with strange appendages, and they didn't survive the extinction event. Same here: many startups with many different ideas, and I think NVIDIA and Google have both succeeded because a lot of those other ideas weren't as capable, and they may go extinct.

12:21Bill Dally: I don't know that we've even converged; I think TPUs and GPUs look very different, except from 20,000 feet. The application drives certain requirements. You have GEMMs, so you need matrix multiply units; you have softmax and norms, so you need vector units that can do transcendental functions; you need a certain memory capacity, memory bandwidth and communication bandwidth. Everything will have that. But there's a lot of nuance in how they're combined. For many years NVIDIA led the way on numerics: we published a paper on vector scaling a few years ago, NVFP4 came out of it, and the MX formats followed. That nuance can give you a 2x advantage over someone doing FP8 at the same accuracy. Similarly with sparsity: Song Han and I wrote a paper in 2015 on how naturally sparse neural networks are, and we put hardware support for sparsity in starting with Ampere. You don't find that in every bug-like creature.

13:41Bill Dally: Google has a big advantage in designing TPUs: until recently they had one customer, Google, so they could decide exactly what they wanted and do it. NVIDIA is fortunate to have a 68% share of both the inference and training markets, but with that comes lots of customers we have to keep happy. They come with feature requests, and if you're a big enough customer we at least have to entertain them and put things in the hardware to make all these different people happy. If we could build exactly what we wanted just for ourselves, we might do better. So I'm a little envious. And I'm fond of their 3D torus interconnect for training TPUs; I built a lot of 3D torus supercomputers at MIT and with Cray in the '90s, and I can see the hand of my co-author Brian Towles in it.

14:58Bill Dally: On the Cambrian explosion: in the late '90s and early 2000s there were no fewer than a hundred graphics chip startups in Silicon Valley. After survival of the fittest there were two left, NVIDIA and ATI, which was acquired by AMD. I think something very similar is the likely outcome of the current explosion.

16:03Norm Jouppi: Bill said it well: there are operations we both have to support, matrix multiply and vector operations. What was different about TPUs is that our first one was a PCIe card that only did inference, but from the second one on they were designed as supercomputers, with the torus network from Bill and Brian's book. So we had no painful steps of adding features that then conflicted with other features. We started with a clean sheet of paper and a supercomputer design, and we've been liquid-cooled for eight years.

17:40 · The system is the moat

17:36Dave Patterson: NVIDIA and Google are both world-class system companies. Has systems engineering, the interconnect, memory, power, cooling and packaging, become an even bigger moat than the silicon itself? How do your interconnect approaches address the critical bottlenecks in scaling?

17:59Bill Dally: The product is the whole system: not just the GPU or the board with GPUs, CPUs and networking gear, but all the hardware, all the software and the configuration that makes it work well together. Starting around our Pascal generation, around 2015, we offered the DGX SuperPOD, before we even had large-scale networking, before we acquired Mellanox. You use this switch, you configure it like this, and you could put together maybe 10,000 GPUs, turn it on, and it would work. Compare the big DOE supercomputers we built, Titan in 2011 or 2012, then Summit and Sierra: after the hardware was working you'd typically spend six months bringing it up because of little things in network configuration. You're delivering an entire system that has to run long training jobs reliably with very high availability. A lot of systems expertise goes into that, and a standardized configuration was critical because we sell into many data centers that do things differently. Where do we differ? Mostly the scale-up network. Google uses the 3D torus with optical circuit switches to route around bad parts. We take a more conventional approach, a Clos network with a proprietary, very low-latency link to avoid the overhead of Ethernet, and a conventional Ethernet scale-out network.

19:57Norm Jouppi: A lot of startups don't have that systems experience. Luiz Barroso and others wrote that the data center is a computer, and many of those lessons aren't obvious until you've felt the pain yourself building these machines. That gives both of us an advantage over startups.

03

The software moat is filling in; precision is nearly used up

With a spec, having Claude Code rebuild a tuned library isn't hard. Of the 2x-a-year gains only 3x came from process; precision has a turn or two left, and good ideas can last eight to ten years.

20:34 · The software moat is filling with sand

20:51Dave Patterson: So there's a lot more to it than chips. How much of NVIDIA's and Google's advantage is hardware versus software, compilers and libraries? Put differently: if you gave a startup the netlist and the layout but not the stack, could it compete?

21:39Norm Jouppi: In designing TPUs we followed my thesis advisor's advice: don't put off until runtime what you can do at compile time. We designed the architecture in a room of ten people, two of them from the compiler team, to make a machine that was easy to compile to. We've kept the same general architecture; memory sizes can grow or shrink, like adding DIMMs to a PC and still running Word. Frameworks keep evolving, so you need good support for them, and some people like writing their own kernels, so you need to support that too.

23:00Bill Dally: The product is the whole system, and software is an integral part of it. But in many ways deep learning is an easier software problem than general GPU computing, which led up to it. When we launched CUDA with G80 in 2006, there were thousands of applications that would benefit from running in parallel but had serial code, many in Fortran, with million-line codes to port. We invested hugely in the CUDA ecosystem, the language and many libraries for FFTs, matrix operations and so on. When we began our first deep learning effort in 2010, a collaboration with Andrew Ng at Stanford that became cuDNN, it was: wow, a tiny application, only a couple of kernels really matter, make those fast and the whole thing is fast. A breath of fresh air after million-line weather codes. It also became critical: we'd get a model running and then double its performance over six months, because there are easy ways to lose performance, so we built tools to profile, fuse kernels and tune. Software is a very critical part. I don't know how much of a barrier it will be to startups these days, because right now you just tell Claude you want a software stack and go away for a weekend.

24:55Dave Patterson: That's my next question. Anyone married to a programmer has heard shouts in the middle of the night: "look what it did." Coding is being done by machines. If there was a software moat, is it about to disappear, or is that a Pollyannaish view?

25:35Bill Dally: There's still an advantage to libraries tuned and honed over the years that are easy to apply. But I think any software moat has been significantly degraded, because if you have the spec of a library, turning Claude Code loose to recreate it is not that difficult.

25:55Norm Jouppi: I agree. The moat is getting shallower. It's filling up with sand; people are starting to be able to walk across.

25:36 · Precision is nearly used up

26:10Dave Patterson: We squeezed a lot of speed in the early years by reducing precision, but we can only go so low. What happens to generation-over-generation gains when we can't shave bits anymore?

26:38Bill Dally: We've pretty much hit 2x per year for the past 14 years, starting with Kepler in 2012, the first generation where we took this seriously as an application, and only 3x of that came from process technology. That's the opposite of the microprocessor heyday of the '90s, when it all came from process. NVIDIA led on numerics, from FP32 in Kepler to shipping NVFP4, and there are some new things in Vera Rubin. There are a couple more turns, but we're getting near the end on numerical precision. There are other axes to keep innovating on: sparsity, circuits, locality, even better models that give more tokens per watt. The low-hanging fruit has been picked, and we need to climb higher up the tree to find those 2x fruits. But we have a bunch of good ideas, and I think they can get us through at least the next four or five generations, and then I can retire.

27:52Dave Patterson: Is that four or five years, or eight to ten?

28:07Bill Dally: Eight to ten years.

28:22Norm Jouppi: We had numerics innovations too. When Jeff Dean was doing the initial AI work, they computed in FP32 but stored values by truncating the low-order 16 bits. To a numerical analyst, truncation is fingernails on a chalkboard, but it worked. So when we built TPUs we could run programs in BF16 that had run on the CPUs of the time and get the same results. That was powerful, because we didn't have to spend time debugging at the system level why a model wasn't working. We could verify correct operation, and then we adopted smaller formats like FP8 and FP4.

04

Chips take three years; models change every three months

Aim ahead of the duck and do the basic operations well; mixture of experts pushes interconnect latency; InferenceMAX has replaced MLPerf as the benchmark people actually watch.

29:46 · Chips take three years; models change every three months

29:31Dave Patterson: One of Google's theoretical advantages is having people pushing the state of the art in AI alongside the people building hardware, and I've read that NVIDIA is building that expertise in-house too. How big an advantage is access to those people, versus just using open models?

30:16Norm Jouppi: Leaderboards show some open models performing quite well, so there's a race for proprietary models to stay ahead, and because you can tune a model better for a particular system, GPU or TPU, there will always be some advantage to the proprietary ones. As for hardware people talking to ML people about what's next: to some extent. The problem is that it takes two and a half or three years from an initial idea to a system manufactured in volume, and the ML people come up with a new idea every three months. You can't design a super-specialized machine for a particular model, because it will be different by then. You have to do the basic operations, matrix and vector operations, and do them well.

31:54Bill Dally: We've built our own models for some time and have a lot of internal expertise. But there's a distinction between us and a company like OpenAI, which had a great paper at Hot Chips on its Jalapeño processor, designed pretty much to run one model, so they know their relative mix of matrix ops, vector ops and memory bandwidth. If you have to support all the models out there, there's a very big variety that shifts the provisioning, especially memory bandwidth versus math, particularly with attention. That started with DeepSeek's MLA attention, and now many models use hybrid schemes, alternating three layers of state-space models with one layer of full n-squared attention, or sparse attention that does a quick filter and attends to the top k. That variation drives very different demands on the hardware. As a silicon supplier supporting all models, you have to look across them and decide the right hardware to make everybody as happy as you can without making anybody really unhappy, and which to emphasize. And as Norm said, you have to aim ahead of the duck, because it will be a couple of years before it's out and people will come up with clever ideas we haven't seen yet. That's what good computer architecture is about.

33:50Norm Jouppi: One of the more disruptive recent developments, though proposed about seven years ago, is mixture of experts, because it requires a lot more interconnect bandwidth, and for quick responses to users you really have to push down latency. Training is more a bandwidth problem; low latency is harder to get in computer systems than high bandwidth. That's what drove the interconnect configuration of our latest TPU.

33:44 · The benchmark wars, again

34:41Dave Patterson: At Hot Chips it felt like the battles of the 1980s. Back then RISC companies quoted MIPS, and it was up to you what you ran; everyone else was lying. Now it's tokens per second: running what, on how big a model? The solution then was SPEC, which companies rallied around. MLPerf was an attempt to anticipate this problem, but nobody mentioned it at Hot Chips. Was it less successful than SPEC, and why?

36:14Bill Dally: I like MLPerf; it did what benchmarks are supposed to do, a level playing field to cut through the BS. But it had two issues. First, almost nobody submitted to it, because it was a lot of work. We did every benchmark every generation. But a lot of startups said: if we played by all the rules we wouldn't look that good, so we'll cherry-pick one result for our slides. Second, what everybody really cares about today is tokens per watt or tokens per dollar, and most MLPerf benchmarks don't address that, and even the closest one doesn't address it quite right. What everybody shows, and you did see slides of it at Hot Chips, is SemiAnalysis's InferenceMAX benchmark, because it hits exactly the data people want, and in a very reproducible way: to get a point on that chart you need a repository of code, on GitHub or Hugging Face, that anyone can load and run on that hardware with that model. SemiAnalysis runs it. I think people use InferenceMAX more than MLPerf today.

38:00Norm Jouppi: MLPerf also required a full-time team, and not a small one. It was both expensive and didn't give you the number you wanted.

05

Energy is the first bottleneck

Transmission lines are hard to permit, so data centers generate on site. Gas turbines are sold out for about five years, lithium is too expensive, thermal batteries look promising, and gas still beats nuclear on levelized cost.

38:34 · Impact and responsibility

38:47Dave Patterson: This gold rush has economic, social and even geopolitical implications beyond the technology. What are the most profound impacts, and what responsibility do hardware architects and the leaders of these companies bear?

39:11Bill Dally: I wouldn't call it the silicon gold rush; it's the AI gold rush, and it benefits almost every part of our lives. In medicine, AI started in image analysis and now helps doctors make more accurate diagnoses, and we can have personal health coaches. In education, every student could have a personalized tutor that understands what motivates them and how they learn. In engineering, AI is already automating many tasks. I recently needed to design a piece of hardware: I wrote the spec, set Claude loose, and after I fixed a couple of errors in my spec, it produced a pretty good design. It produced exactly what I asked for, which wasn't the right thing; you have to learn to write a good spec. But it lets engineers in every field move up the ladder. They're no longer doing junior-level calculations; they decide what needs to be built and manage a team of agent minions that carry out the work. Business processes are moving the same way. In entertainment it helps produce wonderful work in movies, games and music.

40:58Bill Dally: On the flip side, any great technology can be used for good or evil. The most obvious are deepfakes, and we need to move rapidly to provenance and authentication, so that unless an image is properly authenticated, you assume it's fake. Cyber gets a lot of attention, although over 80% of successful attacks are human engineering, phishing, which AI of course makes better. AI can be used for pathogen design. And the big risk is that as everybody becomes more productive, the nature of employment will change, with more people needed in some jobs and fewer in others. We need to find ways to ease that transition for the people moving from one to the other.

42:11Norm Jouppi: One of the biggest impacts will be in science. There have been super-impressive results that don't make the popular press. At Google DeepMind they folded all the proteins known to mankind and got a Nobel Prize. Google also uses AI in virtually all its applications, often subtly. Early on, Google Maps would say "drive 1,000 feet and turn right on Tennyson Street"; by integrating the Street View database it now says "turn right at the Shell gas station", which is much more accessible, especially at night when street signs aren't lit. People won't notice many things like that and will just take them for granted.

43:41 · The biggest bottleneck: energy

44:03Dave Patterson: Here's the poll: the biggest bottleneck to continued widespread deployment of AI chips is energy availability, semiconductor manufacturing capacity, cooling, memory capacity and bandwidth, or the cost of the chips? It's still changing, but right now the audience says energy availability. What do you two think?

45:02Norm Jouppi: Energy availability. People talk about gigawatt or 5-gigawatt data centers, and transmission lines are very hard to get permits for; communities don't like big power lines over their houses. So hyperscalers and others are moving to on-site power generation. Some use natural gas, but if you put the data center in the right place you can get most of the power from wind and solar, which is the approach we're trying to take: carbon-free energy, off the grid, so you can build data centers without clobbering everything else.

46:09Bill Dally: You need steady power, and the sun only shines by day and the wind only when it blows, but in general the supply sources have been reasonably balanced by market forces. Land, power and shell are the three things you need for a data center, and they drive the demand. Many people are co-locating natural gas generators, and gas turbines are sold out for about the next five years because of this demand. Some companies are taking engines off airplanes and converting them into generators.

47:01Dave Patterson: But if we go to natural gas, we'll have a giant carbon footprint.

47:06Bill Dally: It's often used to back up renewables, because you can't count on solar and wind and we don't yet have the storage. Lithium batteries are too expensive per kilowatt-hour to ride through the 20-to-100-hour gaps you need to bridge when you don't have sunny days for a while. Thermal batteries look very promising, and people are starting to build data centers with them; they're cheap enough per kilowatt-hour to fill that gap. And a data center is a capital resource: you spent $10 billion building it and want it busy around the clock, so when consumers here are asleep you sell the GPU cycles to consumers on the other side of the world.

48:18Norm Jouppi: It depends on the use. We like quick responses for our users.

48:25Dave Patterson: Nuclear is also carbon-free, and there are small modular reactors; Google has made an announcement in this space. Thoughts on nuclear for data centers?

48:41Bill Dally: We're looking at all technologies; we have someone whose job is to track energy technologies and advise on data centers. Nuclear looks promising, but it has always been expensive in levelized cost per kilowatt-hour compared with natural gas. Natural gas is very tough to beat, and even with carbon sequestration I think it still beats nuclear on levelized cost. I'm a big fan of pumped hydro, but it's not very practical.

06

DRAM scaling has stopped; demand hasn't

Memory makers can charge several times more for the same part, so they have less reason to build. For decades they overbuilt and prices crashed; whether this time differs is left open.

48:53 · Memory: DRAM scaling has stopped

49:35Dave Patterson: Now audience questions. The most popular: what are the biggest challenges to resolving the supply chain barriers for AI hardware, where demand outruns supply?

49:55Bill Dally: The biggest challenge is that it takes a long time to build a fab. Right now the most pain is on the memory side, and the memory manufacturers are loving it, because they can charge many times more for the same part than if it were plentiful. So maybe they're less motivated to build fabs than they otherwise would be. It's always this: you project demand, demand exceeds the projection, and there's a two- or three-year delay to build production capacity.

50:30Norm Jouppi: If I'd been paying more attention, I might have seen this coming, because DRAM density scaling has basically stopped. Computers keep needing more memory, even ignoring AI, so I think the balance has finally tipped, and for quite a while DRAM will be a better business. It used to be one of the worst businesses.

51:05Dave Patterson: You've taught me that SRAM had already plateaued.

51:10Norm Jouppi: And DRAM has basically plateaued, but demand hasn't. It will be interesting for the memory companies, whose history is to overbuild fabs and then watch prices fall; they've had a couple of decades of that. Will they now just say, we're highly profitable, what's wrong with the way things are?

51:38Bill Dally: Fortunately, competition will hopefully fix that: someone highly profitable will want a bit more share and build out a fab, and then the others will too. But fabs are expensive and take a long time, not only to build but to get the yield up.

07

Quantum won't help, burning models into chips is risky, AI will design chips

Quantum is big computation on small data, AI the reverse; with frontier models updated monthly, weights in ROM are too risky; analog loses; verification takes 75% of the labor.

51:52 · Quantum, models burned into chips, the edge, analog

51:56Dave Patterson: How significant will quantum computing be for tonight's topics once it matures?

52:04Bill Dally: Very, very little. Quantum computing is a great technology, but there are two really good quantum algorithms: simulating quantum chemistry, and Shor's algorithm for factoring the product of two large primes and breaking a lot of modern cryptography. There's a third, Grover's algorithm for optimization, but it only gives a quadratic speedup. All of them exploit the fact that quantum computers are large-computation, small-data machines. In people's wildest dreams they'd get a few thousand error-corrected qubits and do tremendous amounts of computation on a few thousand bits. AI is the other way around: a large-data, relatively small-computation problem. So quantum computing won't have a measurable impact on AI training or inference.

53:35Norm Jouppi: Where it will have an effect is communication rather than computation: quantum transmission and cryptography.

54:46Dave Patterson: What do you think of architectures like Etched, or Taalas, that essentially etch an LLM onto the chip?

55:57Bill Dally: Taalas is the extreme: they plan to burn a set of weights into ROM, which is maybe four times as dense as SRAM and way less dense than DRAM. My understanding is that Etched started with that thesis and has become more and more programmable. With the Taalas approach you have to believe some model isn't changing very fast, but the frontier models come out with a new release every month or so, and even among open-source models there's a clever new idea every couple of weeks that would break anything matched too closely to a model, even with programmable weights.

56:33Norm Jouppi: We're still in the early days, and programmability is critically important.

56:39Dave Patterson: This is a completely different universe from CPUs with million-line legacy code and millions of programmers. It's a small amount of code with tremendous pressure to innovate, which makes it easier for hardware people and algorithm people alike. Picking a model and getting a chip out in time is an interesting bet. What about centralized versus distributed compute, and edge hardware?

57:35Bill Dally: We ship a lot of products into autonomous vehicles; every Mercedes above a certain level has an NVIDIA processor and NVIDIA software for its self-driving features. If you can put the compute in the data center, you do, because it's much more economical. But for latency, or because you have to operate reliably through a network partition, or because you're acquiring so much data you can't send it all back, you may need to compute at the edge. The general rule: do it in the data center unless one of those factors stops you.

58:15Norm Jouppi: Google Pixel phones have an edge TPU that handles relatively simple things, like speech recognition. Anything too big goes over the network if it's available.

59:11Dave Patterson: Since low precision achieves comparable AI results, will there ever be room for analog rather than digital systems? You can multiply so elegantly in analog.

59:35Norm Jouppi: When you build real systems for a living, you have to test them. If something is analog and doesn't give the same result each time, it's really hard to test whether the device works. When you probe wafers, you send test sequences and read back bits at tremendous speed, and they have to match one for one.

1:00:27Bill Dally: Even if you could test it, I've seen few cases where analog has an advantage, and we re-evaluate it continuously, because it's a very seductive argument: put the activations in as current, treat a resistor as a conductance, V equals I times G, and you get a free multiply, and a free sum too. But there's no free multiply. Worse, in analog you have no reliable way to store a value for more than a very short time; a capacitor leaks, especially in modern processes. Moving an analog value any distance is very expensive. So people build a small in-memory array, do a matrix multiply of activations times weights, sum the results, and then need an analog-to-digital conversion. The fundamental energy of that conversion limits you, at 8-bit precision, to something like 5 TOPS per watt, whereas we've built digital accelerators at 100 TOPS per watt. If you can stay analog without converting, you might build an attractive device; if you have to convert, you lose.

1:02:57 · Chips designed by AI

1:02:41Dave Patterson: My younger colleagues are excited about using AI for hardware design. Will the next chip design software be radically different from what we've used for 30 years?

1:02:57Bill Dally: I think chip design software will be an LLM you talk to: I want this chip. The EDA companies have a lot of expertise and many of the pieces that need to be plugged together, and the LLM has to know how to run the tools for place-and-route and verification. When I look at where human time goes in designing a chip, it's verification. LLMs are really good at writing tests: make sure this works, cover all the edge cases, and it writes a whole test suite. That's where 75% of our labor goes.

1:03:47Norm Jouppi: I was going to say the same. The results presented at Hot Chips showed around 10% improvements in the design itself, the layout and circuit performance, but the biggest part of the team is design verification, and they need all the help they can get.

1:04:08Bill Dally: On human productivity we're seeing more than that: a doubling for average engineers and 10x for really good ones. In one benchmark study for a particular task, the average came out around 3x. And again, in AI the real gold miners are the people attacking a vertical, and a whole bunch of startups are attacking the EDA vertical.

1:04:39 · How close should the CPU be?

1:05:15Dave Patterson: What about bringing the CPU closer to the GPU, to make room for memory-bound workloads?

1:05:21Bill Dally: Ours are pretty close: they sit on our chip-to-chip NVLink, a single-ended, very high-speed link. What do you get from closeness? First, we provision that link so that all of the Vera CPU's LPDDR5 bandwidth, something like 1.8 terabytes per second, can be routed to either of the two GPUs attached, so the link is never the bottleneck. Second, latency: fetching from that memory is already relatively high-latency, so the little the link adds isn't material. And for control interactions like launching kernels, there's enough latency elsewhere that a faster link isn't critical. We also have many tens of little RISC-V CPUs on the GPU for housekeeping the user never sees. For the CPUs users program, NVLink chip-to-chip is about as close as you want.

1:07:00Norm Jouppi: Long term, the CPU attached to an accelerator will mostly do housekeeping. The interesting thing now is agentic computing. If you say "write me a program", it has to be compiled somewhere, and you won't compile it on a systolic matrix multiplier. You need real CPUs with serious memory systems and networking in the same data center, doing things like compiling and tool calls, though they don't have to sit right next to the accelerator.

1:08:36 · Advice for the next generation

1:08:28Dave Patterson: A 10-year-old in the audience asks: what advice do you have to prepare for my future?

1:08:50Bill Dally: Study math and science. There's no substitute for a really strong foundation in mathematics and the basic sciences.

1:09:06Norm Jouppi: Communication skills are really important too. Learn to write well. I know LLMs can do it for you, but there will be times when you don't have one handy and you still need to communicate well.

1:09:33Dave Patterson: And for young computer architects and engineers entering the field today?

1:09:57Bill Dally: The world is very different. Early in my career I designed a machine at Bell Labs by drawing all the circuits in pencil on vellum, with a technician doing the layout, and it worked on first silicon without a single simulation. Great manual logic design skills now have no value at all. I have three pieces of advice. First, master a vertical, because a lot of the value today is in understanding what to design, not in manual skill to realize it; the AI will help you realize it. Second, be very broad in your understanding of computer technology. Many architects limit themselves: they ask the memory person whether a memory can be built a certain way and hear no. If you understand circuit design broadly, you realize they aren't thinking far enough outside the box. That Bell Labs machine had a 3T DRAM on it that all the circuit people said wouldn't work; I prototyped it and convinced myself it would. Third, make AI your partner, and figure out the uniquely human part of the human-AI partnership in architecture, and practice being good at it. There's no point doing what the AI can do, because it will do it better. But so far you have to recognize when it has screwed up. I can definitely recognize it, maybe because I can do it manually. Or you can get one AI to check the other.

1:12:23Norm Jouppi: Computer architecture has a long history, 70 years or more, with a lot of lessons learned along the way. One of the best things to do is learn those historical lessons, because some apply today and some don't. There's no point reinventing something that didn't work the last seven or eight times people tried it.

Where Indigo landsFurther

Indigo's conclusion

First-hand confirmation from two architects: energy comes first, and the DRAM super-cycle is structural, not cyclical. Dally admits the CUDA moat is getting shallower and Google's TPUs have a structural single-customer edge. On the bubble question they give the strongest sell-side case that this time the demand is real.

What to remember

  1. Energy is the first bottleneck: transmission is hard to permit, so generate on site; gas turbines sold out for about five years; thermal batteries promising; nuclear still loses to gas on levelized cost.
  2. DRAM scaling stopped while demand didn't, so the memory super-cycle is structural; whether memory makers keep prices high for long, Jouppi leaves open.
  3. The TPU's single-customer edge against NVIDIA's multi-customer tax, which Dally says he's a little envious of: a business-model difference, not technical strength.
  4. SemiAnalysis's InferenceMAX has replaced MLPerf as the winning benchmark: it measures exactly tokens per watt.

Claims you can check later

ClaimWhoWhen we will knowHow firm
Energy availability is the first bottleneck for deploying AI chipsDally, Jouppi and the audience pollNowFirst-hand; high credibility
DRAM scaling has stopped while demand hasn't, so memory stays supply-constrainedJouppiOngoingFirst-hand architectural judgment
Software moats like CUDA are significantly weakened by AIDally, JouppiUnder wayFirst-hand; their own admission
2x-a-year gains can last four or five more generations (eight to ten years); numerical precision is nearly used upDallyUntil about 2034First-hand
Quantum computing will have almost no measurable effect on AI training or inferenceDallyLong termFirst-hand
Gas turbines are sold out for about five yearsDallyNowFirst-hand, reported

Back on the long-running theses

confirms

Constraints are moving from algorithms to physics Two top architects confirm this view at the highest level: energy is the first bottleneck, and DRAM scaling has stopped.

adds to

You don't own the model AI filling in the CUDA moat, plus “the system is the moat”: value moves from software libraries to systems, scale and ecosystem.

adds to

AI capex as a single engine “The gold rush is real, unlike the '90s bubble” is the strongest sell-side technical case on the must-spend, real-demand side.

confirms

Perplexity Lily: the real signal is the memory-bandwidth wall Perplexity measured on the usage side that memory bandwidth is the limit; Jouppi says on the supply side that DRAM scaling has stopped.

adds to

SemiAnalysis: commoditization at the capability layer, value at the harness layer Dally says InferenceMAX has replaced MLPerf; SemiAnalysis's influence, confirmed by two architects.

What would change my mind

memory makers return to overbuilding and crashing prices, and the case for structurally constrained DRAM supply falls apart.

Finished. Indigo's take on this piece is in two places: