2:10Dave Patterson: 我来介绍两位从 1980 年代就认识的朋友。那时我在伯克利当助理教授,Bill 在加州理工读博,Norm 在斯坦福。Bill Dally 从加州理工开始就一直在造联网的计算机:先在 MIT,后回到斯坦福当上计算机系主任,2009 年起任 NVIDIA 首席科学家,现在也是高级副总裁。他最有名的贡献包括虫洞路由,以及和 Brian Towles 合写的互连网络那本书;他在斯坦福的流处理项目,是现代 GPU 计算的重要前身。Norm 在斯坦福读研时参与了 MIPS 项目,毕业后在 DEC 西部研究实验室待了大约十年,之后去了惠普实验室。2013 年,他的朋友 Jeff Dean 请他到 Google 为深度学习造硬件。Norm 见过太多 AI 的炒作,很怀疑,但 Jeff 说服了他:深度学习用在什么上都管用。他现在是 Google 的院士兼副总裁,以 TPU 和受害者缓存、预取缓冲这类存储层级的工作闻名。两人的生涯惊人地平行:都在斯坦福师从 John Hennessy,都是美国国家工程院院士,都拿过计算机体系结构的最高奖 Eckert-Mauchly 奖。
06:10 · 这次淘金热有什么不同
6:35Dave Patterson: 今晚的题目是「硅谷淘金热」:密集的创新、投资和竞争。站在 NVIDIA 和 Google 的位置,你们觉得这个 AI 硬件时代的特征是什么?和以往的计算机体系结构时代有什么根本不同?
6:56Bill Dally: 有三个特征让它非常不同。第一,强烈的经济需求。就像 Jeff Dean 说的,AI 用在什么上都管用,于是对更多 token、更多算力、更多 AI 的需求永远填不满。第二,应用变化很快,但和那些动辄几百万行代码的大型超算应用相比,它相对简单。Transformer 是个相对简单的东西,很容易看出要造什么才能让它跑快。第三,大家愿意很快地改,没有积满灰尘的老代码。对比 90 年代初,那时围绕超算也有一轮体系结构淘金热。但没有强烈的经济需求,它是靠 DARPA 的战略计算计划撑起来的,这笔钱让人以为下游有个大市场,其实没有。很多人创了业,比如 Thinking Machines,大多数最后都倒闭了。那时的应用是积满灰尘的老代码,有的接近一百万行。你可以加速看起来最重要的那个内核,但 Amdahl 定律会反咬你,因为另外 99% 的代码没被加速。所以,一个简单的应用、加速几个基本操作就能得到很好的结果,再加上真实的经济需求,让这一轮站得住。这不是 80 年代的 Lisp 机热潮,也不是 90 年代的超算热潮,它在创造真实的价值。
8:49Norm Jouppi: 我一直关注全球经济,现在投进来的钱多得惊人。纽约有一段百年老隧道通往新泽西,一直漏水,差十亿美元修不了;而 Google 宣布明年资本开支 $1050 亿,其他公司也投类似的数目。会有赢家也会有输家,每个人都想跑得越快越好。
13:41Bill Dally: Google 设计 TPU 有个很大的优势:直到最近,它们只有一个客户,就是 Google 自己,所以能想做什么就做什么。NVIDIA 很幸运,在推理和训练两个市场都占 68% 的份额,但随之而来的是我们有很多客户,要让他们都满意。他们会提功能需求,如果是够大的客户,我们至少得认真对待,在硬件里加进东西,让这些不同的人都满意。如果能只为自己造最想要的东西,我们也许能做得更好。所以我有点羡慕。我也很喜欢他们训练 TPU 用的 3D 环面互连。我在 MIT 和 90 年代在 Cray 造过很多 3D 环面网络的超算,能从里面看到我那本书的合著者 Brian Towles 的手笔。
16:03Norm Jouppi: Bill 说得对,有些操作我们都得支持,矩阵乘法、向量运算。TPU 不同的地方在于,第一代是块只做推理的 PCIe 卡,但从第二代开始,它就是按超算来设计的,用的就是 Bill 和 Brian 书里的环面网络。所以我们没经历过那种痛苦:加一个功能,又和别的功能打架。我们从一张白纸、一个超算设计起步,液冷也已经用了八年。
17:40 · 系统才是护城河
17:36Dave Patterson: NVIDIA 和 Google 都是世界级的系统公司。系统工程,也就是互连、内存、供电、散热、封装,是不是已经成了比芯片本身更大的护城河?你们在互连上的做法,怎样解决扩展时的关键瓶颈?
23:00Bill Dally: 产品是整个系统,软件是其中不可分割的一部分。但从很多方面看,深度学习的软件问题,比之前通用 GPU 计算的问题容易。2006 年我们随 G80 推出 CUDA 时,有成千上万的应用能从并行里受益,但代码是串行的,很多是 Fortran 写的,要移植上百万行代码。我们在 CUDA 生态上投入巨大,包括语言本身,以及做 FFT、做矩阵运算的各种库。2010 年我们开始第一个深度学习项目,和斯坦福的 Andrew Ng 合作,后来成了 cuDNN。当时的感受是:哇,这个应用真小,真正要紧的只有几个内核,把它们做快,整个就快了。看过上百万行的气象代码之后,这简直是一股清风。它也变得非常关键:我们把一个模型跑起来,接下来六个月能把性能翻一倍,因为很容易在里面丢性能,所以我们做了各种工具来分析性能、融合内核、调优。软件非常关键。但它对今天的创业公司还算多大的门槛,我不知道,因为现在你只要跟 Claude 说「我要一套软件栈」,然后出去过个周末就行了。
1:09:57Bill Dally: 今天的世界很不一样。我职业生涯早期在贝尔实验室设计过一台机器,用铅笔在描图纸上画出所有电路,由技术员做版图,第一次流片就成功了,没做过一次仿真。现在,手工逻辑设计的本事已经毫无价值了。我有三条建议。第一,精通一个垂直领域,因为今天很多价值在于懂得该设计什么,而不是手工实现它的本事,AI 会帮你实现。第二,对计算机技术要有很宽的理解。很多架构师把自己限制住了:去问做内存的人能不能造出某种内存,对方说不行。如果你对电路设计有宽泛的理解,就会发现对方的思路不够开阔。贝尔实验室那台机器上有一种 3T DRAM,所有做电路的人都说行不通,我自己做了原型,说服自己它能行。第三,让 AI 成为你的搭档,想清楚在人和 AI 的合作里,哪部分是人独有的,并且练好它。去做 AI 能做的事没有意义,它会做得更好。但至少到现在,你还得能看出它什么时候搞砸了。我肯定看得出来,也许是因为我会手工做。或者,你可以让一个 AI 去检查另一个。
The two people best placed to know: energy is the first bottleneck, DRAM has plateaued, and AI is filling in the CUDA moat.
Part 1 of 7 · 2:10
This gold rush rests on real demand
Real economic demand, an application simple enough that accelerating a few primitives works, no legacy code: unlike the '90s supercomputing boom. Both companies are paying the memory makers a lot.
Real economic demand, an application simple enough that accelerating a few primitives works, no legacy code: unlike the '90s supercomputing boom. Both companies are paying the memory makers a lot. Read this part →
GPUs and TPUs haven't converged; the system is the moat
NVIDIA leads on number formats and sparsity but carries a multi-customer tax; TPUs serve only Google and started as supercomputers from a clean sheet; 10,000-GPU systems that just work are experience startups lack. Read this part →
The software moat is filling in; precision is nearly used up
With a spec, having Claude Code rebuild a tuned library isn't hard. Of the 2x-a-year gains only 3x came from process; precision has a turn or two left, and good ideas can last eight to ten years. Read this part →
Chips take three years; models change every three months
Aim ahead of the duck and do the basic operations well; mixture of experts pushes interconnect latency; InferenceMAX has replaced MLPerf as the benchmark people actually watch. Read this part →
Transmission lines are hard to permit, so data centers generate on site. Gas turbines are sold out for about five years, lithium is too expensive, thermal batteries look promising, and gas still beats nuclear on levelized cost. Read this part →
Memory makers can charge several times more for the same part, so they have less reason to build. For decades they overbuilt and prices crashed; whether this time differs is left open. Read this part →
Quantum won't help, burning models into chips is risky, AI will design chips
Quantum is big computation on small data, AI the reverse; with frontier models updated monthly, weights in ROM are too risky; analog loses; verification takes 75% of the labor. Read this part →
Indigo's conclusion
First-hand confirmation from two architects: energy comes first, and the DRAM super-cycle is structural, not cyclical. Dally admits the CUDA moat is getting shallower and Google's TPUs have a structural single-customer edge. On the bubble question they give the strongest sell-side case that this time the demand is real.
How to read this A public conversation between two top chip architects. Both companies sell shovels, so “the gold rush is real and we win either way” suits them. Take the technical calls on energy, memory, moats and benchmarks as hard material, and discount “the gold rush is real” as shovel-seller optimism. A public stage with no adversarial follow-ups, but both make a few rare admissions.
What to remember
Energy is the first bottleneck: transmission is hard to permit, so generate on site; gas turbines sold out for about five years; thermal batteries promising; nuclear still loses to gas on levelized cost.
DRAM scaling stopped while demand didn't, so the memory super-cycle is structural; whether memory makers keep prices high for long, Jouppi leaves open.
The TPU's single-customer edge against NVIDIA's multi-customer tax, which Dally says he's a little envious of: a business-model difference, not technical strength.
SemiAnalysis's InferenceMAX has replaced MLPerf as the winning benchmark: it measures exactly tokens per watt.
What would change my mind
memory makers return to overbuilding and crashing prices, and the case for structurally constrained DRAM supply falls apart.
How to read this
A public conversation between two top chip architects. Both companies sell shovels, so “the gold rush is real and we win either way” suits them. Take the technical calls on energy, memory, moats and benchmarks as hard material, and discount “the gold rush is real” as shovel-seller optimism. A public stage with no adversarial follow-ups, but both make a few rare admissions.
Real economic demand, an application simple enough that accelerating a few primitives works, no legacy code: unlike the '90s supercomputing boom. Both companies are paying the memory makers a lot.
01:24 · Two architects with parallel lives
2:10Dave Patterson: I get to introduce two friends I've known since the 1980s, when I was an assistant professor at Berkeley, Bill was a PhD student at Caltech and Norm was at Stanford. Bill Dally has been building network-connected computers ever since Caltech: at MIT, then back at Stanford, where he became chair of computer science, and since 2009 as chief scientist at NVIDIA, now also a senior vice president. Among his best-known contributions are wormhole routing and the book on interconnection networks he wrote with Brian Towles, and his stream-processing projects at Stanford were an important precursor to modern GPU computing. Norm worked on the MIPS project as a Stanford graduate student, then spent about a decade at Digital's Western Research Lab and went on to HP Labs. In 2013 his friend Jeff Dean asked him to come to Google and build hardware for deep learning. Norm was skeptical given all the past hype about AI, but Jeff convinced him: deep learning worked on everything they tried. He is now a fellow and vice president at Google, known for the TPUs and for memory-hierarchy work like victim caches and prefetch buffers. Their careers run strikingly parallel: both studied under John Hennessy at Stanford, both are in the National Academy of Engineering, both won the Eckert-Mauchly Award, the highest award in computer architecture.
06:10 · What makes this gold rush different
6:35Dave Patterson: The title is "The Silicon Gold Rush": intense innovation, investment and competition. From your seats at NVIDIA and Google, what defines this era for AI hardware, and what makes it fundamentally different from earlier eras of computer architecture?
6:56Bill Dally: Three characteristics make it very different. First, intense economic demand. AI works on everything it's applied to, as Jeff Dean said, and the demand for more tokens, more flops, more AI is insatiable. Second, the application evolves rapidly but is relatively simple compared with the big supercomputing applications that had millions of lines of code. A transformer is a relatively simple thing, so it's easy to see what you have to build to make it go fast. Third, people are willing to evolve very quickly; there are no dusty decks. Contrast the early '90s, when there was a computer architecture gold rush around supercomputing. There was no intense economic demand; it was fueled by DARPA's strategic computing program, whose funding made people think there was a big downstream market, which there wasn't. Lots of people founded companies, like Thinking Machines, and most went bust. The applications were dusty decks with up to close to a million lines of code. You could accelerate what looked like the important kernel, but Amdahl's law would bite you, because the other 99% of the code wasn't accelerated. So the ease of accelerating a few primitives in a simple application, plus real economic demand, makes this one stick. This is not the Lisp machine frenzy of the '80s or the supercomputing frenzy of the '90s. It delivers real value.
8:49Norm Jouppi: I keep an eye on the global economy, and the amount of money being invested is incredible. For a while New York had a hundred-year-old leaky tunnel to New Jersey and was a billion dollars short, and now Google has announced $105 billion of capex next year, with other companies investing similar amounts. There will be winners and losers, and everyone wants to race as fast as possible.
9:34Dave Patterson: In the actual gold rush, who are the prospectors and who sells the pickaxes?
9:42Bill Dally: The prospectors are the people finding verticals. Find the right vertical to apply AI to and you can strike it rich, just like the prospectors. NVIDIA and companies like Google that make the infrastructure are the Leland Stanfords, selling picks and shovels to the prospectors. Whether they strike it rich or not, we're going to make money.
10:12Norm Jouppi: Well, we're both still paying the memory companies a lot of money.
10:19Bill Dally: I think Micron made as much money last quarter as it had in the previous 19 years. It's interesting to see their revenue go up while the number of parts sold stays level.
10:31Norm Jouppi: Exactly. The margins are through the roof.
02
GPUs and TPUs haven't converged; the system is the moat
NVIDIA leads on number formats and sparsity but carries a multi-customer tax; TPUs serve only Google and started as supercomputers from a clean sheet; 10,000-GPU systems that just work are experience startups lack.
10:29 · Convergent evolution? GPUs and TPUs
10:35Dave Patterson: A colleague accuses training accelerators of convergent evolution: a max-size compute die with a large systolic matrix unit, as many HBM stacks as fit around the perimeter, and the fastest SerDes for a custom link. Why is he wrong? Do GPUs and TPUs take divergent approaches, and is there anything you like about the other's design?
11:22Norm Jouppi: Remember that after the Cambrian explosion there was also an implosion. There were all these weird and wonderful creatures, like giant pill bugs with strange appendages, and they didn't survive the extinction event. Same here: many startups with many different ideas, and I think NVIDIA and Google have both succeeded because a lot of those other ideas weren't as capable, and they may go extinct.
12:21Bill Dally: I don't know that we've even converged; I think TPUs and GPUs look very different, except from 20,000 feet. The application drives certain requirements. You have GEMMs, so you need matrix multiply units; you have softmax and norms, so you need vector units that can do transcendental functions; you need a certain memory capacity, memory bandwidth and communication bandwidth. Everything will have that. But there's a lot of nuance in how they're combined. For many years NVIDIA led the way on numerics: we published a paper on vector scaling a few years ago, NVFP4 came out of it, and the MX formats followed. That nuance can give you a 2x advantage over someone doing FP8 at the same accuracy. Similarly with sparsity: Song Han and I wrote a paper in 2015 on how naturally sparse neural networks are, and we put hardware support for sparsity in starting with Ampere. You don't find that in every bug-like creature.
13:41Bill Dally: Google has a big advantage in designing TPUs: until recently they had one customer, Google, so they could decide exactly what they wanted and do it. NVIDIA is fortunate to have a 68% share of both the inference and training markets, but with that comes lots of customers we have to keep happy. They come with feature requests, and if you're a big enough customer we at least have to entertain them and put things in the hardware to make all these different people happy. If we could build exactly what we wanted just for ourselves, we might do better. So I'm a little envious. And I'm fond of their 3D torus interconnect for training TPUs; I built a lot of 3D torus supercomputers at MIT and with Cray in the '90s, and I can see the hand of my co-author Brian Towles in it.
14:58Bill Dally: On the Cambrian explosion: in the late '90s and early 2000s there were no fewer than a hundred graphics chip startups in Silicon Valley. After survival of the fittest there were two left, NVIDIA and ATI, which was acquired by AMD. I think something very similar is the likely outcome of the current explosion.
16:03Norm Jouppi: Bill said it well: there are operations we both have to support, matrix multiply and vector operations. What was different about TPUs is that our first one was a PCIe card that only did inference, but from the second one on they were designed as supercomputers, with the torus network from Bill and Brian's book. So we had no painful steps of adding features that then conflicted with other features. We started with a clean sheet of paper and a supercomputer design, and we've been liquid-cooled for eight years.
17:40 · The system is the moat
17:36Dave Patterson: NVIDIA and Google are both world-class system companies. Has systems engineering, the interconnect, memory, power, cooling and packaging, become an even bigger moat than the silicon itself? How do your interconnect approaches address the critical bottlenecks in scaling?
17:59Bill Dally: The product is the whole system: not just the GPU or the board with GPUs, CPUs and networking gear, but all the hardware, all the software and the configuration that makes it work well together. Starting around our Pascal generation, around 2015, we offered the DGX SuperPOD, before we even had large-scale networking, before we acquired Mellanox. You use this switch, you configure it like this, and you could put together maybe 10,000 GPUs, turn it on, and it would work. Compare the big DOE supercomputers we built, Titan in 2011 or 2012, then Summit and Sierra: after the hardware was working you'd typically spend six months bringing it up because of little things in network configuration. You're delivering an entire system that has to run long training jobs reliably with very high availability. A lot of systems expertise goes into that, and a standardized configuration was critical because we sell into many data centers that do things differently. Where do we differ? Mostly the scale-up network. Google uses the 3D torus with optical circuit switches to route around bad parts. We take a more conventional approach, a Clos network with a proprietary, very low-latency link to avoid the overhead of Ethernet, and a conventional Ethernet scale-out network.
19:57Norm Jouppi: A lot of startups don't have that systems experience. Luiz Barroso and others wrote that the data center is a computer, and many of those lessons aren't obvious until you've felt the pain yourself building these machines. That gives both of us an advantage over startups.
03
The software moat is filling in; precision is nearly used up
With a spec, having Claude Code rebuild a tuned library isn't hard. Of the 2x-a-year gains only 3x came from process; precision has a turn or two left, and good ideas can last eight to ten years.
20:34 · The software moat is filling with sand
20:51Dave Patterson: So there's a lot more to it than chips. How much of NVIDIA's and Google's advantage is hardware versus software, compilers and libraries? Put differently: if you gave a startup the netlist and the layout but not the stack, could it compete?
21:39Norm Jouppi: In designing TPUs we followed my thesis advisor's advice: don't put off until runtime what you can do at compile time. We designed the architecture in a room of ten people, two of them from the compiler team, to make a machine that was easy to compile to. We've kept the same general architecture; memory sizes can grow or shrink, like adding DIMMs to a PC and still running Word. Frameworks keep evolving, so you need good support for them, and some people like writing their own kernels, so you need to support that too.
23:00Bill Dally: The product is the whole system, and software is an integral part of it. But in many ways deep learning is an easier software problem than general GPU computing, which led up to it. When we launched CUDA with G80 in 2006, there were thousands of applications that would benefit from running in parallel but had serial code, many in Fortran, with million-line codes to port. We invested hugely in the CUDA ecosystem, the language and many libraries for FFTs, matrix operations and so on. When we began our first deep learning effort in 2010, a collaboration with Andrew Ng at Stanford that became cuDNN, it was: wow, a tiny application, only a couple of kernels really matter, make those fast and the whole thing is fast. A breath of fresh air after million-line weather codes. It also became critical: we'd get a model running and then double its performance over six months, because there are easy ways to lose performance, so we built tools to profile, fuse kernels and tune. Software is a very critical part. I don't know how much of a barrier it will be to startups these days, because right now you just tell Claude you want a software stack and go away for a weekend.
24:55Dave Patterson: That's my next question. Anyone married to a programmer has heard shouts in the middle of the night: "look what it did." Coding is being done by machines. If there was a software moat, is it about to disappear, or is that a Pollyannaish view?
25:35Bill Dally: There's still an advantage to libraries tuned and honed over the years that are easy to apply. But I think any software moat has been significantly degraded, because if you have the spec of a library, turning Claude Code loose to recreate it is not that difficult.
25:55Norm Jouppi: I agree. The moat is getting shallower. It's filling up with sand; people are starting to be able to walk across.
25:36 · Precision is nearly used up
26:10Dave Patterson: We squeezed a lot of speed in the early years by reducing precision, but we can only go so low. What happens to generation-over-generation gains when we can't shave bits anymore?
26:38Bill Dally: We've pretty much hit 2x per year for the past 14 years, starting with Kepler in 2012, the first generation where we took this seriously as an application, and only 3x of that came from process technology. That's the opposite of the microprocessor heyday of the '90s, when it all came from process. NVIDIA led on numerics, from FP32 in Kepler to shipping NVFP4, and there are some new things in Vera Rubin. There are a couple more turns, but we're getting near the end on numerical precision. There are other axes to keep innovating on: sparsity, circuits, locality, even better models that give more tokens per watt. The low-hanging fruit has been picked, and we need to climb higher up the tree to find those 2x fruits. But we have a bunch of good ideas, and I think they can get us through at least the next four or five generations, and then I can retire.
27:52Dave Patterson: Is that four or five years, or eight to ten?
28:22Norm Jouppi: We had numerics innovations too. When Jeff Dean was doing the initial AI work, they computed in FP32 but stored values by truncating the low-order 16 bits. To a numerical analyst, truncation is fingernails on a chalkboard, but it worked. So when we built TPUs we could run programs in BF16 that had run on the CPUs of the time and get the same results. That was powerful, because we didn't have to spend time debugging at the system level why a model wasn't working. We could verify correct operation, and then we adopted smaller formats like FP8 and FP4.
04
Chips take three years; models change every three months
Aim ahead of the duck and do the basic operations well; mixture of experts pushes interconnect latency; InferenceMAX has replaced MLPerf as the benchmark people actually watch.
29:46 · Chips take three years; models change every three months
29:31Dave Patterson: One of Google's theoretical advantages is having people pushing the state of the art in AI alongside the people building hardware, and I've read that NVIDIA is building that expertise in-house too. How big an advantage is access to those people, versus just using open models?
30:16Norm Jouppi: Leaderboards show some open models performing quite well, so there's a race for proprietary models to stay ahead, and because you can tune a model better for a particular system, GPU or TPU, there will always be some advantage to the proprietary ones. As for hardware people talking to ML people about what's next: to some extent. The problem is that it takes two and a half or three years from an initial idea to a system manufactured in volume, and the ML people come up with a new idea every three months. You can't design a super-specialized machine for a particular model, because it will be different by then. You have to do the basic operations, matrix and vector operations, and do them well.
31:54Bill Dally: We've built our own models for some time and have a lot of internal expertise. But there's a distinction between us and a company like OpenAI, which had a great paper at Hot Chips on its Jalapeño processor, designed pretty much to run one model, so they know their relative mix of matrix ops, vector ops and memory bandwidth. If you have to support all the models out there, there's a very big variety that shifts the provisioning, especially memory bandwidth versus math, particularly with attention. That started with DeepSeek's MLA attention, and now many models use hybrid schemes, alternating three layers of state-space models with one layer of full n-squared attention, or sparse attention that does a quick filter and attends to the top k. That variation drives very different demands on the hardware. As a silicon supplier supporting all models, you have to look across them and decide the right hardware to make everybody as happy as you can without making anybody really unhappy, and which to emphasize. And as Norm said, you have to aim ahead of the duck, because it will be a couple of years before it's out and people will come up with clever ideas we haven't seen yet. That's what good computer architecture is about.
33:50Norm Jouppi: One of the more disruptive recent developments, though proposed about seven years ago, is mixture of experts, because it requires a lot more interconnect bandwidth, and for quick responses to users you really have to push down latency. Training is more a bandwidth problem; low latency is harder to get in computer systems than high bandwidth. That's what drove the interconnect configuration of our latest TPU.
33:44 · The benchmark wars, again
34:41Dave Patterson: At Hot Chips it felt like the battles of the 1980s. Back then RISC companies quoted MIPS, and it was up to you what you ran; everyone else was lying. Now it's tokens per second: running what, on how big a model? The solution then was SPEC, which companies rallied around. MLPerf was an attempt to anticipate this problem, but nobody mentioned it at Hot Chips. Was it less successful than SPEC, and why?
36:14Bill Dally: I like MLPerf; it did what benchmarks are supposed to do, a level playing field to cut through the BS. But it had two issues. First, almost nobody submitted to it, because it was a lot of work. We did every benchmark every generation. But a lot of startups said: if we played by all the rules we wouldn't look that good, so we'll cherry-pick one result for our slides. Second, what everybody really cares about today is tokens per watt or tokens per dollar, and most MLPerf benchmarks don't address that, and even the closest one doesn't address it quite right. What everybody shows, and you did see slides of it at Hot Chips, is SemiAnalysis's InferenceMAX benchmark, because it hits exactly the data people want, and in a very reproducible way: to get a point on that chart you need a repository of code, on GitHub or Hugging Face, that anyone can load and run on that hardware with that model. SemiAnalysis runs it. I think people use InferenceMAX more than MLPerf today.
38:00Norm Jouppi: MLPerf also required a full-time team, and not a small one. It was both expensive and didn't give you the number you wanted.
05
Energy is the first bottleneck
Transmission lines are hard to permit, so data centers generate on site. Gas turbines are sold out for about five years, lithium is too expensive, thermal batteries look promising, and gas still beats nuclear on levelized cost.
38:34 · Impact and responsibility
38:47Dave Patterson: This gold rush has economic, social and even geopolitical implications beyond the technology. What are the most profound impacts, and what responsibility do hardware architects and the leaders of these companies bear?
39:11Bill Dally: I wouldn't call it the silicon gold rush; it's the AI gold rush, and it benefits almost every part of our lives. In medicine, AI started in image analysis and now helps doctors make more accurate diagnoses, and we can have personal health coaches. In education, every student could have a personalized tutor that understands what motivates them and how they learn. In engineering, AI is already automating many tasks. I recently needed to design a piece of hardware: I wrote the spec, set Claude loose, and after I fixed a couple of errors in my spec, it produced a pretty good design. It produced exactly what I asked for, which wasn't the right thing; you have to learn to write a good spec. But it lets engineers in every field move up the ladder. They're no longer doing junior-level calculations; they decide what needs to be built and manage a team of agent minions that carry out the work. Business processes are moving the same way. In entertainment it helps produce wonderful work in movies, games and music.
40:58Bill Dally: On the flip side, any great technology can be used for good or evil. The most obvious are deepfakes, and we need to move rapidly to provenance and authentication, so that unless an image is properly authenticated, you assume it's fake. Cyber gets a lot of attention, although over 80% of successful attacks are human engineering, phishing, which AI of course makes better. AI can be used for pathogen design. And the big risk is that as everybody becomes more productive, the nature of employment will change, with more people needed in some jobs and fewer in others. We need to find ways to ease that transition for the people moving from one to the other.
42:11Norm Jouppi: One of the biggest impacts will be in science. There have been super-impressive results that don't make the popular press. At Google DeepMind they folded all the proteins known to mankind and got a Nobel Prize. Google also uses AI in virtually all its applications, often subtly. Early on, Google Maps would say "drive 1,000 feet and turn right on Tennyson Street"; by integrating the Street View database it now says "turn right at the Shell gas station", which is much more accessible, especially at night when street signs aren't lit. People won't notice many things like that and will just take them for granted.
43:41 · The biggest bottleneck: energy
44:03Dave Patterson: Here's the poll: the biggest bottleneck to continued widespread deployment of AI chips is energy availability, semiconductor manufacturing capacity, cooling, memory capacity and bandwidth, or the cost of the chips? It's still changing, but right now the audience says energy availability. What do you two think?
45:02Norm Jouppi: Energy availability. People talk about gigawatt or 5-gigawatt data centers, and transmission lines are very hard to get permits for; communities don't like big power lines over their houses. So hyperscalers and others are moving to on-site power generation. Some use natural gas, but if you put the data center in the right place you can get most of the power from wind and solar, which is the approach we're trying to take: carbon-free energy, off the grid, so you can build data centers without clobbering everything else.
46:09Bill Dally: You need steady power, and the sun only shines by day and the wind only when it blows, but in general the supply sources have been reasonably balanced by market forces. Land, power and shell are the three things you need for a data center, and they drive the demand. Many people are co-locating natural gas generators, and gas turbines are sold out for about the next five years because of this demand. Some companies are taking engines off airplanes and converting them into generators.
47:01Dave Patterson: But if we go to natural gas, we'll have a giant carbon footprint.
47:06Bill Dally: It's often used to back up renewables, because you can't count on solar and wind and we don't yet have the storage. Lithium batteries are too expensive per kilowatt-hour to ride through the 20-to-100-hour gaps you need to bridge when you don't have sunny days for a while. Thermal batteries look very promising, and people are starting to build data centers with them; they're cheap enough per kilowatt-hour to fill that gap. And a data center is a capital resource: you spent $10 billion building it and want it busy around the clock, so when consumers here are asleep you sell the GPU cycles to consumers on the other side of the world.
48:18Norm Jouppi: It depends on the use. We like quick responses for our users.
48:25Dave Patterson: Nuclear is also carbon-free, and there are small modular reactors; Google has made an announcement in this space. Thoughts on nuclear for data centers?
48:41Bill Dally: We're looking at all technologies; we have someone whose job is to track energy technologies and advise on data centers. Nuclear looks promising, but it has always been expensive in levelized cost per kilowatt-hour compared with natural gas. Natural gas is very tough to beat, and even with carbon sequestration I think it still beats nuclear on levelized cost. I'm a big fan of pumped hydro, but it's not very practical.
06
DRAM scaling has stopped; demand hasn't
Memory makers can charge several times more for the same part, so they have less reason to build. For decades they overbuilt and prices crashed; whether this time differs is left open.
48:53 · Memory: DRAM scaling has stopped
49:35Dave Patterson: Now audience questions. The most popular: what are the biggest challenges to resolving the supply chain barriers for AI hardware, where demand outruns supply?
49:55Bill Dally: The biggest challenge is that it takes a long time to build a fab. Right now the most pain is on the memory side, and the memory manufacturers are loving it, because they can charge many times more for the same part than if it were plentiful. So maybe they're less motivated to build fabs than they otherwise would be. It's always this: you project demand, demand exceeds the projection, and there's a two- or three-year delay to build production capacity.
50:30Norm Jouppi: If I'd been paying more attention, I might have seen this coming, because DRAM density scaling has basically stopped. Computers keep needing more memory, even ignoring AI, so I think the balance has finally tipped, and for quite a while DRAM will be a better business. It used to be one of the worst businesses.
51:05Dave Patterson: You've taught me that SRAM had already plateaued.
51:10Norm Jouppi: And DRAM has basically plateaued, but demand hasn't. It will be interesting for the memory companies, whose history is to overbuild fabs and then watch prices fall; they've had a couple of decades of that. Will they now just say, we're highly profitable, what's wrong with the way things are?
51:38Bill Dally: Fortunately, competition will hopefully fix that: someone highly profitable will want a bit more share and build out a fab, and then the others will too. But fabs are expensive and take a long time, not only to build but to get the yield up.
07
Quantum won't help, burning models into chips is risky, AI will design chips
Quantum is big computation on small data, AI the reverse; with frontier models updated monthly, weights in ROM are too risky; analog loses; verification takes 75% of the labor.
51:52 · Quantum, models burned into chips, the edge, analog
51:56Dave Patterson: How significant will quantum computing be for tonight's topics once it matures?
52:04Bill Dally: Very, very little. Quantum computing is a great technology, but there are two really good quantum algorithms: simulating quantum chemistry, and Shor's algorithm for factoring the product of two large primes and breaking a lot of modern cryptography. There's a third, Grover's algorithm for optimization, but it only gives a quadratic speedup. All of them exploit the fact that quantum computers are large-computation, small-data machines. In people's wildest dreams they'd get a few thousand error-corrected qubits and do tremendous amounts of computation on a few thousand bits. AI is the other way around: a large-data, relatively small-computation problem. So quantum computing won't have a measurable impact on AI training or inference.
53:35Norm Jouppi: Where it will have an effect is communication rather than computation: quantum transmission and cryptography.
54:46Dave Patterson: What do you think of architectures like Etched, or Taalas, that essentially etch an LLM onto the chip?
55:57Bill Dally: Taalas is the extreme: they plan to burn a set of weights into ROM, which is maybe four times as dense as SRAM and way less dense than DRAM. My understanding is that Etched started with that thesis and has become more and more programmable. With the Taalas approach you have to believe some model isn't changing very fast, but the frontier models come out with a new release every month or so, and even among open-source models there's a clever new idea every couple of weeks that would break anything matched too closely to a model, even with programmable weights.
56:33Norm Jouppi: We're still in the early days, and programmability is critically important.
56:39Dave Patterson: This is a completely different universe from CPUs with million-line legacy code and millions of programmers. It's a small amount of code with tremendous pressure to innovate, which makes it easier for hardware people and algorithm people alike. Picking a model and getting a chip out in time is an interesting bet. What about centralized versus distributed compute, and edge hardware?
57:35Bill Dally: We ship a lot of products into autonomous vehicles; every Mercedes above a certain level has an NVIDIA processor and NVIDIA software for its self-driving features. If you can put the compute in the data center, you do, because it's much more economical. But for latency, or because you have to operate reliably through a network partition, or because you're acquiring so much data you can't send it all back, you may need to compute at the edge. The general rule: do it in the data center unless one of those factors stops you.
58:15Norm Jouppi: Google Pixel phones have an edge TPU that handles relatively simple things, like speech recognition. Anything too big goes over the network if it's available.
59:11Dave Patterson: Since low precision achieves comparable AI results, will there ever be room for analog rather than digital systems? You can multiply so elegantly in analog.
59:35Norm Jouppi: When you build real systems for a living, you have to test them. If something is analog and doesn't give the same result each time, it's really hard to test whether the device works. When you probe wafers, you send test sequences and read back bits at tremendous speed, and they have to match one for one.
1:00:27Bill Dally: Even if you could test it, I've seen few cases where analog has an advantage, and we re-evaluate it continuously, because it's a very seductive argument: put the activations in as current, treat a resistor as a conductance, V equals I times G, and you get a free multiply, and a free sum too. But there's no free multiply. Worse, in analog you have no reliable way to store a value for more than a very short time; a capacitor leaks, especially in modern processes. Moving an analog value any distance is very expensive. So people build a small in-memory array, do a matrix multiply of activations times weights, sum the results, and then need an analog-to-digital conversion. The fundamental energy of that conversion limits you, at 8-bit precision, to something like 5 TOPS per watt, whereas we've built digital accelerators at 100 TOPS per watt. If you can stay analog without converting, you might build an attractive device; if you have to convert, you lose.
1:02:57 · Chips designed by AI
1:02:41Dave Patterson: My younger colleagues are excited about using AI for hardware design. Will the next chip design software be radically different from what we've used for 30 years?
1:02:57Bill Dally: I think chip design software will be an LLM you talk to: I want this chip. The EDA companies have a lot of expertise and many of the pieces that need to be plugged together, and the LLM has to know how to run the tools for place-and-route and verification. When I look at where human time goes in designing a chip, it's verification. LLMs are really good at writing tests: make sure this works, cover all the edge cases, and it writes a whole test suite. That's where 75% of our labor goes.
1:03:47Norm Jouppi: I was going to say the same. The results presented at Hot Chips showed around 10% improvements in the design itself, the layout and circuit performance, but the biggest part of the team is design verification, and they need all the help they can get.
1:04:08Bill Dally: On human productivity we're seeing more than that: a doubling for average engineers and 10x for really good ones. In one benchmark study for a particular task, the average came out around 3x. And again, in AI the real gold miners are the people attacking a vertical, and a whole bunch of startups are attacking the EDA vertical.
1:04:39 · How close should the CPU be?
1:05:15Dave Patterson: What about bringing the CPU closer to the GPU, to make room for memory-bound workloads?
1:05:21Bill Dally: Ours are pretty close: they sit on our chip-to-chip NVLink, a single-ended, very high-speed link. What do you get from closeness? First, we provision that link so that all of the Vera CPU's LPDDR5 bandwidth, something like 1.8 terabytes per second, can be routed to either of the two GPUs attached, so the link is never the bottleneck. Second, latency: fetching from that memory is already relatively high-latency, so the little the link adds isn't material. And for control interactions like launching kernels, there's enough latency elsewhere that a faster link isn't critical. We also have many tens of little RISC-V CPUs on the GPU for housekeeping the user never sees. For the CPUs users program, NVLink chip-to-chip is about as close as you want.
1:07:00Norm Jouppi: Long term, the CPU attached to an accelerator will mostly do housekeeping. The interesting thing now is agentic computing. If you say "write me a program", it has to be compiled somewhere, and you won't compile it on a systolic matrix multiplier. You need real CPUs with serious memory systems and networking in the same data center, doing things like compiling and tool calls, though they don't have to sit right next to the accelerator.
1:08:36 · Advice for the next generation
1:08:28Dave Patterson: A 10-year-old in the audience asks: what advice do you have to prepare for my future?
1:08:50Bill Dally: Study math and science. There's no substitute for a really strong foundation in mathematics and the basic sciences.
1:09:06Norm Jouppi: Communication skills are really important too. Learn to write well. I know LLMs can do it for you, but there will be times when you don't have one handy and you still need to communicate well.
1:09:33Dave Patterson: And for young computer architects and engineers entering the field today?
1:09:57Bill Dally: The world is very different. Early in my career I designed a machine at Bell Labs by drawing all the circuits in pencil on vellum, with a technician doing the layout, and it worked on first silicon without a single simulation. Great manual logic design skills now have no value at all. I have three pieces of advice. First, master a vertical, because a lot of the value today is in understanding what to design, not in manual skill to realize it; the AI will help you realize it. Second, be very broad in your understanding of computer technology. Many architects limit themselves: they ask the memory person whether a memory can be built a certain way and hear no. If you understand circuit design broadly, you realize they aren't thinking far enough outside the box. That Bell Labs machine had a 3T DRAM on it that all the circuit people said wouldn't work; I prototyped it and convinced myself it would. Third, make AI your partner, and figure out the uniquely human part of the human-AI partnership in architecture, and practice being good at it. There's no point doing what the AI can do, because it will do it better. But so far you have to recognize when it has screwed up. I can definitely recognize it, maybe because I can do it manually. Or you can get one AI to check the other.
1:12:23Norm Jouppi: Computer architecture has a long history, 70 years or more, with a lot of lessons learned along the way. One of the best things to do is learn those historical lessons, because some apply today and some don't. There's no point reinventing something that didn't work the last seven or eight times people tried it.
Where Indigo landsFurther
Indigo's conclusion
First-hand confirmation from two architects: energy comes first, and the DRAM super-cycle is structural, not cyclical. Dally admits the CUDA moat is getting shallower and Google's TPUs have a structural single-customer edge. On the bubble question they give the strongest sell-side case that this time the demand is real.
What to remember
Energy is the first bottleneck: transmission is hard to permit, so generate on site; gas turbines sold out for about five years; thermal batteries promising; nuclear still loses to gas on levelized cost.
DRAM scaling stopped while demand didn't, so the memory super-cycle is structural; whether memory makers keep prices high for long, Jouppi leaves open.
The TPU's single-customer edge against NVIDIA's multi-customer tax, which Dally says he's a little envious of: a business-model difference, not technical strength.
SemiAnalysis's InferenceMAX has replaced MLPerf as the winning benchmark: it measures exactly tokens per watt.
Claims you can check later
Claim
Who
When we will know
How firm
Energy availability is the first bottleneck for deploying AI chips
Dally, Jouppi and the audience poll
Now
First-hand; high credibility
DRAM scaling has stopped while demand hasn't, so memory stays supply-constrained
Jouppi
Ongoing
First-hand architectural judgment
Software moats like CUDA are significantly weakened by AI
Dally, Jouppi
Under way
First-hand; their own admission
2x-a-year gains can last four or five more generations (eight to ten years); numerical precision is nearly used up
Dally
Until about 2034
First-hand
Quantum computing will have almost no measurable effect on AI training or inference
Dally
Long term
First-hand
Gas turbines are sold out for about five years
Dally
Now
First-hand, reported
Back on the long-running theses
confirms
Constraints are moving from algorithms to physics Two top architects confirm this view at the highest level: energy is the first bottleneck, and DRAM scaling has stopped.
adds to
You don't own the model AI filling in the CUDA moat, plus “the system is the moat”: value moves from software libraries to systems, scale and ecosystem.
adds to
AI capex as a single engine “The gold rush is real, unlike the '90s bubble” is the strongest sell-side technical case on the must-spend, real-demand side.
SemiAnalysis: commoditization at the capability layer, value at the harness layer Dally says InferenceMAX has replaced MLPerf; SemiAnalysis's influence, confirmed by two architects.
What would change my mind
memory makers return to overbuilding and crashing prices, and the case for structurally constrained DRAM supply falls apart.
Finished. Indigo's take on this piece is in two places: