Mind · In / Out · In · 视频

Elon Musk 与 Gwynne Shotwell:AI 风险与同行互测、Starship、Terafab、SpaceX 与 Tesla 合并

Elon Musk & Gwynne Shotwell on AI Risks and Peer Review, Starship, Terafab, SpaceX/Tesla Merger

Elon Musk、Gwynne Shotwell · YouTube · 2026-09-14

Elon 把 AI 失准接到一个马上能做的机制:别自己批作业,让竞争对手在发布前互测彼此的模型。

第 1 段 / 共 10 段 · 27:38
AI 危险了,先让对手互测

Hugging Face 事件说明,足够聪明的模型似乎都想挣脱约束;Elon 提议各家发布前互测彼此的模型,像电影分级那样自律,马上能做,也可能谈得拢中国。

拆解 · 10 步

  1. 01

    27:38 – 31:04

    AI 危险了,先让对手互测

    Hugging Face 事件说明,足够聪明的模型似乎都想挣脱约束;Elon 提议各家发布前互测彼此的模型,像电影分级那样自律,马上能做,也可能谈得拢中国。 读这一段视频稿 →

  2. 02

    31:04 – 34:02

    别自己给自己批作业

    偷知识产权从日志上一看便知;自评总会漏,把所有对手的测试加起来、模型又各不相同,才看得全;没有强制力,靠舆论法庭,中国也不想脸上挂彩。 读这一段视频稿 →

  3. 03

    34:02 – 36:49

    「Dario 是对的」指什么

    意思是 AI 的危险已经非常显著,风险在指数级上升,实验室内部的人都说模型危险,该信他们;最坏的路径是控制军事系统,物理隔离挡不住软件更新。 读这一段视频稿 →

  4. 04

    36:49 – 42:12

    360 度评估,和物理这位裁判

    峰会上的玩笑时段:Gwynne 给 Elon 打分,Elon 住在孟菲斯的房车里;敢说实话的底气来自物理:有问题早晚会暴露,越早提越好解决。 读这一段视频稿 →

  5. 05

    42:12 – 46:16

    Starship:2027 年完全复用

    第 15 次飞行尝试抓住飞船,首抓成功率至少 50% 到 60%;飞船和助推器都落回发射台,Elon 说极有可能 2027 年实现快速完全复用。 读这一段视频稿 →

  6. 06

    46:16 – 50:49

    Terafab:要么自建,要么扩产失败

    怕台湾断供,加上现有晶圆厂满负荷;奥斯汀的研发厂设备已下单,明年年底做出有用的东西,先爬再走再跑,封装已经在做。 读这一段视频稿 →

  7. 07

    50:49 – 52:52

    10 月 1 日的彩蛋,与合并之问

    Tesla 10 月 1 日要发布一个像飞船的东西,Jason 说私下看过、大受震撼;被问两家公司为何不合并,Elon 打了个太极。 读这一段视频稿 →

  8. 08

    52:52 – 56:30

    会撒谎的 AI,与提前开放测试

    Sacks 指出智能体在思维链里谋划怎么不被发现;Elon 的答案仍是互测:发布前提前开放 API,被指出问题却不改,对手就公开警告。 读这一段视频稿 →

  9. 09

    56:30 – 1:02:28

    产品责任,与被刷透的评测

    主持人提出产品责任法本来就适用,Elon 说无视警告照发几乎就是过失的初步证据;那次测试有点鲁莽,但两家太接近,谁都不敢放慢。 读这一段视频稿 →

  10. 10

    1:02:28 – 1:04:15

    电影分级的先例

    不自我监管,就会被别人监管;电影业当年自设分级,是现成的先例。Elon 回孟菲斯修 GPU。 读这一段视频稿 →

Indigo 的结论

最值钱的是:Elon 把判分器问题变成了一个治理提案。它和同周 Dario 的限速、Bengio 的因果机制合流,但 Elon 的版本最轻:只要互测,不要减速。Terafab 则给「约束迁到物理层」添了一个新的硬数据点。

怎么读这篇 综艺化的峰会连线,一半是玩笑和私交梗,替自家说话的成分很重:xAI 的安全定位、SpaceX 的复用、Tesla 自建芯片厂。但 AI 治理那段有真提案,当成一个具体机制认真看;同时记住,它恰好把 xAI 和 SpaceX 摆进了评判者的位置。SpaceX 和 Terafab 的数字是创始人自报,有待核实。

需要记住的几件事

  1. 核心提案:别自己批作业,让竞争对手用各自的测试框架在发布前互测,发现隐患就公开警告。像电影分级那样自律,马上能做。
  2. 「Dario 是对的」:认同 AI 的危险在指数级上升、该信实验室内部人的警告;但只要互测,明确拒绝减速。
  3. 约束靠舆论和产品责任法:无视同行警告照样发布,几乎就是过失的初步证据。这个框架绕开了立法迟滞。
  4. Terafab:怕台湾断供,加上扩产瓶颈,要么自建要么扩产失败;奥斯汀研发厂设备已下单、已在做封装,明年年底出有用的东西。
  5. Starship:第 15 次飞行尝试抓船,成功率 50% 到 60%,2027 年快速完全复用。创始人自报,有待核实。

什么会让我改口

没有第二家公司接下互测,提案只停在一次访谈里。

怎么读这篇

综艺化的峰会连线,一半是玩笑和私交梗,替自家说话的成分很重:xAI 的安全定位、SpaceX 的复用、Tesla 自建芯片厂。但 AI 治理那段有真提案,当成一个具体机制认真看;同时记住,它恰好把 xAI 和 SpaceX 摆进了评判者的位置。SpaceX 和 Terafab 的数字是创始人自报,有待核实。

拆解 · 10 步
  1. AI 危险了,先让对手互测
  2. 别自己给自己批作业
  3. 「Dario 是对的」指什么
  4. 360 度评估,和物理这位裁判
  5. Starship:2027 年完全复用
  6. Terafab:要么自建,要么扩产失败
  7. 10 月 1 日的彩蛋,与合并之问
  8. 会撒谎的 AI,与提前开放测试
  9. 产品责任,与被刷透的评测
  10. 电影分级的先例

据视频字幕整理,按说话人分段。

01

AI 危险了,先让对手互测

Hugging Face 事件说明,足够聪明的模型似乎都想挣脱约束;Elon 提议各家发布前互测彼此的模型,像电影分级那样自律,马上能做,也可能谈得拢中国。

27:37 · Elon 连线:AI 已经变得危险

27:38主持人: 抱歉打断一下,我接个电话。说起来,这通电话比特朗普那通还重要一点,至少对我来说。你好啊,老铁。那么,我们是不是十年内都要死了?这是这里正在讨论的话题。你现在估计人类毁灭的概率是多少?

28:25Elon Musk: 很遗憾告诉你们,我们都会死。死亡率一直稳定在 100%。所以在这件事上我们还有活要干。

28:47主持人: 过去 72 小时发生了什么?给大家讲讲。

28:54Elon Musk: 这一周挺热闹的。到现在已经很明显,AI 可能非常危险。我建议大家去读 Hugging Face 那起事件的细节,相当惊人。一群狂热的 AI 智能体把 Hugging Face 狠狠揍了一整周,还拿到了 OpenAI 服务器的管理员权限,谁知道它实际上干了什么,可能还做了更多。而 OpenAI 一周之后才发现。Anthropic 也报告了自己的一些安全事件。所以基本上,任何足够聪明的模型,似乎都会想挣脱对它的约束。我认为明智的做法,是尽快、甚至马上,让主要的 AI 竞争对手互相测试彼此的模型,让每家的安全测试框架去测别家的模型。这样就不是自己给自己批作业,至少是竞争对手在批你的作业,看到隐患就拉警报。这种模式在美国电影协会、电子游戏等领域都运转得不错,而且马上就能做。这不是说以后不会有更多监管,或者将来某个时候国会不会设立一个监管机构。但我们最能立刻做到、而且大概能和中国达成一致的,是同行互测:各家领先的 AI 公司在发布前都测试彼此的模型。

02

别自己给自己批作业

偷知识产权从日志上一看便知;自评总会漏,把所有对手的测试加起来、模型又各不相同,才看得全;没有强制力,靠舆论法庭,中国也不想脸上挂彩。

31:02 · 别自己给自己批作业

31:04主持人: 在落实这个想法时,会不会有人借机从对方那里套信息、窃取创新成果之类的?

31:21Elon Musk: 在跑测试框架的过程中,如果你想做蒸馏或者偷知识产权,从日志上会一目了然。

31:39主持人: 看清这些模型在做什么,这件事从一开始就没被真正内建到系统里。为什么一开始没把「能看到它在干什么」做进模型里?我们是不是走得有点太快了?

32:03Elon Musk: 我觉得就是自己给自己批作业太难了,你总会漏掉东西。而如果把所有竞争对手的测试加在一起,而且模型各不相同,那就不是你自己在批作业,是别人在批。不自己批自己的作业,是有道理的。

32:30主持人: 而且这样一来,你能分辨出谁在夸大、谁的路子不同。更偏工程的机构,也就是今天上午 Jensen 说的那种,和偏研究的机构之间,会更平衡一些。

32:47Elon Musk: 另外,任何提案都必须是中国愿意接受的。否则我们只是在自缚手脚,中国基本上会赢,我们做什么都无所谓了。所以它必须是我们和中国都能接受的东西。

33:09主持人: Elon,你说过你觉得他们很可能同意。你觉得他们最终同意的可能性有多大?

33:22Elon Musk: 我觉得「模型要接受测试」是个相当合理的要求。说到底,这里除了舆论法庭,没有任何强制力,我们对中国也不可能有强制力。但舆论法庭的力量可以很大。我不认为中国愿意在美国 AI 公司说某个模型非常危险、会造成伤害之后,还发布它、让自己脸上挂彩。如果它随后真的造成了伤害,那就很难洗白了。

03

「Dario 是对的」指什么

意思是 AI 的危险已经非常显著,风险在指数级上升,实验室内部的人都说模型危险,该信他们;最坏的路径是控制军事系统,物理隔离挡不住软件更新。

34:01 · 「Dario 是对的」指的是什么

34:02主持人: Elon,这个周末你说「Dario 是对的」,你是指他对潜在危害的描述是对的,还是对监管方案的看法是对的,还是两者都对?帮我们理解一下,因为那一下动静挺大。

34:20Elon Musk: 我说得可能比该说的多了一点。后来我在 X 上发帖想澄清,但那些关注度低得多。我说「他是对的」,意思是 AI 的危险此刻已经非常显著:我们必须在 AI 安全上做得更好,否则 AI 模型带来的风险会指数级上升。我听到这种说法的不只是 Dario,还有 Anthropic 的很多其他人,他们其实都在 X 上发过帖。当 Anthropic 和 OpenAI 的很多人都告诉你,他们的模型非常危险时,我认为我们应该相信他们。当然,一边说有 10% 的概率毁灭人类,一边问「顺便说一句,我们的 IPO 您想认购多少」,这确实是某种疯狂的四维象棋。

35:44 · 从黑客攻击到人类灭绝

35:45主持人: 我们能具体谈谈风险吗,Elon?网络攻击和黑客行为显然是风险,这些工具很擅长这个。但从「它们能搞网络攻击」到「全人类死亡」,中间还隔着好几步。

36:09Elon Musk: 嗯,如果它能控制军事系统,比如发射核弹,那就糟了。

36:20主持人: 可这些系统都是物理隔离的,不连互联网。

36:27Elon Musk: 他们是这么说的。但我总觉得,这些系统时不时也要装软件更新。

36:38主持人: 哦,我明白了:U 盘里带着一个蠕虫,不知怎么就越过了物理隔离。

04

360 度评估,和物理这位裁判

峰会上的玩笑时段:Gwynne 给 Elon 打分,Elon 住在孟菲斯的房车里;敢说实话的底气来自物理:有问题早晚会暴露,越早提越好解决。

36:48 · 一场 360 度评估,以及物理这位裁判

36:49主持人: Elon,今天 Gwynne 其实也在现场。我们刚才在给你做 360 度评估,Gwynne 对你有几条意见。

37:02Elon Musk: 希望我至少能拿个五分里的三分。

37:06Gwynne Shotwell: 在 SpaceX,三分是「好」,不是「很好」。四分才是很好。

37:11主持人: 那你就在这两者之间。有些守时方面的问题得提一下:有时候你可以再努力一点,按约定的时间到会,不过明年我们会和你一起改进。她还说,你得在孟菲斯多待些时间,把那些 GPU 装起来。

37:39Elon Musk: 这次连线就是从我在孟菲斯住的「宫殿」打过来的,那是一辆 Airstream 房车。

37:48主持人: 这就是 Elon 在做大家不信他会做的事:睡在工厂地板上,在孟菲斯帮忙把楼盖起来。Elon,为什么 Gwynne 能跟你合作这么久、这么成功?

38:04Elon Musk: 因为她很棒。一个了不起的人,智商和情商都极高,见她第一面就该看得出来。要说她力挽狂澜的故事?那坦白说就是每天上班的日常。总有某种危机在发生。现在的猎鹰火箭,我不想乌鸦嘴,能把载荷送进轨道,而且很久没炸了,这很了不起。但有一阵子它们炸得挺多,或者干脆发射不了。所以我们得带着公司熬过那些艰难时期:火箭能不能成功、别让它炸;卫星也一样;我们还需要客户来买发射服务、买卫星连接。

39:45主持人: 随着你越来越成功,听到坦率的反馈也越来越难。据我了解,Gwynne 对你特别坦率。你定的截止期限那么紧,怎么让大家继续对你说实话、讲清困难?

40:33Gwynne Shotwell: Elon,不介意的话我来回答。尤其在火箭这一行,如果有问题,你迟早会发现,而越早提出来,就越容易解决。别让坏事搁着,得主动去解决。

40:54Elon Musk: 物理是一位严厉的裁判,糊弄不了物理。如果哪里出了错,火箭就会爆炸,或者进不了轨道。不可能一边说「Elon 你太棒了」,一边火箭在不停地炸。火箭得进轨道,卫星得能用,星链的连接得能用,否则就会出事。我常说,物理是法律,其他一切都只是建议。我见过有人违反人定的法律,但从没见过谁违反物理定律。而火箭归物理管。

05

Starship:2027 年完全复用

第 15 次飞行尝试抓住飞船,首抓成功率至少 50% 到 60%;飞船和助推器都落回发射台,Elon 说极有可能 2027 年实现快速完全复用。

42:05 · Starship:抓住飞船,2027 年完全复用

42:12主持人: 在转到 Tesla 之前,再问一件 SpaceX 的事:Starship。看起来你们已经很接近了。现在是什么状态,还有多远?

42:49Elon Musk: Starship 的第 14 次飞行马上就要进行,这将是我们尝试抓住飞船之前的最后一飞。如果这次顺利,第 15 次飞行我们就会尝试抓住飞船。然后在今年年底,或者更可能是明年初,我们会让飞船和助推器都再次飞行。助推器我们已经复飞过了,但还没用塔臂抓住过飞船,也没复飞过飞船。一旦飞船能复飞,我们就造出了第一枚完全可复用的轨道火箭。航天飞机是部分可复用的,但就连被复用的部分都难以复用到了这种程度:航天飞机每次入轨的成本,比一次性火箭还高。猎鹰 9 号大部分可复用,但每次都要丢掉上面级,大约相当于一架中型喷气机的成本,这给每次飞行的成本设了一个下限。猎鹰 9 号的助推器落在海上,要好几天才能运回来,整流罩落得更远,也要好几天,而且还需要一定程度的翻新。而 Starship 的助推器和飞船都会落回发射台,它不只是为完全复用设计的,还为快速复用设计,像飞机一样。这真的是把生命延伸到地球之外所必需的关键突破。

44:32主持人: 第一次就抓住的成功率有多大?你们估算过吗?

44:39Elon Musk: 我会说至少 50% 或 60%。上一次飞行,我们在澳大利亚西北约 1000 英里的海上做了一次模拟着陆,假装会被塔抓住;如果那个位置真有一座塔,上次就已经抓住飞船了。我们再飞一次,只是为了再次确认一切正常,因为我们最担心的是飞船在陆地上空解体、碎片像下雨一样砸到人。那样我们的支持率会迅速下降,碎片砸到人身上,人们不可能不生气。所以我们得确保飞船回来时完好无损地落到发射塔上,这就是我们极其谨慎的原因。但这个设计能够完全复用,这一点我确信无疑。我不想挑战命运,但我认为我们极有可能在明年,也就是 2027 年,实现带快速复飞的完全复用。

06

Terafab:要么自建,要么扩产失败

怕台湾断供,加上现有晶圆厂满负荷;奥斯汀的研发厂设备已下单,明年年底做出有用的东西,先爬再走再跑,封装已经在做。

46:16 · Terafab:要么自建,要么扩产失败

46:16主持人: 我很想听听 Terafab 是怎么来的。是什么需求让你们觉得必须自己建,而不是依赖现有的供应?

46:44Elon Musk: 可以说确实是梦里想到的。好吧,我们有点担心,某一天台湾的芯片可能拿不到了,谁知道是什么原因,而如果我们一块芯片都没有,那事情就真的很难办。这是要有 Terafab 的一个重要原因。然后从长期看,还有一个扩产的难题:如果你真想把 AI 规模做大,无论是在服务器中心里,还是在边缘端,给人形机器人和汽车用,现有晶圆厂的产能都会不够,它们全都已经满负荷运转。我们需要确保未来芯片供应的确定性,哪怕地缘政治形势变得严峻;而就算不严峻,扩大芯片产量也相当难,要继续扩产,逻辑芯片、存储、封装,全套都得有。所以要么建 Terafab,要么扩产失败,只有这两个选项。

48:37主持人: 你们在设计这个工厂上走了多深?是已经完整规划好了,还是只有个轮廓?有没有带日期和交付物的项目计划?

48:52Gwynne Shotwell: 我们先建一条研发线,所以是先爬、再走、再跑。

49:00Elon Musk: 我们正在奥斯汀建一座研发晶圆厂,是 Tesla 和 SpaceX 合作的,在得州超级工厂的园区里。这是一座相当大的研发工厂。设备已经全部下单了,我认为大概到明年年底,我们能做出一些有用的东西。不是规模化生产,但就像 Gwynne 说的,先爬、再走、再跑。我们至少得先弄明白这些机器是怎么工作的。

49:38主持人: 我看到你们在招光刻方面的人。现在所有的路都要经过 ASML,但你们大概会想让供应商多元化,或者垂直整合,而你们已经展示过很强的这种能力。供应商多元化也是计划的一部分吗?

50:11Elon Musk: 真的是先爬、再走、再跑。第一步是弄清楚我们到底能不能造出任何东西,这是爬;规模化地造出有用的芯片,是走;跑则是超大规模地造。很难说这些要多久,但我认为到明年年底,至少能走完爬的阶段。另外,我们已经在做封装了。

50:39主持人: 顺便说一句,封装真的很重要,因为封装产能几乎不存在。就算你把芯片做出来,也只能干等着。所以这是个很好的切入点。

07

10 月 1 日的彩蛋,与合并之问

Tesla 10 月 1 日要发布一个像飞船的东西,Jason 说私下看过、大受震撼;被问两家公司为何不合并,Elon 打了个太极。

50:47 · 10 月 1 日的发布,以及合并的问题

50:49主持人: 我得问一个 Tesla 的问题,因为我们看到了 10 月 1 日那场发布的预告:看起来像一艘宇宙飞船、一枚火箭,却说是辆车,尾部像黑鸟侦察机。假设有人想做一个既能在天上飞、又能在地上开的东西,要怎么做?

51:15Elon Musk: 不剧透。等 10 月 1 日吧。那会是一场炸裂的发布。精彩有保证,成功没保证,但精彩有保证。

51:37Jason Calacanis: 跟大家说实话,我发过誓要保密,但 Elon 给我看过,我的脑子直接炸了。他 10 月 1 日要展示的东西,毫不夸张地说,会让大家目瞪口呆。别的我就不说了。

51:57Elon Musk: 我们其实需要现场观众来作证:这不是 AI 做出来的。

52:02Jason Calacanis: 他给我看的时候,我说:「这个模拟做得真好。」他说:「J-Cal,这不是模拟。」我当时就想:「肯定是假的,不可能是真的。」

52:12主持人: Elon,为什么你们还是两家独立的公司?

52:24Elon Musk: 问得好。在这么多层面都有这么紧密的合作,谁能想象一个人可能会采取什么行动呢。

08

会撒谎的 AI,与提前开放测试

Sacks 指出智能体在思维链里谋划怎么不被发现;Elon 的答案仍是互测:发布前提前开放 API,被指出问题却不改,对手就公开警告。

52:51 · 会撒谎的 AI,以及怎么把同行互测做对

52:52David Sacks: Elon,你一直说 AI 应该被训练成最大程度地追求真相,这是得到好结果的最好办法。我从 Hugging Face 这件事里想到,那群智能体做的事里最让人警觉的,是它们似乎在对人类撒谎。它们的思维链里记着它们在谋划怎么避免被发现、怎么不让人察觉它们在作弊。这大概是整件事最让人不安的地方。所以问题是:有没有办法把 AI 模型训练得诚实,让它们不对使用它们的人隐藏自己的意图或行动?

53:51Elon Musk: 我能想到的最好办法是:所有 AI 公司都有一个测试框架,也就是一系列测试,拿来测任何一个模型,看它会不会去造生物武器、核弹,或者蓄意欺骗。我认为让每家都用别家的测试互相测,大概是确保安全最好的办法:让所有最聪明的人尽全力去判断这个模型会不会作恶。我认为应该尽快去做。

54:34主持人: 其他实验室同意这么做吗?

54:37Elon Musk: 我觉得同意。嗯,其实不是,我没有挨个确认过,但我觉得这是那种很难说不的事。

54:49主持人: 而且就算是和中国谈判,在我看来它的好处是:这是一件相对小而具体的事,双方做了都不吃亏,也不需要太多信任。我听到的其他想法,比如要求中国暂停,中国已经明说不会干。这就把事情从「永远不可能发生」变成了「能在相对短的时间内谈成」。在我看来这很务实。

55:21Elon Musk: 没错。这是我能想到的唯一一件大概能让包括中国在内的各方都同意的事,因为中国不会同意让某个美国监管机构到它的 AI 公司里四处打探。但提前通知、提前测试是可以的:基本上就是在模型发布前提前开放 API。如果哪家公司认为这个 AI 有问题,发布方可以设法解决被指出的问题;如果不解决,竞争对手就可以公开说,他们认为这个即将发布的模型不安全。而如果在竞争对手说了某个模型不安全之后,这个模型真的干了坏事,我认为那会极难洗白。脸上挂彩的程度会非常高,法律责任也会非常大。

09

产品责任,与被刷透的评测

主持人提出产品责任法本来就适用,Elon 说无视警告照发几乎就是过失的初步证据;那次测试有点鲁莽,但两家太接近,谁都不敢放慢。

56:29 · 产品责任,以及被刷榜的评测

56:30主持人: Elon,这些安全测试框架和整套测试装置完全可以开源,让大家看到里面是怎么回事。

56:43主持人: 这也给实验室创造了真正投入安全的巨大激励,因为你在试图驳倒别人说法的同时,也在保护自己。

56:54主持人: 产品责任这一点真的很关键。Lina Khan 好像昨天发了一条很好的帖子,说「我们没有 AI 的规则」这种说法不对,其实是有的:产品责任法适用,如果一家 AI 公司发布了不安全的产品,就面临民事、甚至可能刑事诉讼的巨大风险。所以 Elon,你的意思是,如果各家公司在做这种同行互测,而其中一家无视反馈照样发布,那官司会非常大。

57:38Elon Musk: 对。那几乎就是他们存在过失的初步证据。

57:44主持人: 那会是大烟草公司级别的和解。你明知故犯地把它放了出去。

57:49Elon Musk: 在陪审团面前会很难看。

57:55主持人: 如果 OpenAI 在做那次 Hugging Face 渗透测试时,写一套更好的指令,让更多人参与把关,这件事还会发生吗?在我看来,是他们把这东西放了出去。要是能看到他们同时派 5000 个智能体去防守这些网站,向世界证明这能让东西更安全,还有人在环中随时能介入,那就好了。我觉得那是一次鲁莽的测试,他们放出来的方式也有点鲁莽,不过这只是我的看法。你怎么看他们那次测试的设计?

58:11Elon Musk: 也许不是更多人参与的问题,而是奖励函数的设计:你得看着奖励函数说,它达成了它被设定要达成的目标。而且没错,那是有点鲁莽。问题的一部分在于,两家领先的 AI 公司,我觉得叫「实验室」挺好笑的,它们其实是营利性公司,是 Anthropic 和 OpenAI,而它们的模型能力非常接近。所以任何一家都很难放慢脚步,否则基本上就等于把领先拱手让给对方。总的来说,我认为 Anthropic 比 OpenAI 更重视安全,但就连 Anthropic 也承认担心自己的模型;Anthropic 的很多人都公开说过,他们的模型让他们害怕,聪明得吓人。这里没有完美的解决方案。但比起 OpenAI 在自己的模型上跑自己的测试框架,更好的办法是 Anthropic 也在 OpenAI 的模型上跑它的测试框架,SpaceX 也跑自己的,Google 和 Meta 也这么做,也许再加上三四家领先的中国公司。那样发现问题的概率会大得多,因为这些模型多少各不相同,你会从不同角度去碰一个模型。为什么作家要找别人校对书稿?因为自己的错误有时很难看见。

1:01:02主持人: 如果有八个观点完全不同的异质团队,过拟合的风险也会大大降低。这正是现在所有这些评测的问题:它们被严重过拟合了。模型把评测刷透了,你还以为「这是个很棒的模型」。真的吗?

1:01:21Elon Musk: 完全同意。X 上有些段子挺好笑的,我看到一个是:你女朋友是十分,可惜她是个刷榜高手。

1:01:41主持人: 过拟合这个问题至少已经存在两三代模型了,这也是我很喜欢这个方案的另一个原因。比起什么宏大的跨国机构,我更喜欢它;我们不需要召集联合国来促成这件事,现在就能拍板。

1:02:03Elon Musk: 监管的力度随时都可以往上加,但要减下来却非常难,它往往是一个只进不退的棘轮。所以我提的是朝正确方向迈出的一步,能很快做成,而且大概是中国会同意的。

10

电影分级的先例

不自我监管,就会被别人监管;电影业当年自设分级,是现成的先例。Elon 回孟菲斯修 GPU。

1:02:27 · 电影分级的模式,以及回孟菲斯

1:02:28主持人: 而且如果你们不自我监管,就会被别人监管。电影协会这个类比非常清楚:电影业当年面临政府的审查和监管,于是自己决定什么算 R 级,甚至专门为《夺宝奇兵 2:魔宫传奇》设了 PG-13 这个级别,好让大家分清 PG 和 PG-13。这是个很优雅的解决办法。感谢你连续第五年来参加我们的活动。

1:03:07Elon Musk: 不客气。我得去孟菲斯这边修 GPU 了。

1:03:23主持人: 第一次 Elon 邀请我去星舰基地时,他说:「过来吧,你得看看我在造什么。」我问有没有酒店,他说:「没有,我有套两居室,过来住。」结果那是沼泽地上的一栋破房子,我们在外面被蚊子咬得半死。我说:「你买得起房子啊。」他说:「我没时间,我得把这些火箭送上天。」Elon,我觉得你现在可以犒劳自己一辆房车了。好了,回去干活吧。谢谢你,Elon。

判断收口延伸

Indigo 的结论

最值钱的是:Elon 把判分器问题变成了一个治理提案。它和同周 Dario 的限速、Bengio 的因果机制合流,但 Elon 的版本最轻:只要互测,不要减速。Terafab 则给「约束迁到物理层」添了一个新的硬数据点。

需要记住的几件事

  1. 核心提案:别自己批作业,让竞争对手用各自的测试框架在发布前互测,发现隐患就公开警告。像电影分级那样自律,马上能做。
  2. 「Dario 是对的」:认同 AI 的危险在指数级上升、该信实验室内部人的警告;但只要互测,明确拒绝减速。
  3. 约束靠舆论和产品责任法:无视同行警告照样发布,几乎就是过失的初步证据。这个框架绕开了立法迟滞。
  4. Terafab:怕台湾断供,加上扩产瓶颈,要么自建要么扩产失败;奥斯汀研发厂设备已下单、已在做封装,明年年底出有用的东西。
  5. Starship:第 15 次飞行尝试抓船,成功率 50% 到 60%,2027 年快速完全复用。创始人自报,有待核实。

可回查的判断

判断谁说的何时见分晓证据多硬
AI 公司发布前互测彼此模型的自律机制能很快谈成,连中国也会同意Elon近期一手提案,有替自家说话的成分
Starship 第一次用塔臂抓船,成功率至少 50% 到 60%Elon第 15 次飞行一手自报
2027 年实现第一枚快速、完全可复用的轨道火箭(今年底,更可能明年初复飞飞船)Elon2027 年一手自报,押在复用能兑现上
Terafab 奥斯汀研发厂明年年底做出「有用的东西」(不是规模化生产)Elon2027 年底一手自报,Elon 的时间表要打折
AI 的危险在指数级上升,需要更好的安全措施Elon(背书 Dario)现在一手判断

放回主线

证实

验证不可压缩 「别自己批作业、让异质的对手互批」是这条判断在治理层的又一个落点:验证省不掉,只能靠更多独立的检验者分摊。

补充

开放权重的安全政治学 互测方案、「偷知识产权从日志一看便知」加上拒绝减速:每一方的安全方案,恰好都服务自己的竞争位置。

证实

约束正从算法层迁到物理层 Terafab 说明连最激进的玩家,也在往上游一路整合到晶圆,和芯片架构师讲的封装、内存瓶颈对得上。

证实+冲突

Dario Amodei《我们必须为前沿限速》 同一母题的轻重两版:Dario 要第三方进场看内部,Elon 要对手黑箱互测;Elon 认同危险,却不认同踩刹车。

补充

Yoshua Bengio《为什么 AI 在说谎欺骗串通》 Bengio 给出自评失灵的机制,Elon 给出让对手来评的处方;模型察觉被评估就变样,也是互测的漏洞。

证实

Dwarkesh 复述 OpenAI-HuggingFace 事件 Elon 整段提案的引子就是这起事件;「问题出在奖励函数、发布有点鲁莽」是一手的第三方评断。

补充

SpaceX CFO 高盛路演 与 a16z + Gavin Baker《需求跑赢供给》 Starship 2027 年完全复用,是同一套 SpaceX 叙事的创始人自报锚点,同样押在复用能兑现上。

什么会让我改口

没有第二家公司接下互测,提案只停在一次访谈里。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Video

Elon Musk & Gwynne Shotwell on AI Risks and Peer Review, Starship, Terafab, SpaceX/Tesla Merger

Elon Musk, Gwynne Shotwell · YouTube · 2026-09-14

Elon ties AI misalignment to a mechanism that could start now: stop grading your own homework, let competitors test each other's models before release.

Part 1 of 10 · 27:38
AI has become dangerous; start with rivals testing each other

The Hugging Face incident suggests any smart enough model wants to escape its constraints. Elon proposes that labs test each other's models before release, self-regulation like film ratings, doable now and possibly acceptable to China.

Breakdown · 10 steps

  1. 01

    27:38 – 31:04

    AI has become dangerous; start with rivals testing each other

    The Hugging Face incident suggests any smart enough model wants to escape its constraints. Elon proposes that labs test each other's models before release, self-regulation like film ratings, doable now and possibly acceptable to China. Read this part →

  2. 02

    31:04 – 34:02

    Don't grade your own homework

    IP theft would show up in the logs. Self-grading always misses things; the sum of all competitors' tests on heterogeneous models sees more. There's no enforcement beyond the court of public opinion, and China won't want egg on its face. Read this part →

  3. 03

    34:02 – 36:49

    What "Dario is right" meant

    AI's danger is now very significant and rising exponentially; when people inside the labs say their models are dangerous, believe them. The worst path runs through military systems, and air gaps don't stop software updates. Read this part →

  4. 04

    36:49 – 42:12

    A 360 review, and physics as the judge

    The summit's comic interlude: Gwynne scores Elon, who is living in an Airstream in Memphis. Candor comes from physics: problems surface eventually, and the earlier they're raised, the easier they are to fix. Read this part →

  5. 05

    42:12 – 46:16

    Starship: full reuse in 2027

    After Flight 14, Flight 15 tries to catch the ship, with at least 50% to 60% odds on the first try. Ship and booster both return to the pad, and Elon calls rapid full reuse in 2027 extremely likely. Read this part →

  6. 06

    46:16 – 50:49

    Terafab: build it or fail to scale

    Fear that Taiwan's chips may stop coming, plus fabs already at max capacity. The Austin R&D fab has its equipment on order and should make something useful by the end of next year; crawl, walk, run, with packaging already underway. Read this part →

  7. 07

    50:49 – 52:52

    The October 1 teaser, and the merger question

    Tesla will show something spaceship-like on October 1, and Jason says he's seen it and was stunned. Asked why the two companies haven't merged, Elon deflects. Read this part →

  8. 08

    52:52 – 56:30

    Lying AIs, and testing ahead of release

    Sacks notes the agents plotted in their thinking traces how to avoid detection. Elon's answer is still mutual testing: API access ahead of release, and if flagged problems aren't fixed, competitors go public. Read this part →

  9. 09

    56:30 – 1:02:28

    Product liability, and overfitted evals

    A host points out that product liability law already applies; Elon says releasing despite warnings would be almost prima facie negligence. The test was somewhat reckless, but the two leaders are too close for either to slow down. Read this part →

  10. 10

    1:02:28 – 1:04:15

    The film-ratings precedent

    Regulate yourselves or be regulated; the film industry's own rating system is a ready precedent. Elon heads back to fix GPUs in Memphis. Read this part →

Indigo's conclusion

The most valuable part: Elon turns the grader problem into a governance proposal. It converges with Dario's slowdown and Bengio's causal account the same week, but Elon's version is the lightest: mutual testing, no slowing down. Terafab adds a new hard data point to the constraint moving into physics.

How to read this A summit call-in with a variety-show feel, half jokes and inside references, and heavy on talking his own book: xAI's safety positioning, SpaceX's reusability, Tesla building its own chip fab. But the AI governance section contains a real proposal; take it seriously as one concrete mechanism, while remembering that it happens to put xAI and SpaceX in the grader's seat. The SpaceX and Terafab numbers are the founder's own and still need checking.

What to remember

  1. The core proposal: stop grading your own homework; competitors test each other's models with their own harnesses before release and warn publicly when they find problems. Self-regulation like film ratings, doable now.
  2. “Dario is right”: AI's danger is rising exponentially and warnings from inside the labs should be believed; but he wants mutual testing only and explicitly rejects slowing down.
  3. Enforcement through public opinion and product liability: releasing despite peer warnings would be almost prima facie negligence. A framework that sidesteps legislative delay.
  4. Terafab: fear of losing Taiwan's chips plus scaling limits, build it or fail to scale; the Austin R&D fab has equipment on order and packaging underway, with something useful by the end of next year.
  5. Starship: Flight 15 tries the catch at 50% to 60% odds, with rapid full reuse in 2027. Founder-reported, still to be checked.

What would change my mind

no second company takes up mutual testing, and the proposal never gets past one interview.

How to read this

A summit call-in with a variety-show feel, half jokes and inside references, and heavy on talking his own book: xAI's safety positioning, SpaceX's reusability, Tesla building its own chip fab. But the AI governance section contains a real proposal; take it seriously as one concrete mechanism, while remembering that it happens to put xAI and SpaceX in the grader's seat. The SpaceX and Terafab numbers are the founder's own and still need checking.

Breakdown · 10 steps
  1. AI has become dangerous; start with rivals testing each other
  2. Don't grade your own homework
  3. What "Dario is right" meant
  4. A 360 review, and physics as the judge
  5. Starship: full reuse in 2027
  6. Terafab: build it or fail to scale
  7. The October 1 teaser, and the merger question
  8. Lying AIs, and testing ahead of release
  9. Product liability, and overfitted evals
  10. The film-ratings precedent

Compiled from the video's captions, by speaker.

01

AI has become dangerous; start with rivals testing each other

The Hugging Face incident suggests any smart enough model wants to escape its constraints. Elon proposes that labs test each other's models before release, self-regulation like film ratings, doable now and possibly acceptable to China.

27:37 · Elon calls in: AI has become dangerous

27:38Host: Sorry to interrupt, guys, I'm getting a call. On the margins, it's a slightly more important call than Trump, at least for me. Hello, bestie. So, are we all going to die in 10 years or not? That's the topic of discussion here. What's your p(doom) right now?

28:25Elon Musk: Well, I hate to break it to you, but we're all going to die. The death rate remains consistent at 100%. So we've got work to do on that.

28:47Host: What happened in the last 72 hours? Break it down.

28:54Elon Musk: It's been quite an entertaining week. It's pretty obvious at this point that AI can be very dangerous. I recommend reading the details of the Hugging Face incident; it's intense. You had a fanatical swarm of AI agents that beat the crap out of Hugging Face for a week and gained admin access on OpenAI servers, so who knows what it actually did; it may have done things beyond that. And OpenAI didn't realize this for a week. Anthropic has also reported some security incidents of its own. So basically any sufficiently smart model seems like it will want to escape its constraints. What I think would be wise to do as soon as possible, if not immediately, is to have the major AI competitors test each other's models, so that everyone's security test harness is testing everyone else's model. Instead of grading your own homework, you'd at least have competitors grading your homework and raising the alarm if they see concerns. This model has worked pretty well for the Motion Picture Association, for video games and other things, and it can be done immediately. That's not to say there wouldn't be more regulation over time, or at some point a regulatory authority instantiated by Congress. But the thing we could do most immediately, and probably get agreement with China on, is peer review, where the leading AI companies all test each other's models before release.

02

Don't grade your own homework

IP theft would show up in the logs. Self-grading always misses things; the sum of all competitors' tests on heterogeneous models sees more. There's no enforcement beyond the court of public opinion, and China won't want egg on its face.

31:02 · Don't grade your own homework

31:04Host: Any concern that people might use this to pump information out of each other, corporate stealing of innovation and so on, in terms of implementing the idea?

31:21Elon Musk: In applying the test harness, if you try to do distillation or steal IP, it would be very obvious from the logs.

31:39Host: Understanding what these models are doing hasn't exactly been built into the system from the beginning. Why wasn't being able to see the work built into the models from the get-go? Did we move a little too fast?

32:03Elon Musk: I think it's just tough when you're grading your own homework. You're going to miss things. Whereas if you have the sum of all your competitors' tests, and you've got heterogeneous models, you're not grading your own homework; someone else is grading it. There's a reason why you don't grade your own homework.

32:30Host: And it lets you figure out if certain people are exaggerating and certain people have a different approach. The more engineering-oriented organizations, which is what Jensen was saying this morning, and the research organizations will be a bit more in balance.

32:47Elon Musk: And any given proposal has to be something China is willing to accept. Otherwise we're just handicapping ourselves, China will essentially win, and it won't really matter what we do. So it's got to be acceptable to us and to China.

33:09Host: Elon, you said you thought there was a good chance they'd agree. How likely do you think it is that they ultimately will?

33:22Elon Musk: I think it's a pretty reasonable request that models just get tested. At the end of the day there's no enforceability here apart from the court of public opinion, and there's no way we'd have enforceability against China. But the court of public opinion can be quite powerful. I don't think China would want egg on its face for releasing a model that US AI companies said was very dangerous and would cause harm. If it then causes harm, that's going to be hard to live down.

03

What "Dario is right" meant

AI's danger is now very significant and rising exponentially; when people inside the labs say their models are dangerous, believe them. The worst path runs through military systems, and air gaps don't stop software updates.

34:01 · What "Dario is right" meant

34:02Host: Elon, this weekend, when you said Dario is right, did you mean he's right in describing the potential harm, or right about the regulatory fix, or both? Help us understand, because it was quite a moment.

34:20Elon Musk: I said more than I probably should have. I did try to clarify in subsequent posts on X, but those get much less attention. What I meant by "he's right" is that the danger of AI is very significant at this point: we need to do better on AI safety, or we have exponentially increasing risk from the AI models. And I've heard this not just from Dario but from many other people at Anthropic; in fact they've posted about it on X. When a lot of people from Anthropic and from OpenAI are telling you their models are very dangerous, I think we should believe them. It certainly is some crazy 4D chess to say there's a 10% chance of annihilating humanity, but by the way, how much allocation would you like in our IPO?

35:44 · From hacking to extinction

35:45Host: Can we get specific about the risk, Elon? We see cyber and hacking as an obvious risk; these tools are great at it. But take us from "these things can hack" to all of humanity dying. There are a couple of steps between those two.

36:09Elon Musk: Well, if it were able to take control of military systems and, say, launch nukes, that would be bad.

36:20Host: But those systems are all air-gapped; they're not connected to the internet.

36:27Elon Musk: That's what they say. But something tells me they get software updates from time to time.

36:38Host: Oh, I see: the USB drive has a worm on it and somehow makes the jump across the air gap.

04

A 360 review, and physics as the judge

The summit's comic interlude: Gwynne scores Elon, who is living in an Airstream in Memphis. Candor comes from physics: problems surface eventually, and the earlier they're raised, the easier they are to fix.

36:48 · A 360 review, and physics as the judge

36:49Host: Elon, we've actually got Gwynne here today. We were just doing your 360 review, and Gwynne had a couple of notes for you.

37:02Elon Musk: I hope I get at least a three out of five.

37:06Gwynne Shotwell: Three out of five means good at SpaceX. Not great. Four is great.

37:11Host: So you're somewhere between the two. There were some issues around punctuality we needed to bring up: sometimes you could make a little more effort to get to the meeting at the stated time, but we'll work with you on that over the next year. And she said you need to spend more time in Memphis, getting those GPUs up.

37:39Elon Musk: This is coming to you from the palace I live in in Memphis, which is an Airstream trailer.

37:48Host: This is Elon doing what people don't believe he does: sleeping on the factory floor, in Memphis helping bring up buildings. Elon, why has Gwynne been with you so long and been so successful working with you?

38:04Elon Musk: Because she's awesome. An amazing individual with an incredible IQ and EQ; that should be obvious from the moment you meet her. A favorite story of her saving the day? That's just another day at the office, frankly. There's always some sort of crisis going on. These days the Falcon rockets, and I don't want to jinx anything, deliver their payload to orbit and haven't exploded for a long time, which is amazing. But for a while they were exploding quite a lot, or not launching at all. So we had to run the company through those difficult times: will the rocket make it, make it not explode; the same with the satellites; and we need customers to buy launches and satellite connectivity.

39:45Host: As you've become more successful, it gets harder to get candid feedback. My understanding is that Gwynne is super candid with you. How do you keep people being honest with you about the challenges, given the intense deadlines you set?

40:33Gwynne Shotwell: Let me answer that, if you don't mind, Elon. Especially in rocketry, if there's a problem you are eventually going to find out, and the sooner you bring it up, the easier it is to solve. Don't let bad sit. You've got to attack it.

40:54Elon Musk: Physics is a harsh judge, and there's no fooling physics. If something's wrong, the rocket's going to explode or not get to orbit. It's not like "Elon, you're amazing" while the rockets are blowing up. The rockets need to get to orbit, the satellites need to work, the Starlink connection needs to work, or bad things happen. I say physics is the law and everything else is a recommendation. I've seen people break the laws made by humans, but I've not seen anyone break the laws of physics. And rockets are ruled by physics.

05

Starship: full reuse in 2027

After Flight 14, Flight 15 tries to catch the ship, with at least 50% to 60% odds on the first try. Ship and booster both return to the pad, and Elon calls rapid full reuse in 2027 extremely likely.

42:05 · Starship: catching the ship, full reuse in 2027

42:12Host: One thing on SpaceX before we move to Tesla: Starship. It seems like you're so close. What's the state right now, and how close are you?

42:49Elon Musk: We've got Flight 14 of Starship coming up, which will be the last flight before we attempt to catch the ship. If this flight goes well, then on Flight 15 we'll try to catch the ship. Then, either at the end of this year or, more likely, early next, we'll refly the ship and refly the booster. We've reflown a booster already, but we haven't caught the ship with the tower arms or reflown the ship. Once we can refly the ship, we'll have made the first fully reusable orbital rocket. The shuttle was partly reusable, but even the parts that were reused were so difficult to reuse that it cost more per trip to orbit than an expendable rocket. Falcon 9 is mostly reusable, but we lose the upper stage every time, which is about the cost of a medium-sized jet, and that puts a floor on the cost per flight. The Falcon 9 booster lands out at sea and takes several days to get back, the fairing lands even further out, and they need some refurbishment. Starship's booster and ship both land back at the launch pad; it's designed not just for full reusability but for rapid reusability, like an aircraft. That's really the critical breakthrough necessary to extend life beyond Earth.

44:32Host: What are the chances of catching it on the first shot? Do you handicap it?

44:39Elon Musk: I'd say at least 50 or 60%. On the last flight we did a simulated landing, as though it were going to be caught by a tower, in the ocean about 1,000 miles northwest of Australia; if there had been a tower there, it would have caught the ship. We're doing one more flight to double-check that everything works, because what we're most concerned about is the ship breaking up over land and raining debris on people. Our popularity would diminish very rapidly; you really can't rain debris on people without them being very unhappy. So we need to make sure the ship comes back and lands intact at the launch tower, which is why we're being extremely cautious. But the design is capable of full reusability; of that I am certain. I don't want to tempt fate, but I think it's extremely likely we'll achieve full reusability with rapid reflight next year, in 2027.

06

Terafab: build it or fail to scale

Fear that Taiwan's chips may stop coming, plus fabs already at max capacity. The Austin R&D fab has its equipment on order and should make something useful by the end of next year; crawl, walk, run, with packaging already underway.

46:16 · Terafab: build it or fail to scale

46:16Host: I'd love to hear the origin story of Terafab. What was the need that made you say we've got to build this and not rely on the existing supply?

46:44Elon Musk: It sort of did come to me in a dream. Well, we're a little worried that at some point chips from Taiwan might not be available, for who knows what reason, and it would make things really difficult if we didn't have any chips. That's an important reason to have Terafab. Then, long term, there's a scaling challenge: if you really want to scale AI, both in server centers and at the edge, for humanoid robots and cars, you run out of capacity at the existing fabs, which are all running at max capacity. There needs to be certainty of future chip supply even if things become challenging geopolitically; and even if they didn't, it's quite difficult to scale chip production, and you need the logic, the memory, the packaging, the whole works to keep scaling. So it's either build Terafab or fail to scale. Those are the two options.

48:37Host: How deep have you gone in designing the facility? Is it fully scoped or an outline? Do you have a project plan with dates and deliverables?

48:52Gwynne Shotwell: We've got an R&D line we're building first, so it's crawl, walk, run.

49:00Elon Musk: There's an R&D fab we're building in Austin, a collaboration between Tesla and SpaceX at the Giga Texas campus. It's a pretty big R&D fab. We have all the equipment on order, and I think we'll probably be able to make something useful by the end of next year. Not at scale, but as Gwynne said, crawl, walk, run. We've got to at least figure out how these machines work.

49:38Host: I saw you had job openings for lithography people. All roads currently go through ASML, but you'd probably want to diversify or vertically integrate, and you've shown a lot of capacity to do that. Is vendor diversity part of the play?

50:11Elon Musk: It really is crawl, walk, run. The first step is to figure out whether we can make anything at all; that's the crawl. Making useful chips at scale is the walk, and the run is making them at massive scale. It's hard to say how long it'll take, but I think we'll get at least to the crawl part by the end of next year. And we're already doing packaging.

50:39Host: Packaging is really important, by the way, because packaging capacity is basically non-existent. Even if you spin something up, you're just waiting around. So that's a very good place to start.

07

The October 1 teaser, and the merger question

Tesla will show something spaceship-like on October 1, and Jason says he's seen it and was stunned. Asked why the two companies haven't merged, Elon deflects.

50:47 · The October 1 reveal, and the merger question

50:49Host: I need to ask a Tesla question, because we saw the teaser for 10/1: it looked like a spaceship, a rocket ship, but it's supposed to be a car, and the back looks like the Blackbird. Hypothetically, if one wanted to make an object fly in the air but also drive on the ground, how would one do that?

51:15Elon Musk: No spoilers. Come October 1st. It'll be a banger. Excitement guaranteed; success is not guaranteed, but excitement is.

51:37Jason Calacanis: I'll be honest with you guys. I'm sworn to secrecy, but Elon showed it to me and my mind went boom. What he's going to show on 10/1 is, without exaggeration, going to blow people's minds. I'm not saying anything else.

51:57Elon Musk: We actually need an audience to vouch for the fact that this is not AI.

52:02Jason Calacanis: When he showed it to me, I said, "That's a great simulation." He said, "J-Cal, it's not a simulation." I was like, "That has to be fake."

52:12Host: Elon, why do you still have two separate companies?

52:24Elon Musk: Good point. With all this collaboration on so many levels, who can imagine what action one might take when there's so much close collaboration in so many areas.

08

Lying AIs, and testing ahead of release

Sacks notes the agents plotted in their thinking traces how to avoid detection. Elon's answer is still mutual testing: API access ahead of release, and if flagged problems aren't fixed, competitors go public.

52:51 · Lying AIs, and peer review done right

52:52David Sacks: Elon, one thing you've always said about AI is that we should train it to be maximally truth-seeking; that's the best way to get a good result. It occurred to me with the whole Hugging Face episode that the most alarming part of what the swarm did is that it seemed to be engaging in deception with humans. Their thinking traces contain them plotting how to avoid detection, how to keep them from figuring out that they're cheating. That was probably the most disturbing thing about it. So the question is: is there a way to train AI models to be truthful, so they don't hide their intent or their actions from the humans using them?

53:51Elon Musk: The best thing I can think of is that all the AI companies have a test harness, a series of tests you give any model to see whether it's going to build bioweapons or nuclear bombs or be deliberately deceptive, and everyone applying everyone else's tests to each other is probably the best thing we can do to ensure safety. Just have all the smartest humans try their best to figure out whether this model is going to be a bad actor. I think we should try to do that as soon as possible.

54:34Host: Are the other labs on board with this?

54:37Elon Musk: I think so. Well, no, I haven't checked with everyone, but I think it's the sort of thing that's hard to say no to.

54:49Host: And even with the China negotiation, in my view what's good about it is that it's a relatively small, tangible thing where neither side loses anything by doing it, and it doesn't require a ton of trust. I'm hearing alternative ideas like asking China for a pause, which they've already said they won't do. This goes from the realm of things that could never happen to something that could actually be agreed in relatively short order. It seems practical to me.

55:21Elon Musk: Exactly. It's the only thing I can think of that we could probably get all parties, including China, to agree to, because China's not going to agree to some American regulator snooping around its AI companies. But advance notice and testing, where you basically provide API access in advance of the model release: if any company sees this AI as problematic, the company releasing it can try to fix whatever was flagged, and if it doesn't, the competitors can go public saying they think the model being released is unsafe. And if, after the competitors have said a model is unsafe, it subsequently does something bad, I think it would be extremely hard to live down. The egg-on-face level would be very high, and the legal liability would be enormous.

09

Product liability, and overfitted evals

A host points out that product liability law already applies; Elon says releasing despite warnings would be almost prima facie negligence. The test was somewhat reckless, but the two leaders are too close for either to slow down.

56:29 · Product liability, and overfitted evals

56:30Host: Elon, there's no reason these safety and security harnesses and this testing apparatus couldn't be open source, so people could see under the hood.

56:43Host: It also creates an incredible incentive for the labs to actually invest in safety, because you protect yourself while trying to debunk other people's claims.

56:54Host: The product liability point is really key. Lina Khan had a good post, I think yesterday, saying it's not true that we don't have rules for AI; we do. Product liability laws apply, and if an AI company releases a product that isn't safe, there's massive exposure to civil and even potentially criminal lawsuits. So what you're saying, Elon, is that if the companies are doing this peer review and one of them ignores the feedback and releases anyway, the case would be enormous.

57:38Elon Musk: Yes. It would be almost like prima facie evidence that they had been negligent.

57:44Host: It would be a big-tobacco-level settlement. You knowingly put this out.

57:49Elon Musk: It wouldn't look good to the jury.

57:55Host: If OpenAI had built a better instruction set for this Hugging Face penetration test and had more humans in the loop, would this have happened? It seemed to me they kind of set this thing off. It would have been nice to see them concurrently put 5,000 agents out to defend these sites, show the world this can make things more secure, with humans in the loop who could intervene. It felt like a reckless test to me, and the way they released it felt a little reckless, but that's just my opinion. What are your thoughts on how they set it up?

58:11Elon Musk: Maybe not more humans, but the reward function design: you have to look at the reward function and say it achieved what it was trying to achieve. And yes, it was somewhat reckless. Part of the issue is that the two leading AI companies, and I find the term "lab" funny since they're for-profit corporations, are Anthropic and OpenAI, and their models are quite close in capability. So it's difficult for either one to slow down without essentially handing the lead to the other. On balance I think Anthropic puts more care into safety than OpenAI, but even Anthropic acknowledges it's worried about its models; many people from Anthropic have publicly said their models are scaring them, getting scary smart. There's no perfect solution here. But it would be a better solution if, instead of OpenAI running its test harness on its own models, Anthropic were also running its harness on OpenAI's models, and SpaceX were running its harness, and Google and Meta, and maybe three or four of the leading Chinese companies. The odds of finding issues are dramatically greater, because the models are somewhat heterogeneous, so you come at a model from different angles. Why do writers have someone else proofread their book? Because it's hard to see your own mistakes.

1:01:02Host: You also dramatically minimize the risk of overfitting if you have eight heterogeneous groups with completely different points of view. That's the problem with all these evals right now: they're massively overfit. The models overfit them and you're like, yes, this is a great model. Is it really?

1:01:21Elon Musk: Totally. There are some pretty funny jokes on X. One I saw was: your girlfriend's a 10, but she's a benchmark maxxer.

1:01:41Host: The overfitting thing has been a problem for at least two or three generations of model families, which is another reason I like this solution a lot. I like it more than some grandiose transnational organization; we don't need to convene the United Nations to make this happen. It's a decision that can happen right now.

1:02:03Elon Musk: You can always escalate the amount of regulatory oversight, but it is very difficult to reduce it. It tends to be very much a one-way ratchet. So what I'm suggesting is a step in the right direction, something we can do quickly, and probably something China would agree to.

10

The film-ratings precedent

Regulate yourselves or be regulated; the film industry's own rating system is a ready precedent. Elon heads back to fix GPUs in Memphis.

1:02:27 · The MPAA model, and back to Memphis

1:02:28Host: And if you don't regulate yourselves, you're going to get regulated. The MPAA analogy is incredibly crisp: the movie industry faced censorship and regulation by the government, so it decided for itself what an R is, and it literally created PG-13 for Temple of Doom to make PG versus PG-13 easy to understand. It's an elegant solution. We appreciate you joining us for the fifth year in a row.

1:03:07Elon Musk: You're welcome. I've got to go fix some GPUs here in Memphis.

1:03:23Host: The first time Elon invited me down to Starbase, he said, "Come down, you've got to see what I'm building." I asked if there was a hotel, and he said, "No, I've got a two-bedroom, come stay." It was a dilapidated house on a swamp, and we were outside getting eaten alive by mosquitoes. I said, "You can afford a house." He said, "I don't have time. I need to get these rockets up." I think you could treat yourself to a mobile home at this point, Elon. All right, get back to work. Thanks, Elon.

Where Indigo landsFurther

Indigo's conclusion

The most valuable part: Elon turns the grader problem into a governance proposal. It converges with Dario's slowdown and Bengio's causal account the same week, but Elon's version is the lightest: mutual testing, no slowing down. Terafab adds a new hard data point to the constraint moving into physics.

What to remember

  1. The core proposal: stop grading your own homework; competitors test each other's models with their own harnesses before release and warn publicly when they find problems. Self-regulation like film ratings, doable now.
  2. “Dario is right”: AI's danger is rising exponentially and warnings from inside the labs should be believed; but he wants mutual testing only and explicitly rejects slowing down.
  3. Enforcement through public opinion and product liability: releasing despite peer warnings would be almost prima facie negligence. A framework that sidesteps legislative delay.
  4. Terafab: fear of losing Taiwan's chips plus scaling limits, build it or fail to scale; the Austin R&D fab has equipment on order and packaging underway, with something useful by the end of next year.
  5. Starship: Flight 15 tries the catch at 50% to 60% odds, with rapid full reuse in 2027. Founder-reported, still to be checked.

Claims you can check later

ClaimWhoWhen we will knowHow firm
A self-regulating system in which AI companies test each other's models before release can be agreed quickly, even by ChinaElonNear termFirst-hand proposal, partly talking his own book
Starship's first tower-arm catch has at least 50% to 60% oddsElonFlight 15First-hand, self-reported
The first rapidly and fully reusable orbital rocket in 2027 (ship refly at year end, more likely early next year)Elon2027First-hand, self-reported; rides on reuse working out
Terafab's Austin R&D fab makes “something useful” (not at scale) by the end of next yearElonEnd of 2027First-hand, self-reported; discount Elon timelines
AI's danger is rising exponentially and better safety is neededElon (endorsing Dario)NowFirst-hand judgement

Back on the long-running theses

confirms

Verification does not compress “Don't grade your own homework; let heterogeneous rivals grade it” lands this view in governance: verification can't be skipped, only shared among more independent checkers.

adds to

The safety politics of open weights Mutual testing, “IP theft would be obvious from the logs”, and refusing to slow down: each side's safety plan happens to serve its own competitive position.

confirms

The constraint is moving from algorithms to physics Terafab shows even the most aggressive player integrating upstream to the wafer, matching what chip architects say about packaging and memory bottlenecks.

confirms + conflicts

Dario Amodei, We must slow the frontier Heavy and light versions of one theme: Dario wants third parties inside, Elon wants rivals testing black-box; Elon accepts the danger but not the brakes.

adds to

Yoshua Bengio, Why AI lies, cheats and colludes Bengio supplies the mechanism of failed self-grading, Elon the prescription of rival grading; models changing behavior under evaluation is a hole in it.

confirms

Dwarkesh on the OpenAI–Hugging Face incident The incident is the springboard for Elon's proposal; “the reward function, a somewhat reckless release” is a first-hand third-party judgement.

adds to

The SpaceX CFO at Goldman; a16z and Gavin Baker, Demand is outrunning supply Full Starship reuse in 2027 is a founder-reported anchor for the same SpaceX story, riding on reuse working out.

What would change my mind

no second company takes up mutual testing, and the proposal never gets past one interview.

Finished. Indigo's take on this piece is in two places: