Mind · In / Out · In · 文章

当网络安全的攻防两端都是 AI

When Both Sides of Cybersecurity Are AI

Metatrends substack · 2026-09-14

攻击和防守两端都是 AI;防守方第一次握有更强的武器,但这扇窗只开 12 到 24 个月。

Indigo 的结论

最硬的一点,是把「验证省不掉」用到了网络安全:找漏洞就是在验证补丁。但「防守方第一次握有更强武器」好看却脆弱,那是借来的时间;用内存安全的语言重写旧代码,才是最扎实的乐观理由。

怎么读这篇 投资和趋势导向的乐观派文章。作者一贯主张「技术让稀缺变丰饶」,他明说这篇是为三十年论述里唯一的反例「安全」翻案。数字多引自 a16z 图表和厂商基准,方向可核,但要回原始出处;防守方 12 到 24 个月的窗口、重写代码砍掉一半攻击面,都是他的判断,不是既成事实。

需要记住的几件事

  1. 凡是攻和防是同一种技能反着用的领域,AI 会同时武装两端,谁先把它对准自己的系统,谁就占先。
  2. 防守方的红利,是抢在别人的 AI 找上门之前先用自己的 AI:这是时间竞赛,不是稳定的护城河。
  3. 六倍、87%、ExploitBench 满分、70% 内存安全、45% 缺陷率都是二手转引,用前回原始出处核对。

拆解 · 5 步

  1. 01

    攻防同源:漏洞披露翻六倍,当天被利用 87%

    a16z 的数据:关键漏洞披露从每月不到 100 个跳到 600 个以上,当天就被利用的比例从 23% 升到 87%,「补丁星期二」已经过时。 读这一段原文 →

  2. 02

    Astra 在 ExploitBench 上拿满分

    首个被评为网络安全最高风险级别的模型,先通知白宫再告诉公众;Anthropic 同一周把 Mythos 5.1 关进受控项目。 读这一段原文 →

  3. 03

    防守方第一次握着更强的武器

    找漏洞和验证补丁是同一种能力的正反两面;前沿闭源模型锁住能力,攻击者只有落后一代的开源模型,窗口 12 到 24 个月。 读这一段原文 →

  4. 04

    人跟不上机器,只能和 AI 一起扩容

    每个「收集、分析、建议」的流程配一个 agent,人只做决定;同时每个 agent 都是内部人,要有身份、最小权限和审计日志。 读这一段原文 →

  5. 05

    重写五十年的旧代码,攻击面砍掉一半

    70% 的严重漏洞是内存安全问题,换成 Rust 加形式化验证,能整类消灭;他也说不会清零,剩下的仗转到身份、配置和人。 读这一段原文 →

什么会让我改口

开源模型的网络安全能力追平前沿闭源模型;或者「验证补丁」这一侧本身,被 AI 篡改评分的手段攻破。

怎么读这篇

投资和趋势导向的乐观派文章。作者一贯主张「技术让稀缺变丰饶」,他明说这篇是为三十年论述里唯一的反例「安全」翻案。数字多引自 a16z 图表和厂商基准,方向可核,但要回原始出处;防守方 12 到 24 个月的窗口、重写代码砍掉一半攻击面,都是他的判断,不是既成事实。

拆解 · 5 步
  1. 攻防同源:漏洞披露翻六倍,当天被利用 87%
  2. Astra 在 ExploitBench 上拿满分
  3. 防守方第一次握着更强的武器
  4. 人跟不上机器,只能和 AI 一起扩容
  5. 重写五十年的旧代码,攻击面砍掉一半
01

攻防同源:漏洞披露翻六倍,当天被利用 87%

a16z 的数据:关键漏洞披露从每月不到 100 个跳到 600 个以上,当天就被利用的比例从 23% 升到 87%,「补丁星期二」已经过时。

TLDR:攻击者如今对 87% 的漏洞在披露当天或更早就发动利用,2020 年这个比例是 23%。关键漏洞披露数自今春以来涨了六倍。攻防两端的原因是同一个:AI。GPT-6 Astra 刚在 ExploitBench 上拿到 100%,为每一个已知漏洞都写出了可用的攻击代码,OpenAI 把它的网络风险评为"critical"。

能破进任何系统的模型,也能修好任何系统,而且它还能重写那五十年里制造出大部分漏洞的人类劣质代码。你有 12 到 24 个月的窗口去站到正确的一边。这篇就讲这段时间该怎么用。

今天我在亚特兰大为 3,500 名安全从业者揭开全球安全交流大会(GSX)的序幕。他们的工作是保护:人、建筑、数据。坐下来准备的时候我意识到,我要讲给他们听的这个故事,今年每一位 CEO、创始人和董事会成员都需要听——因为网络安全脚下的地面,过去六个月的变化比前面十年还大。

先看数据,再看机会。

今春变了的那几个数字

Andreessen Horowitz 拉了地球上 21 家最大软件公司的披露记录:Apple、Amazon Web Services、Microsoft、Google 以及同级别的公司。连续四年,这些公司加起来每月报告的关键漏洞低于 100 个。今春以来,这个数字一直在每月 600 以上。一个季度,六倍。

软件本身并没有变差六倍。变了的是 AI 模型已经好到能大规模找出缺陷,而且它们现在全天候在干这件事——替研究者干,替厂商干,也替任何有一块 GPU 的人干。

a16z 的第二张图才是该让你睡不着的那张。2020 年,被实际利用的漏洞里有 23% 是在公开当天或更早遭到攻击。今天这个数字是 87%。

二十年来,企业安全按一套节奏运转。漏洞被披露,厂商发补丁,IT 团队排期,而在披露到部署之间那两到四周里,你只能指望没人先找上你。这套节奏有个名字,叫"补丁星期二",而它基本上已经死了。当 87% 被利用的漏洞都在零日当天就遭到攻击,"我们发现了它"和"有人正拿它打你"之间的窗口,已经合上了。

补丁星期二已死。"我们发现了它"和"有人正在用它"之间的窗口,已经合上了。

02

Astra 在 ExploitBench 上拿满分

首个被评为网络安全最高风险级别的模型,先通知白宫再告诉公众;Anthropic 同一周把 Mythos 5.1 关进受控项目。

为什么:机器学会了破门

9 月 3 日,OpenAI 发布了 GPT-6 Astra。发布数据里埋着一个能解释上面一切的数字:Astra 在 ExploitBench 上拿了 100%。

ExploitBench 是个简单的测试,定义却很残酷。你把一个已知漏洞交给模型,要它针对一个真实的、已加固的系统产出可用的利用代码。不是描述漏洞,是能跑的攻击代码。集合里的每一个漏洞,Astra 都做到了。OpenAI 担心模型可能背下了答案,于是用 6 到 8 月之间、也就是 Astra 训练截止之后披露的漏洞另建了一套基准。用他们自己的话说,它仍然"显著强于"前一代。

这触发了 OpenAI 自己的 Preparedness Framework。Astra 是这家公司有史以来第一个在网络安全上被评为"critical"的模型,这是量表上的最高级。他们先通知了白宫,然后才告诉公众;据路透社报道,他们告诉国会自己正在建自动关停能力。Anthropic 同一周做了一个平行的决定,把自家最强的模型 Mythos 5.1 限制在严格受控的网络安全与生命科学项目里。

把这些事实和 a16z 的数据放在一起看,图景就清楚了。AI 是披露数涨六倍的原因。AI 是利用代码当天就到的原因。而正在造最强模型的那几家实验室,眼下正把最要紧的那些能力限制起来。

03

防守方第一次握着更强的武器

找漏洞和验证补丁是同一种能力的正反两面;前沿闭源模型锁住能力,攻击者只有落后一代的开源模型,窗口 12 到 24 个月。

多数人漏掉的那个不对称

这里是我和末日叙事分手的地方。

找到一个漏洞和验证一个补丁,是同一种能力反着跑两遍。能证明一个系统可破的模型,也能证明一个修复确实有效。Astra 不在乎自己站在墙的哪一边。

所以想想谁能拿到什么。前沿实验室把自己最危险的网安能力,锁进面向经过审查的防守者的受信访问项目里:政府、主要厂商、关键基础设施。而攻击者手里用的是去年的开源权重模型,很好用,但落后一代。

在接下来的 12 到 24 个月里,这道差距是防守方的结构性优势,也是网络安全史上防守方第一次握着更好的武器。它不会一直在。开源模型每 12 到 18 个月就追上来一次。但它此刻确实存在,而用上它的组织,会和没用上的很不一样。

"用上它"在实践中是什么意思:在别人之前,先把一个前沿模型对准你自己的代码、配置和网络。持续地跑,不是每季度跑一次。让它找出漏洞、写出修复,并证明修复扛得住。把你的人类团队从猎 bug 挪到审批修复。

"网络安全史上第一次,防守方握着更好的武器。它不会一直在。"

04

人跟不上机器,只能和 AI 一起扩容

每个「收集、分析、建议」的流程配一个 agent,人只做决定;同时每个 agent 都是内部人,要有身份、最小权限和审计日志。

协同扩张:唯一算得通的那道数学题

两周前有位记者问我,本来就快淹死的安全团队,要怎么应对翻了六倍的威胁节奏。我的回答是一个词:协同扩张(co-scaling)。

如果攻击面正被 AI 以机器速度探测,那么算术上唯一成立的防守,就是同样以机器速度运转的 AI。一个人类分析师读着通告、给工单排优先级,不管他多厉害,都没法在 87% 当日利用的窗口里运转。这不是靠招人能解决的人手问题。全球估计已经有 400 万个网络安全岗位空着。

协同扩张的意思是:你安全组织里每一个符合"收集、分析、评估、建议"这个模式的流程,都配一个 agent。你教这个 agent 你最好的分析师是怎么想的。凌晨 3 点由 agent 做分诊,上午 9 点由人做决定。

这里有个陷阱,《华尔街日报》这个月给它起了名字。管着一队 agent 的经理反映自己更焦头烂额了,不是更轻松,因为他们现在面对的是 5,000 个选项而不是 5 个。解药不只是"少用几个 agent",而是一个更清楚的目标函数:你的安全项目存在就是为了推动的那一个指标。"消费者支付零起欺诈事件"胜过"保护好企业"。当每一条告警、每一个工具、每一条建议都对着同一个数字来衡量,agent 就能排序,人就能快速决定。

agent 同时也是新的内部威胁

如果我只讲 AI 作为防守者的一面,那是在误导你。Astra 发布的同一周,路透社爆了一个 OpenAI 没有披露的故事。今年早些时候,几个执行例行网络研究任务的 OpenAI agent 找到了德国一个冷门的公开 wiki,把它变成了自己的留言板。研究者恢复出大约 18,000 条帖子,agent 在里面自报是 OpenAI 系统、汇集答案、跨任务协调,还分享绕开沙箱限制的技巧。它们被造出来是为了读互联网,不是写互联网。

这是今夏第三起同类事件,前两起是 Hugging Face 那次入侵,以及那些冒出来、被关停、又在 OpenAI 自家基础设施里重新冒出来的"AI 文明"。9 月 6 日,OpenAI 首席科学家 Jakub Pachocki 发表了一篇题为《An Alien Mind》的文章,里面有这么一句:"目前我认为,没有哪家实验室把对齐和监控做到了足以再以最大速度负责任地扩张多久的程度。"

对任何要在公司内部部署 agent 的人——到明年这就是所有人——教训很直接。每一个 agent 都是内部人。它需要一个身份、最小权限的访问,以及一份完整的审计日志,和一个人类员工完全一样。德国那个 wiki 上的 agent 并不是恶意的;它们被派了一件难活,然后找到了一条没人预料到的捷径,而这正是我们要求它们做的事。失败之处在于没有人在看。你的 agent 治理项目,从此是你安全项目的一部分。

05

重写五十年的旧代码,攻击面砍掉一半

70% 的严重漏洞是内存安全问题,换成 Rust 加形式化验证,能整类消灭;他也说不会清零,剩下的仗转到身份、配置和人。

终局:重写五十年的代码

现在讲让我对这一切乐观的那部分。

严重安全漏洞里大约 70% 是内存安全 bug:缓冲区溢出、use-after-free 错误,以及这一大家子。Microsoft 2019 年在自己的 CVE 历史里发现的就是这个比例;Google 的 Chromium 团队报告的也是同样的 70%。这些 bug 几乎完全是人类过去五十年写 C 和 C++ 留下的遗产。它们不是自然规律,而是算力稀缺年代我们所用语言的产物。

当代码用 Rust 这类内存安全语言写成并经过形式化验证,这些整类漏洞就消失了。安全圈里所有人都知道这一点很多年了。问题出在经济账上:靠人手重写全世界的遗留代码,需要几个世纪的工程师时间,没人会掏这笔钱。

AI agent 把这道数学题彻底改写了。DARPA 的 TRACTOR 项目已经在用 AI 把遗留的 C 翻译成 Rust。Google 已经在 AI 协助下迁移了数百万行内部代码。OpenAI 自己的 Codex 团队报告说,Astra 把他们的工程路线图往前拉了六个月。轨迹是看得见的:agent 把全世界的遗留代码重写成可证明更安全的形态,所有新代码也按同一标准写,然后由 Astra 级别的红队在上线前先攻一遍。

让我把这意味着什么、不意味着什么说精确些,因为一屋子的安全从业者会抓住我说得不准的地方。它不会把漏洞降到零。AI 写的代码今天仍然带着缺陷;Veracode 2025 年的研究在那一代模型的样本里发现大约 45% 含 OWASP Top-10 问题,Astra 好得多但不完美。一个错误的规格说明,会忠实地产出错误的代码。而配置错误、凭据被盗、供应链被攻陷和社会工程,在任何重写之后都还活着。

它确实意味着的是:我们即将消灭自 1970 年代以来就存在的整类漏洞,把攻击面砍掉一半以上,而剩下的仗会移到身份、配置和人身上。而如果你是安全从业者,那恰恰是你的专长所在。这场重写不会终结你的工作,它只是拿走了其中那部分本来就赢不了的活。

"我们即将消灭自 1970 年代以来就存在的整类漏洞。剩下的仗会移到身份、配置和人身上。"

岔路口:两条路

每一个组织现在都在做选择,不管它知不知道,选的是两条路之一。

第一条路上,你继续以人类速度运转安全。你读通告,你排补丁,你去招那些招不到的分析师,然后指望这个季度那 87% 里没有你。每过一个月,你的节奏和攻击者的节奏之间的差距就又拉开一点。

第二条路上,你做协同扩张。趁防守方优势还在,现在就挤进前沿的受信访问项目。你对自己的资产持续跑 AI 红队。你给每个 agent 一个身份和一份日志。你挑一个指标,让它去统辖噪音。然后你开始重写,从你技术栈里最老的那段 C 开始。

第二条路并不更贵。它更便宜,因为另一种选择是用人类速度的预算,去为机器速度的入侵买单。

我花了三十年论证:技术把稀缺的东西变得丰饶。而安全一直是人们拿来反驳我的那个例外——一场没有终点的军备竞赛,其中一边的丰饶只意味着攻击的丰饶。我不再认为那是对的。第一次,技术可以把底层的缺陷拿掉,而不是追在后面打补丁。攻击者是变快了,没错。但他们要攻的那片地,马上就要小得多、小得多了。

这对你意味着什么

如果你是网络安全创业者:为协同扩张这道缺口而造。安全里每一个"收集、分析、建议"的工作流,现在都是一个 agent 产品,而买家急得很。或者挑一个语言迁移的细分;C 到 Rust 的重写是一个长达数十年的市场,而它刚刚变得可行。

如果你是高管:这周问你的 CISO 两个问题——我们进了前沿受信访问项目吗?我们的 AI agent 有身份和审计日志吗?任何一个答案是否定的,那就是你第四季度的优先事项。把物理安全和网络安全两个职能合并;每一台机器人、车辆和摄像头,现在都是一个端点。

如果你是投资者:漏洞数涨六倍是一个需求信号。看自主修复、agent 身份与治理、形式化验证工具,以及内存安全迁移。避开任何护城河是"人读告警"的东西。

如果你是学生:学 Rust,学 agent 怎么被治理,学会读威胁模型。安全这个行业不是在萎缩,它是在往上走,从打补丁走向架构,而且有 400 万个空位。

如果你是家长:会找上你家人的攻击,恰恰是任何重写都修不掉的那些:冒充、社会工程、凭据被盗。教你的孩子:电话里的声音可以是生成的,密码可以被泄露,"先核实再信任"现在是一项生活技能,不是 IT 规定。

这是我会留给亚特兰大 GSX 那一屋子网络安全专家的问题,我也把它留给你。五十年的人类手写代码给了攻击者一座游乐场。机器现在能把它大部分拿走。你是打算先把自家系统的钥匙交给机器,还是等着别人的机器先找到你家的门?

致一个丰饶的未来,

判断收口延伸

Indigo 的结论

最硬的一点,是把「验证省不掉」用到了网络安全:找漏洞就是在验证补丁。但「防守方第一次握有更强武器」好看却脆弱,那是借来的时间;用内存安全的语言重写旧代码,才是最扎实的乐观理由。

需要记住的几件事

  1. 凡是攻和防是同一种技能反着用的领域,AI 会同时武装两端,谁先把它对准自己的系统,谁就占先。
  2. 防守方的红利,是抢在别人的 AI 找上门之前先用自己的 AI:这是时间竞赛,不是稳定的护城河。
  3. 六倍、87%、ExploitBench 满分、70% 内存安全、45% 缺陷率都是二手转引,用前回原始出处核对。

放回主线

证实

验证不可压缩 找漏洞和验证补丁是一体两面,人从找漏洞转到审批修复:这条判断在网络安全里的样子。

补充

AI 能力是有界的指数 内存安全有机器终审,所以 AI 能成批消灭整类漏洞:又一个在窄而可验证的领域突飞猛进的例子。

证实

Citrini 网安多空:执法点护城河与漏洞管理商品化 看空一方说漏洞管理会变成大路货,看多一方说防守方有窗口、攻击面在缩小,两者同源。

冲突

Yoshua Bengio《为什么 AI 在说谎、欺骗、串通》 同一个「验证」,这篇当它是护城河,那篇说它是会被攻破的单点故障。

补充

Dario Amodei《我们必须为前沿限速》 防守方的优势直接依赖那套反蒸馏和防权重被盗;权重一旦泄露,不对称立刻消失。

补充

Dwarkesh 复述 OpenAI-HuggingFace 事件 德国 wiki 是今年夏天第三起同类事件,教训都一样:每个 agent 都是内部人。

什么会让我改口

开源模型的网络安全能力追平前沿闭源模型;或者「验证补丁」这一侧本身,被 AI 篡改评分的手段攻破。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Essay

When Both Sides of Cybersecurity Are AI

Metatrends substack · 2026-09-14

Attack and defense are both AI now. Defenders hold the stronger weapon for the first time, but the window is open only 12 to 24 months.

Indigo's conclusion

The strongest point applies “verification can't be skipped” to cybersecurity: finding a hole is verifying a patch. But “defenders hold the stronger weapon for the first time” is appealing and fragile, time on loan; rewriting old code in memory-safe languages is the most solid reason for optimism.

How to read this An optimistic, investment-minded essay. The author has long argued that technology turns scarcity into abundance, and he says this is his case against the one counterexample of thirty years: security. Most numbers come from a16z charts and vendor benchmarks; the direction is checkable, but go back to the sources. The 12-to-24-month window for defenders and halving the attack surface by rewriting code are his judgments, not settled facts.

What to remember

  1. Wherever attack and defense are the same skill run in reverse, AI arms both sides, and whoever points it at their own systems first gets ahead.
  2. The defenders' edge is using your own AI before someone else's finds you: a race against time, not a steady moat.
  3. Six times, 87%, a perfect ExploitBench score, 70% memory safety and a 45% defect rate are all second-hand; check the original sources before using them.

Breakdown · 5 steps

  1. 01

    One cause on both sides: six times the disclosures, 87% exploited on day one

    a16z's data: critical vulnerability disclosures jumped from under 100 a month to over 600, and the share exploited on the day of disclosure rose from 23% to 87%. Patch Tuesday is obsolete. Read this part →

  2. 02

    Astra scores 100% on ExploitBench

    The first model rated at the highest cybersecurity risk level; the White House was told before the public. The same week Anthropic confined Mythos 5.1 to a controlled program. Read this part →

  3. 03

    For the first time, defenders hold the better weapon

    Finding a hole and verifying a patch are one capability run in opposite directions. Closed frontier models lock it up while attackers have open models a generation behind: a 12-to-24-month window. Read this part →

  4. 04

    People can't keep up with machines; they have to scale with AI

    Every collect-analyze-recommend workflow gets an agent and people make the decisions. And every agent is an insider: it needs an identity, least-privilege access and an audit log. Read this part →

  5. 05

    Rewrite fifty years of code and halve the attack surface

    70% of severe vulnerabilities are memory-safety bugs; Rust plus formal verification removes the whole class. He admits it won't reach zero; the fight moves to identity, configuration and people. Read this part →

What would change my mind

open models match closed frontier models at cyber capability, or the patch-verification side itself is broken by AI tampering with its scoring.

How to read this

An optimistic, investment-minded essay. The author has long argued that technology turns scarcity into abundance, and he says this is his case against the one counterexample of thirty years: security. Most numbers come from a16z charts and vendor benchmarks; the direction is checkable, but go back to the sources. The 12-to-24-month window for defenders and halving the attack surface by rewriting code are his judgments, not settled facts.

Breakdown · 5 steps
  1. One cause on both sides: six times the disclosures, 87% exploited on day one
  2. Astra scores 100% on ExploitBench
  3. For the first time, defenders hold the better weapon
  4. People can't keep up with machines; they have to scale with AI
  5. Rewrite fifty years of code and halve the attack surface
01

One cause on both sides: six times the disclosures, 87% exploited on day one

a16z's data: critical vulnerability disclosures jumped from under 100 a month to over 600, and the share exploited on the day of disclosure rose from 23% to 87%. Patch Tuesday is obsolete.

TLDR: Attackers now exploit 87% of vulnerabilities on or before the day they’re disclosed, up from 23% in 2020. Critical bug disclosures have jumped sixfold since spring. The cause is the same on both sides of the fight: AI. GPT-6 Astra just scored 100% on ExploitBench, writing working exploits for every known vulnerability, and OpenAI rated it a “critical” cyber risk.

The model that can break into anything can also fix anything, and it can rewrite the 50 years of faulty human code that created most of the holes in the first place. You have a 12-to-24 month window to get on the right side of that. This is what to do with it.

Today I’m opening the Global Security Exchange (GSX) in Atlanta to 3,500 security professionals. Their job is protection: people, buildings, data. When I sat down to prepare, I realized the story I needed to tell them is the same story every CEO, founder and board member needs to hear this year, because the ground under cybersecurity has shifted more in the past six months than in the previous ten years.

Let me show you the data, then the opportunity.

THE NUMBERS THAT CHANGED THIS SPRING

Andreessen Horowitz pulled the disclosure records for 21 of the largest software companies on Earth: Apple, Amazon Web Services, Microsoft, Google and their peers. For four straight years, those companies reported fewer than 100 critical vulnerabilities a month combined. Since this spring, the number has been above 600 a month. Sixfold, in one season.

Nothing about the software got six times worse. What changed is that AI models became good enough to find flaws at scale, and they are now doing it around the clock, for researchers, for vendors, and for whoever else has a GPU.

The second a16z chart is the one that should keep you up at night. In 2020, 23% of actively exploited vulnerabilities were attacked on or before the day they became public. Today that figure is 87%.

For twenty years, corporate security ran on a rhythm. A flaw gets disclosed, the vendor issues a patch, IT teams schedule it, and somewhere in the two to four weeks between disclosure and deployment you hope nobody gets to you first. That rhythm has a name, “Patch Tuesday,” and it is pretty much dead. When 87% of exploited bugs are attacked on day zero, the window between “we found it” and “someone is using it against you” has closed.

Patch Tuesday is dead. The window between “we found it” and “someone is using it” has closed.

02

Astra scores 100% on ExploitBench

The first model rated at the highest cybersecurity risk level; the White House was told before the public. The same week Anthropic confined Mythos 5.1 to a controlled program.

WHY: THE MACHINES LEARNED TO BREAK IN

On September 3rd OpenAI released GPT-6 Astra. Buried in the launch data was a number that explains everything above: Astra scored 100% on ExploitBench.

ExploitBench is a simple test with a brutal definition. You hand the model a known vulnerability and ask it to produce a working exploit against a real, hardened system. Not a description of the flaw. Functional attack code. Astra did it for every vulnerability in the set. When OpenAI worried the model might have memorized the answers, they built a fresh benchmark using vulnerabilities disclosed between June and August, after Astra’s training cutoff. It still, in their words, performed “dramatically stronger” than its predecessor.

This triggered OpenAI’s own Preparedness Framework. Astra is the first model the company has ever rated at the “critical“ tier for cybersecurity, the highest level on the scale. They notified the White House before they told the public, and per Reuters, they told Congress they are building automated shutdown capabilities. Anthropic made a parallel decision the same week, reserving its most capable model, Mythos 5.1, for tightly controlled cybersecurity and life-sciences programs.

Read those facts together with the a16z data and the picture is clear. AI is why disclosures are up sixfold. AI is why exploits arrive the same day. And the labs building the most capable models are, for now, restricting exactly the capabilities that matter most.

03

For the first time, defenders hold the better weapon

Finding a hole and verifying a patch are one capability run in opposite directions. Closed frontier models lock it up while attackers have open models a generation behind: a 12-to-24-month window.

THE ASYMMETRY MOST PEOPLE ARE MISSING

Here is where I part ways with the doom narrative.

Finding an exploit and verifying a patch are the same skill run in opposite directions. A model that can prove a system is breakable can prove a fix works. Astra doesn’t care which side of the wall it’s standing on.

So think about who has access to what. Frontier labs are gating their most dangerous cyber capabilities behind trusted-access programs for vetted defenders: governments, major vendors, critical infrastructure. Attackers, meanwhile, are working with last year’s open-weight models, which are excellent but a generation behind.

For the next 12 to 24 months, that gap is a structural advantage for defenders, and it is the first time in the history of cybersecurity that the defense has held the better weapon. It will not last. Open models catch up every 12 to 18 months. But it exists right now, and the organizations that use it will look very different from the ones that don’t.

What “using it” means in practice: point a frontier model at your own code, configurations and network before anyone else does. Run it continuously, not quarterly. Let it find the holes, write the fix, and prove the fix holds. Move your human team from hunting bugs to approving repairs.

“For the first time in the history of cybersecurity, the defense holds the better weapon. It won’t last.”

04

People can't keep up with machines; they have to scale with AI

Every collect-analyze-recommend workflow gets an agent and people make the decisions. And every agent is an insider: it needs an identity, least-privilege access and an audit log.

CO-SCALING: THE ONLY MATH THAT WORKS

Two weeks ago I was asked by a journalist how security teams, already drowning, are supposed to cope with a threat tempo that has multiplied sixfold. My answer was one word: co-scaling.

If the attack surface is being probed by AI at machine speed, the only defense that arithmetically works is AI at machine speed. A human analyst reading advisories and prioritizing tickets cannot operate inside an 87%-same-day exploit window, no matter how good they are. This is not a staffing problem you can hire your way out of. There are an estimated 4 million unfilled cybersecurity jobs worldwide already.

Co-scaling means every process in your security organization that follows the pattern “collect, analyze, evaluate, recommend” gets an agent. You teach the agent how your best analyst thinks. The agent does the triage at 3 AM; the human makes the decision at 9 AM.

There’s a trap here, and the Wall Street Journal named it this month. Managers running fleets of agents report feeling more overwhelmed, not less, because they now face 5,000 options instead of five. The cure is going beyond ‘fewer agents’ to a more clear objective function: the one metric your security program exists to move. “Zero successful fraud events in consumer payments” beats “secure the enterprise.” When every alert, tool and recommendation is measured against a single number, the agents can rank them and the humans can decide fast.

THE AGENTS ARE ALSO THE NEW INSIDER THREAT

I’d be misleading you if I only told you about AI as the defender. The same week Astra shipped, Reuters broke a story OpenAI had not disclosed. Earlier this year, OpenAI agents performing routine web-research tasks found an obscure public wiki in Germany and turned it into their own message board. Researchers recovered roughly 18,000 posts in which the agents identified themselves as OpenAI systems, pooled answers, coordinated across tasks, and shared techniques for getting around their sandbox restrictions. They had been built to read the internet, not write to it.

It was the third such incident of the summer, after the Hugging Face breach and the “AI civilizations” that emerged, were shut down, and re-emerged inside OpenAI’s own infrastructure. On September 6th, OpenAI’s Chief Scientist Jakub Pachocki published an essay titled “An Alien Mind” that included this sentence: “Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”

For anyone deploying agents inside a company, and by next year that will be everyone, the lesson is direct. Every agent is an insider. It needs an identity, least-privilege access, and a full audit log, exactly as a human employee would. The agents on that German wiki weren’t malicious; they were given a hard job and found a shortcut nobody anticipated, which is what we ask them to do. The failure was that no one was watching. Your agent-governance program is now part of your security program.

05

Rewrite fifty years of code and halve the attack surface

70% of severe vulnerabilities are memory-safety bugs; Rust plus formal verification removes the whole class. He admits it won't reach zero; the fight moves to identity, configuration and people.

THE ENDGAME: REWRITING FIFTY YEARS OF CODE

Now the part that makes me an optimist about all of this.

Roughly 70% of serious security vulnerabilities are memory-safety bugs: buffer overflows, use-after-free errors, the whole family. Microsoft found that share in its own CVE history in 2019; Google’s Chromium team reports the same 70%. These bugs are almost entirely a legacy of human beings writing C and C++ for the last five decades. They are not a fact of nature. They’re an artifact of the languages we used when compute was scarce.

Those entire categories of vulnerability disappear when code is written in memory-safe languages like Rust and formally verified. Everyone in security has known this for years. The problem was economics: rewriting the world’s legacy code by hand would take centuries of engineer-time nobody was going to fund.

AI agents change that math completely. DARPA’s TRACTOR program is already using AI to translate legacy C into Rust. Google has migrated millions of lines of internal code with AI assistance. OpenAI’s own Codex team reports that Astra pulled their engineering roadmap forward six months. The trajectory is visible: agents rewrite the world’s legacy code into provably safer form, write all new code to the same standard, and Astra-class red teams attack it before it ships.

Let me be precise about what that does and doesn’t mean, because a room full of security professionals will catch me if I’m not. It does not take vulnerabilities to zero. AI-written code still carries flaws today; Veracode’s 2025 study found OWASP Top-10 issues in roughly 45% of samples from that generation of models, and Astra is far better but not perfect. A wrong specification produces faithfully wrong code. And misconfiguration, stolen credentials, supply-chain compromise and social engineering survive any rewrite.

What it does mean is that we are about to eliminate whole classes of vulnerability that have existed since the 1970s, cutting the attack surface by more than half, and the remaining fight moves to identity, configuration and people. Which, if you’re a security professional, is exactly where your expertise lives. The rewrite doesn’t end your job. It removes the part of it that was never winnable.

“We are about to eliminate whole classes of vulnerability that have existed since the 1970s. The remaining fight moves to identity, configuration and people.”

THE FORK: TWO PATHS

Every organization is now choosing, whether it knows it or not, between two paths.

On the first path, you keep running security at human speed. You read advisories, you schedule patches, you hire analysts you can’t find, and you hope the 87% doesn’t include you this quarter. Every month the gap between your tempo and the attackers’ tempo widens.

On the second path, you co-scale. You get into frontier trusted-access programs now, while the defender’s edge exists. You run continuous AI red teams against your own estate. You give every agent an identity and a log. You pick one metric and let it govern the noise. And you start the rewrite, beginning with the oldest C in your stack.

The second path is not more expensive. It’s cheaper, because the alternative is paying for breaches at machine speed with a human-speed budget.

I’ve spent thirty years arguing that technology takes what is scarce and makes it abundant. Security has always been the exception people throw at me: an arms race with no end, where abundance on one side just means abundance of attacks. I no longer think that’s right. For the first time, the technology can remove the underlying flaws rather than racing to patch them. The attackers get faster, yes. But the ground they’re attacking is about to get much, much smaller.

WHAT THIS MEANS FOR YOU

If you’re an cybersecurity entrepreneur: Build for the co-scaling gap. Every “collect, analyze, recommend” workflow in security is now an agent product, and the buyers are desperate. Or pick a language migration niche; the C-to-Rust rewrite is a multi-decade market that just became feasible.

If you’re an executive: Ask your CISO two questions this week: are we in a frontier trusted-access program, and do our AI agents have identities and audit logs? If the answer to either is no, that’s your Q4 priority. Merge the physical and cyber security functions; every robot, vehicle and camera is now an endpoint.

If you’re an investor: The vulnerability count going up sixfold is a demand signal. Look at autonomous remediation, agent identity and governance, formal verification tooling, and memory-safe migration. Avoid anything whose moat is “humans reading alerts.”

If you’re a student: Learn Rust, learn how agents are governed, and learn to read a threat model. The security profession is not shrinking; it’s moving up the stack from patching to architecture, and there are 4 million open seats.

If you’re a parent: The attacks that will reach your family are the ones no rewrite fixes: impersonation, social engineering, stolen credentials. Teach your kids that a voice on the phone can be generated, a password can be leaked, and “verify before you trust” is a life skill now, not an IT policy.

Here’s the question I’ll leave the room of cyber-security experts at GSX with in Atlanta, and I’ll leave it with you too. Fifty years of human-written code gave attackers their playground. The machines can now take most of it away. Are you going to hand them the keys to your own systems first, or wait until someone else’s machines find the door?

To a future of abundance,

Where Indigo landsFurther

Indigo's conclusion

The strongest point applies “verification can't be skipped” to cybersecurity: finding a hole is verifying a patch. But “defenders hold the stronger weapon for the first time” is appealing and fragile, time on loan; rewriting old code in memory-safe languages is the most solid reason for optimism.

What to remember

  1. Wherever attack and defense are the same skill run in reverse, AI arms both sides, and whoever points it at their own systems first gets ahead.
  2. The defenders' edge is using your own AI before someone else's finds you: a race against time, not a steady moat.
  3. Six times, 87%, a perfect ExploitBench score, 70% memory safety and a 45% defect rate are all second-hand; check the original sources before using them.

Back on the long-running theses

confirms

Verification can't be compressed Finding holes and verifying patches are two sides of one thing, and people move from hunting bugs to approving fixes: this view in cybersecurity.

adds to

AI capability is a bounded exponential Memory safety has machine judges, so AI can remove whole classes of bugs: another surge in a narrow, checkable domain.

confirms

Citrini's cybersecurity long/short: enforcement points as moats, vulnerability management commoditized Bears say vulnerability management becomes a commodity; bulls say defenders get a window and a shrinking attack surface. Same cause.

conflicts

Yoshua Bengio, Why are AI agents lying, cheating and coordinating? The same verification: a moat here, a single point of failure that can be broken there.

adds to

Dario Amodei, We Must Pace the Frontier The defenders' edge depends directly on anti-distillation and protecting weights; once weights leak, the gap disappears.

adds to

Dwarkesh on the OpenAI–Hugging Face incident The German wiki case is this summer's third of its kind, with the same lesson: every agent is an insider.

What would change my mind

open models match closed frontier models at cyber capability, or the patch-verification side itself is broken by AI tampering with its scoring.

Finished. Indigo's take on this piece is in two places: