Mind · In / Out · In · 文章

组织级第二大脑:打造能向专家学习的 AI

An Organizational Second Brain: Building an AI That Learns From Experts

Shaurya Sengar、Jason Nawrocki 等 · Engineering at Meta · 2026-09-19

Meta 把机构的持续学习做成「只改文本、不碰权重」的编译流程,正面顶住了「终点在权重」的说法。

Indigo 的结论

在需要审计、能回滚的高风险机构知识领域,终点明确在文本、不在权重:这是治理上的选择,不是能力上的妥协。整个改进飞轮把「验证省不掉」做成了产品:验证不是附属工序,而是这套系统全部的难点和价值。

怎么读这篇 Meta 官方工程博客的系统披露,有秀 AI 工程的用意:结果说法偏软(「几乎总是有用」「零回归」,没有硬数字),只真正做了一个合规领域,却说可以推广。但架构写得具体、可复核,和 Karpathy 的 LLM Wiki、Google 的 Open Knowledge Format 对得上:工程是实的,效果数字要打折。

需要记住的几件事

  1. 知识和推理彻底分开:知识文件没有流程,流程文件没有事实;出错能干净归因,按需加载省下约 80% 的 token。
  2. 验证即编译是整套系统的核心:诊断根因、编译最小改动、盲评加对抗性复核,再进回归测试,收益永久保留。
  3. 「价值转向租不到的东西」又一例:底座模型是大路货,护城河在私有的结构化知识、流程和评估套件。

拆解 · 4 步

  1. 01

    先把专家脑子里的东西写成文件

    真正的知识是隐性的:专家怎么推理、优先什么、怎么消除歧义。他们预先整理出 200 多个带依赖关系的结构化文件,不让 agent 每次从头推。 读这一段原文 →

  2. 02

    「知道什么」和「怎么推理」分开放

    流程文件是多步操作说明,只引用知识、不含事实;知识文件只写立场、不含流程。按需加载,每次消耗的 token 减少约 80%。 读这一段原文 →

  3. 03

    把专家纠错当成编译问题

    检查点让人握着方向盘;每次纠错先找根因,编译成最小的验证过的改动,回放加回归(盲评)、对抗性复核,修好后永久进回归测试。 读这一段原文 →

  4. 04

    组织的记忆放在文本里,不放在权重里

    6 周三轮迭代:专家评价「几乎总是有用」,单次评估从几天降到几分钟,零回归。每次改进都是专家 30 秒能看懂、能回滚的文本修改。 读这一段原文 →

什么会让我改口

前沿实验室把权重层面的持续学习做到能精确回滚、能通过合规审计;或者这套架构搬到别的领域后出现严重退步。

怎么读这篇

Meta 官方工程博客的系统披露,有秀 AI 工程的用意:结果说法偏软(「几乎总是有用」「零回归」,没有硬数字),只真正做了一个合规领域,却说可以推广。但架构写得具体、可复核,和 Karpathy 的 LLM Wiki、Google 的 Open Knowledge Format 对得上:工程是实的,效果数字要打折。

拆解 · 4 步
  1. 先把专家脑子里的东西写成文件
  2. 「知道什么」和「怎么推理」分开放
  3. 把专家纠错当成编译问题
  4. 组织的记忆放在文本里,不放在权重里
01

先把专家脑子里的东西写成文件

真正的知识是隐性的:专家怎么推理、优先什么、怎么消除歧义。他们预先整理出 200 多个带依赖关系的结构化文件,不让 agent 每次从头推。

- 我们构建了一个 AI 智能体,作为特定领域的次级专家,使深厚的专业知识能够被轻易获取和保存,供组织内的任何人访问、共享和进一步拓展。

- 结构化且可审计的知识架构将智能体“所知道的内容”与其“推理的方式”分离开来。

- 随后,一个自我改进循环会在无需重新训练模型的情况下,将专家反馈编译为经过验证且通过回归测试的更新。

- 这两个图层共同将一次性的专家纠错转化为永久的、不断复利的制度记忆;该模式旨在泛化到其他由可检索文本而非模型权重主导的领域。

- 该系统正在为 Meta 的领域业务专家(SME)节省大量时间,使他们能够更加专注于其知识最能发挥价值的工作。

许多大型组织在对待专业知识时都面临着同样的问题。虽然部分知识以模型、剧本、检查清单和框架的形式被记录了下来,但最宝贵的专业知识却存在于人们的大脑中,很少能在任何持久的地方被捕获。例如在合规领域,数百次产品审查中可能会出现同类型的问题,专家的评估需要数天的手动研究,而评估之间的不一致则会带来真正的组织风险。

专家花在回答日常例行问题上的时间,往往比花在真正新颖且模棱两可、最需要其判断力的工作上的时间还要多,这种情况并不少见。我们需要能够捕获组织专家如何进行推理,并将这些知识提供给所有需要之人的系统,从而让专业知识更容易被共享、传承和保存。为了解决这一挑战,我们着手将制度智慧代码化为一个针对特定合规领域的 AI 智能体。该智能体结合了一个扮演组织“第二大脑”角色的知识系统、一个镜像反映领域专家真实思考方式的推理层,以及一个能够将专家的努力永久复利的自动化改进流水线。这些模式可以泛化到任何拥有深厚专业知识的企业领域,无论是金融、安全还是工程。

架构一览

现成的 LLM 提供了坚实的基础,但它们往往需要更深层次的组织上下文,才能在专业领域充分发挥作用。如果没有这种锚定,通用模型的价值将非常有限,因为它们无法区分组织“可以做什么”(通用信息的总结)与“应该考虑做什么”(基于历史立场、公司方向、业务上下文等)。在关键高风险领域,填补这一差距需要向模型提供组织自身的知识和优先级,使其分析能够反映组织真正的推理方式。

我们设计的系统包含四个图层,每个图层解决一个独特的问题:

这些图层相互依赖。知识系统的文件结构使自动化编辑成为可能;推理层的明确程序使失败归因变得可追踪;评估框架对每一次变更进行把关;而改进循环则反哺知识层和推理层。移除其中任何一个图层,其他图层都会随之退化。

构建组织级第二大脑

作为专家工作的副产品,大型组织可以积累数以千计的文档。将这些文档视为组织知识固然很诱人,但真正的知识是隐性的:专家如何推理、他们优先考虑什么,以及他们如何解决模棱两可的问题。一个在推理阶段检索文档块的智能体必须在每次运行时从原始素材中重新推导这种推理,这不仅缓慢、容易出错,而且缺乏一致性。

我们提前将这些隐性知识显性化。一个长期运行的离线进程对源文档进行推理,并将它们精炼为结构化的知识文件——即关于组织如何解释其领域的精选陈述,并将约束条件、边界和路由含义转化为机器可读的形式。

最重要的是,这些知识随后构成了一个反馈循环的基础,使智能体能够学习并实施来自人类专家的反馈,而无需对底层模型进行重新训练。

业界已经趋同于类似的理念。Andrej Karpathy 的 LLM Wiki 将智能体知识构建为可导航的文件图,而 Google 的 Open Knowledge Format 为跨智能体互操作性对此进行了标准化。共同的洞察在于:知识应该被预先提取、明确结构化并逐步披露,而不是在每次查询时重新推导。我们将这些原则扩展到了一个对引用保真度和制度一致性绝不妥协的系统中,将 200 多个文件组织进严格的分类体系中:

- 立场文件(Position files)记录权威的组织立场:组织决定如何解释给定的领域问题,以及其约束条件、边界条件和可由机器执行的路由含义,告诉推理层何时适用该立场。

- 分类与词汇文件(Taxonomy and vocabulary files)作为组织用来描述其领域的术语的权威词汇表,例如实体类型、活动类别和分类层级。每一项都被维护为唯一的单一事实来源,以便智能体和组织使用一致的语言。

- 路由索引(Routing indexes)将输入特征映射到相关的立场和程序,在不单依赖嵌入相似度的情况下确定适用哪些文件。这使得检索具有确定性和可审计性。

- 网关文件(Gateway files)定义了智能体在进入分析领域之前必须通过的门槛测试,防止其在不属于专门知识的地方应用该知识。

每个文件都在 YAML 前置元数据中声明其依赖项(depends_on)和使用者(referenced_by),形成一个双向依赖图。当一个文件发生变化时,你可以精确追踪到还有哪些内容可能会受到影响,这在自我改进循环提出自动化编辑建议时至关重要。

知识系统示意图:文件被组织为一个可导航的文件系统(左),每个文件的 YAML 前置元数据(右)声明了其适用时机(触发场景)及其依赖项和使用者,形成一个智能体可以轻松遍历和维护的双向依赖图。

按密度和使用频率组织知识

一个关键的架构决策在于如何在精选的 wiki 和补充检索(RAG)之间划分知识。我们根据信息密度和预期使用频率来进行切分。

高密度、频繁引用的来源放入 wiki 中:这些是捕获组织如何推理的精炼文件,例如立场、决策框架、边界示例和战略解释。智能体几乎在每一轮交互中都会查阅这些内容。由于它们编码了组织不断演进的思想,因此需要保持最新状态,而 wiki 结构使它们易于更新、版本控制和验证。

稀疏的、视情况相关的来源则通过语义或词法检索(RAG)提供:这些是在适用时非常重要、但在大多数运行中不需要详细了解的文档,例如详细参考资料、单个产品规格、历史决策记录和特定领域的外部知识。将所有这些内容加载到 wiki 中会导致系统膨胀并分散注意力。

结果是,智能体的核心推理始终基于最精炼、最新颖的组织知识,同时在场景需要时仍能获取支持性证据。这种结合产生了一个组织级第二大脑,它编码了组织如何解释和应用信息,而不仅仅是去哪里寻找信息。

02

「知道什么」和「怎么推理」分开放

流程文件是多步操作说明,只引用知识、不含事实;知识文件只写立场、不含流程。按需加载,每次消耗的 token 减少约 80%。

通过可组合配方实现专家推理

单靠知识是不够的。领域专家不仅是回忆事实,他们还遵循结构化的方法论:财务分析师按部就班地梳理估值模型,安全工程师遵循威胁建模程序。挑战在于以 LLM 能够可靠执行的形式捕获这些方法论。

我们通过被称为“配方”(recipes)的可组合程序来解决这个问题。知识文件是声明式的,而配方则是指令式的。每个配方都规定了一个多步骤的分析工作流,指定首先检查什么、每一步要加载什么知识、遵循什么决策程序,以及什么构成了一份完整的分析。

关键的设计选择是将智能体“知道什么”与其“如何推理”分离开来。配方引用知识文件,但本身不包含领域事实;知识文件阐述立场,但不上报程序。这意味着:

- 添加一个组织立场意味着添加一个知识文件并更新路由索引。配方无需修改。

- 修复智能体方法论中的缺陷意味着编辑一个配方。知识文件无需修改。

- 失败可以干净地归因于某一个图层。是知识错了,还是程序错了?

配方组合成流水线,就像主厨针对晚宴的主配方会下发给每个组件(酱汁、蛋白质、装饰)的子配方,而其自身并不包含那些细节一样。我们的顶层路由配方会检查输入,并选择调用哪些下游配方,每个下游配方处理一个分析阶段。

这也是实现渐进式披露的原因。每个配方步骤仅包含与该阶段相关的指令和知识,而不是预先加载覆盖所有可能场景的单体指令集。早期版本使用单个扁平的指令文件,并通过语义检索加载所有来源,在每次运行时将大量相关性参差不齐的文件拉入上下文窗口。在重构为配方驱动的阶段后,每次查询仅触及一小部分有针对性的子集,使每轮消耗的 token 减少了约 80%。上下文窗口是有限的,且注意力会随着容量的增加而退化,因此在恰当的时间提供恰当的指令可以直接提升推理质量。

保持人类掌控

03

把专家纠错当成编译问题

检查点让人握着方向盘;每次纠错先找根因,编译成最小的验证过的改动,回放加回归(盲评)、对抗性复核,修好后永久进回归测试。

人类专家在整个过程中保持对该系统的控制。智能体加速并结构化他们的工作;它不会取代他们的判断,也不会取代他们对结果的权威。

我们通过两种机制来实现这一点:

检查点(Checkpoints)是分析中定义的节点,智能体在继续推进之前在此处呈现其中间推理供专家审查,专家可以确认、纠正或重定向。

升级机制(Escalations)在智能体遇到真正的模棱两可时触发,无论这种不确定性来自未明确说明的输入,还是支持不止一种合理解读的证据。智能体不会强行做出决策,而是将问题交由专家处理,专家的选择将决定分析的方向。

检查点和升级机制同时服务于三个目的:

- 质量与方向控制:专家在错误向下游传递并累积之前捕获它们,并使分析保持在他们认为最相关的轨道上。

- 训练信号:每一次纠错和每一次升级都会成为自我改进循环的输入。

- 信任校准:专家通过观察智能体的推理过程(而非仅仅是最终输出),以及看到它标记不确定性(而非掩盖不确定性),逐步建立信任。

决策越重要,这一点就越关键,这就是为什么我们建议在合规、金融风险评估、安全审查和工程安全等领域默认保持人机回环(Human-in-the-loop)。

自我改进飞轮

我们认为自我改进飞轮是该系统中最具特色的部分。尽管结构化知识系统和可组合配方打造出了一个人类与智能体均可读、可测试且模块化的系统,但大量相互依赖的文件使得人工维护无法实现规模化。当领域专家向智能体提供反馈时,该反馈必须转化为精准的文件编辑。这一过程可能需要数周时间,因为它需要理解整个依赖图,验证没有其他内容被破坏,并确认修复确实有效。

从 RAG 记忆系统到模型权重知识编辑,已有大量工作致力于解决智能体如何存储、检索和更新知识。然而,随着基于文档的制度知识库的扩展以及专家立场的演变,如何保持其准确性所获得的关注要少得多。能够自动草拟自身修复方案的智能体越来越普遍;但我们尚未看到在无需模型重新训练的情况下,将这种级别的验证严谨度应用于结构化知识库。

我们将这种维护视为一个编译问题并将其自动化。每一次专家纠错都会经历四个阶段:

- 将专家反馈诊断为包含根本原因的可操作问题。

- 将问题编译为经过验证的最小化编辑。

- 验证修复有效且无回归问题。

- 由领域专家对其进行审查。

一旦循环完成,刚修复的问题就会被补充到回归测试套件中,以便未来的更新能够保留这一正确行为。

自我改进循环:专家纠错被诊断至根本原因,编译为经过验证的最小化编辑,并在经过审查和落地之前根据重放测试和回归测试进行评估。随后,每次修复都会被并回回归套件中,从而使收益成为永久性的。

诊断:将每次纠错归因于根本原因

原始的专家反馈来自于领域 SME 与智能体交互并提供纠错的对话轨迹。诊断阶段从这些对话中提取结构化的信号。

我们的第一种方法是按对话形式对反馈进行分类。如果专家提供了信息,那必定是知识缺失;如果他们重定向了智能体,那必定是程序问题。这种启发式方法失败了,因为对话形式并不能很好地代表根本原因。专家纠正一个结论,可能是在揭示知识缺失、配方缺陷,或是真正的模棱两可。

行之有效的方法是将提取与分类分离开来。首先,从专家那里提取每一个实质性信号,同时结合智能体的完整知识清单(加载的每个文件、何时加载以及如何使用)。其次,阅读实际的知识文件并应用单一归因测试:智能体能否从其源材料中得出正确的结论?

- 如果素材包含正确答案但智能体仍然出错:配方问题。

- 如果素材未包含正确答案:知识缺失。

- 如果专家本身对正确答案存在分歧:模棱两可,标记供人类讨论。

编译:精密的多智能体编辑

编译器将每个诊断出的问题转化为最小化的文件编辑。子智能体并行分析影响,审查交叉引用、与现有立场的冲突、token 预算影响、测试覆盖率以及重复风险。

两项设计选择使这一过程值得信赖:

独立的对抗性审查。一个单独的智能体在一个全新的上下文中运行,对改进原理一无所知,仅接收对知识库拟议的差异(diffs)。它的任务是发现引入矛盾、破坏边缘情况或削弱已有立场等问题。因为它与提出修改的智能体不共享上下文,所以不会继承它们的盲区。

确定性的结构校验。Linter 通过程序化方式捕获问题,如悬空的交叉引用、违反文件大小预算、标识符冲突以及循环依赖。这一层不是概率性的,只有通过或失败。

评估:证明修复有效

每一个拟议的变更都会经历两阶段的验证:

针对性重放(Targeted replay)在触发反馈的原始场景上运行智能体。智能体并不知道自己正在接受测试。一个独立的裁判智能体在不知道修改了什么的情况下,根据原始专家反馈对新的输出进行评估。这种刻意设计的盲审阻止了确认偏误。如果针对性重放失败,将重试编译。

回归测试(Regression testing)为该领域运行多个基准测试,这些基准测试通常是由问答对组成的结构化测试套件。对于可能存在多个正确答案的分析领域,独立的 LLM 裁判会根据特定标准为每个测试用例判定通过/失败。智能体在独立的会话中针对基准问题并行运行,并检测性能是否存在回归。如果回归测试失败,将使用更新后的提示词重试编译,该提示词描述了智能体在哪里发生了回归,以及原始问题和尝试过的修复方案。

落地与丰富:产生复利回报

该流水线的输出是一个附带完整审计追踪的 Pull Request(diff)。人类专家审查的是一个经过证明有效的修复,而不是去调试原始故障。一旦获得批准并落地(知识文件或配方完成更新),原始的失败场景及其经过验证的正确答案就会自动添加到回归测试套件中。这意味着每次修复都会永久提高门槛,未来对知识系统的改变必须保留刚刚纠正的行为。

04

组织的记忆放在文本里,不放在权重里

6 周三轮迭代:专家评价「几乎总是有用」,单次评估从几天降到几分钟,零回归。每次改进都是专家 30 秒能看懂、能回滚的文本修改。

结果

经过历时六周的三个开发迭代(sprints),该系统达到了以下成果:

- 领域 SME 评价智能体的输出几乎始终有用,相比早期版本中输出频繁需要大量重做的情况有了显著改善。

- 单项评估时间从数天缩短至数分钟。

- 自动化自我改进生成经过验证的知识编辑,其速率达到了此前需要整个工程迭代才能完成的水平。

- 在多个改进周期中实现零回归,每次修复都自动强化了回归测试套件。

- 领域专家一致反馈智能体承担了绝大部分分析工作,使他们能够专注于真正需要人类判断的模棱两可案例。

应用这一架构

我们为其构建该系统的特定领域需要将数十个来源(包括内部立场和外部材料)综合为风险加权的评估。但该架构是独立于领域的。它适用于以下任何场景:

- 专业知识作为隐性知识存在于专家的脑海中。

- 评估之间的一致性至关重要。

- 工作量超过了现有的专家能力范围。

- 现成的 LLM 产生的分析不够充分。

适合这种模式的具体领域包括合规监管、协议遵循、金融风险评估、安全审查、工程标准合规和采购评估。共同的主线是:组织需要拥有真正制度专业知识的 AI 系统,而不仅仅是通用知识。

采用这种架构的要求包括:

- 一个具有明确文件边界、交叉引用和依赖图的结构化知识系统(该领域的组织级第二大脑)。

- 一个将领域知识与分析方法论(配方)分离开来的程序层。

- 一个随着每个改进周期不断增长的自动化评估套件。

- 根据领域的风险承受能力进行校准的人机回环(Human-in-the-loop)检查点。

更深层次的原则非常直接:将复杂性保持在人类和智能体均可读的文本文件中,而不是微调后的模型权重中。每一次改进都是一次文本编辑,领域专家可以在 30 秒内进行审查。每一次变更都是经过版本控制的、可进行差异对比的且可逆的。编译流水线固然复杂,但其输出始终保持透明。

我们的目标是建立一个能够让专家的努力实现永久复利的系统。专家的每一次交互都会使系统变得更好。每一次纠错都会作为经过验证的改进持久保留下来。组织的集体知识不再被困于个体身上,而是开始以一致且具备规模效益的方式提供给每一个需要它的人。

判断收口延伸

Indigo 的结论

在需要审计、能回滚的高风险机构知识领域,终点明确在文本、不在权重:这是治理上的选择,不是能力上的妥协。整个改进飞轮把「验证省不掉」做成了产品:验证不是附属工序,而是这套系统全部的难点和价值。

需要记住的几件事

  1. 知识和推理彻底分开:知识文件没有流程,流程文件没有事实;出错能干净归因,按需加载省下约 80% 的 token。
  2. 验证即编译是整套系统的核心:诊断根因、编译最小改动、盲评加对抗性复核,再进回归测试,收益永久保留。
  3. 「价值转向租不到的东西」又一例:底座模型是大路货,护城河在私有的结构化知识、流程和评估套件。

放回主线

冲突

持续学习的终局在权重:直觉大于记忆 Meta 在机构知识上选择不碰权重、存可审计的文本:权重是模型直觉层的终点,可检索、可撤回的组织知识,终点在文本。

证实

验证不可压缩:生成归零后,验证成为瓶颈 整个改进飞轮就是把验证做成产品:盲评、对抗复核、回归测试,是没有便宜终审的领域里最工程化的实现。

证实

你拥有的不是模型:价值上移到不可租用的东西 现成的大模型没有机构的上下文;这套系统真正值钱的,是私有的结构化知识、流程和评估套件。

证实

Ashwin Gopinath:引擎已商品化,记忆才是护城河,而记忆是编译器不是数据库 Meta 把知识维护当编译问题来做,「记忆是编译器不是数据库」被字面兑现。

补充

Tara Seshan:知识工作的验证难题是编码没有的 Tara 说知识工作没有机器终审;Meta 的答案是自己造一套,再把人在回路里写进制度。

补充

Ivan Zhao(Notion CEO)钢铁蒸汽与无限心智 Ivan 点的两个卡点是上下文散落和无法验证;Meta 用结构化知识和验证飞轮,给了企业内部的一版实现。

什么会让我改口

前沿实验室把权重层面的持续学习做到能精确回滚、能通过合规审计;或者这套架构搬到别的领域后出现严重退步。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Essay

An Organizational Second Brain: Building an AI That Learns From Experts

Shaurya Sengar, Jason Nawrocki et al. · Engineering at Meta · 2026-09-19

Meta turns institutional continual learning into a compile process that edits text and never touches weights, a direct challenge to “it all ends in the weights”.

Indigo's conclusion

For high-stakes institutional knowledge that must be auditable and reversible, the end state is clearly text, not weights: a governance choice, not a capability compromise. The flywheel turns “verification can't be skipped” into the product itself: verification isn't a side step, it is the whole difficulty and the whole value.

How to read this A system write-up on Meta's official engineering blog, partly a showcase: the results are soft (“almost always useful”, “zero regressions”, no hard numbers), and only one compliance domain was actually built while generality is claimed. But the architecture is specific and checkable, and it matches Karpathy's LLM Wiki and Google's Open Knowledge Format: the engineering is real; discount the results.

What to remember

  1. Knowledge and reasoning fully separated: no procedures in knowledge files, no facts in recipes; failures trace cleanly, and loading on demand saves about 80% of tokens.
  2. Verification as compilation is the core: diagnose the root cause, compile the smallest edit, blind evaluation plus adversarial review, then into the regression suite, so gains are kept for good.
  3. Another case of value moving to what can't be rented: the base model is a commodity; the moat is private structured knowledge, recipes and evaluation suites.

Breakdown · 4 steps

  1. 01

    First, write down what's in experts' heads

    Real knowledge is tacit: how experts reason, what they prioritize, how they resolve ambiguity. They pre-extract 200+ structured files with dependency links so agents don't re-derive it every time. Read this part →

  2. 02

    Keep “what we know” apart from “how we reason”

    Recipes are step-by-step procedures that reference knowledge but contain no facts; knowledge files state positions with no procedure. Loading only what's needed cuts tokens by about 80%. Read this part →

  3. 03

    Expert corrections treated as a compile problem

    Checkpoints keep people at the wheel. Each correction is traced to a root cause, compiled into the smallest verified edit, replayed and regression-tested blind, reviewed adversarially, then added to the regression suite for good. Read this part →

  4. 04

    Organizational memory lives in text, not weights

    Six weeks, three sprints: experts rate the output “almost always useful”, a single review drops from days to minutes, zero regressions. Every improvement is a text edit an expert can review in 30 seconds and roll back. Read this part →

What would change my mind

frontier labs make weight-level continual learning precisely reversible and audit-ready, or this architecture regresses badly when moved to other domains.

How to read this

A system write-up on Meta's official engineering blog, partly a showcase: the results are soft (“almost always useful”, “zero regressions”, no hard numbers), and only one compliance domain was actually built while generality is claimed. But the architecture is specific and checkable, and it matches Karpathy's LLM Wiki and Google's Open Knowledge Format: the engineering is real; discount the results.

Breakdown · 4 steps
  1. First, write down what's in experts' heads
  2. Keep “what we know” apart from “how we reason”
  3. Expert corrections treated as a compile problem
  4. Organizational memory lives in text, not weights
01

First, write down what's in experts' heads

Real knowledge is tacit: how experts reason, what they prioritize, how they resolve ambiguity. They pre-extract 200+ structured files with dependency links so agents don't re-derive it every time.

- We’ve built an AI agent that acts as a secondary expert for a given domain, making deep specialist knowledge readily available and preserved for anyone in an organization to access, share, and build upon.

- A structured, auditable knowledge architecture separates what the agent knows from how it reasons.

- A self-improvement loop then compiles expert feedback into verified, regression-tested updates without model retraining.

- Together, these two layers turn one-off expert corrections into permanent, compounding institutional memory, and the pattern is designed to generalize to other domains governed by retrievable text rather than model weights.

- This system is saving domain subject matter experts (SME)s at Meta substantial time, allowing them to focus more on the work where their knowledge matters most.

Many large organizations have the same problem when it comes to specialist knowledge. While some of it is written down in the form of models, playbooks, checklists, and frameworks, the most valuable specialist knowledge lives in people’s heads and rarely gets captured anywhere durable. In compliance domains, for example, the same types of questions can arise across hundreds of product reviews, expert assessments take days of manual research, and inconsistency between assessments creates real organizational risk.

It’s not uncommon for experts to spend more time answering routine questions than on genuinely novel and ambiguous work where their judgment matters most. We need systems that can capture how an organization’s experts reason and make that knowledge available to everyone who needs it, so that expertise is easier to share, build on, and preserve. We set about solving this challenge by codifying institutional intelligence into an AI agent for a specific compliance domain. The agent combines a knowledge system that acts as the organization’s “second brain,” a reasoning layer that mirrors how domain experts actually think, and an automated improvement pipeline that compounds expert effort permanently. The patterns generalize to any enterprise domain with deep specialist knowledge, whether that is finance, security, or engineering.

The Architecture at a Glance

Off-the-shelf LLMs provide a strong foundation, but they often need deeper institutional context to be fully effective in specialist domains. Without that grounding, a general purpose model has limited value given it will not be able to distinguish between what an organization could do (a summary of general information) and what it should consider doing (based on historic positions, company direction, business context, etc.). In high-stakes domains, closing this gap requires supplying the model with the organization’s own knowledge and priorities so its analysis reflects how the organization actually reasons.

The system we’ve designed has four layers, each solving a distinct problem:

These layers depend on each other. The knowledge system’s file structure makes automated editing possible. The reasoning layer’s explicit procedures make failure attribution tractable. The evaluation framework gates every change. And the improvement loop feeds back into both knowledge and reasoning. Remove any one layer and the others degrade.

Building the Organizational Second Brain

Large organizations can accumulate thousands of documents as a byproduct of expert work. It is tempting to treat those documents as organizational knowledge, but the real knowledge is implicit: how experts reason, what they prioritize, and how they resolve ambiguity. An agent that retrieves document chunks at inference time has to re-derive that reasoning from raw sources on every run, which is slow, error-prone, and inconsistent.

We make that implicit knowledge explicit ahead of time. A long-running offline process reasons through source documents and distills them into structured knowledge files – curated statements of how the organization interprets its domain, with constraints, boundaries, and routing implications made machine-readable.

Most significantly, that knowledge then forms the basis of a feedback loop that allows the agent to learn from and implement feedback from human experts without the underlying model having to be retrained.

The industry has converged on a similar idea. Andrej Karpathy’s LLM Wiki structures agent knowledge as a navigable graph of files, and Google’s Open Knowledge Format standardizes this for cross-agent interoperability. The shared insight is that knowledge should be pre-extracted, explicitly structured, and progressively disclosed rather than re-derived on every query. We extended these principles into a system where citation fidelity and institutional consistency are non-negotiable, organizing 200+ files into a strict taxonomy:

- Position files capture authoritative organizational stances: how the organization has decided to interpret a given domain question, along with its constraints, boundary conditions, and machine-actionable routing implications that tell the reasoning layer when to apply it.

- Taxonomy and vocabulary files act as an authoritative glossary for the terms the organization uses to describe its domain, such as entity types, activity categories, and classification tiers. Each is maintained as a single source of truth so the agent and the organization use language consistently.

- Routing indexes map input characteristics to the relevant positions and procedures, determining which files apply without relying on embedding similarity alone. This makes retrieval deterministic and auditable.

- Gateway files define threshold tests the agent must pass before entering an analytical domain, preventing it from applying specialized knowledge where it does not belong.

Every file declares its dependencies (depends_on) and consumers (referenced_by) in YAML frontmatter, forming a bidirectional dependency graph. When one file changes, you can trace exactly what else might be affected, which matters when the self-improvement loop proposes automated edits.

An illustration of the knowledge system: files are organized as a navigable filesystem (left), and each file’s YAML frontmatter (right) declares when it applies (the triggering scenarios) plus its dependencies and consumers, forming a bidirectional dependency graph the agent can traverse and maintain easily.

Organizing Knowledge by Density and Usage Frequency

A key architectural decision is how to partition knowledge between the curated wiki and supplementary retrieval (RAG). We split on information density and expected usage frequency.

High-density, frequently referenced sources go into the wiki: Distilled files capturing how the organization reasons, such as positions, decision frameworks, boundary examples, and strategic interpretations. The agent consults these on nearly every turn. Because they encode the organization’s evolving thinking, they need to stay current, and the wiki structure makes them easy to update, version, and validate.

Sparse, situationally relevant sources are served through semantic or lexical search (RAG): documents that matter deeply when they apply but are not needed in detail on most runs, such as detailed reference material, individual product specifications, historical decision records, and niche external knowledge. Loading all of them into the wiki would bloat the system and dilute attention.

The result is that the agent’s core reasoning is always grounded in the most refined, current organizational knowledge, while it can still reach for supporting evidence when a scenario demands it. The combination produces an organizational second brain that encodes how the organization interprets and applies information, rather than only where to find it.

02

Keep “what we know” apart from “how we reason”

Recipes are step-by-step procedures that reference knowledge but contain no facts; knowledge files state positions with no procedure. Loading only what's needed cuts tokens by about 80%.

Expert Reasoning via Composable Recipes

Knowledge alone is not enough. Domain experts do not simply recall facts, they follow structured methodologies: a financial analyst works through a valuation model step by step, a security engineer follows a threat modeling procedure. The challenge is capturing those methodologies in a form an LLM can execute reliably.

We solve this with composable procedures we call recipes. Where knowledge files are declarative, recipes are imperative. Each one prescribes a multi-step analytical workflow, specifying what to examine first, which knowledge to load at each step, what decision procedures to follow, and what constitutes a complete analysis.

The critical design choice is separating what the agent knows from how it reasons. Recipes reference knowledge files but contain no domain facts; knowledge files state positions but prescribe no procedures. This means:

- Adding an organizational position means adding a knowledge file and updating a routing index. No recipe changes.

- Fixing a flaw in the agent’s methodology means editing a recipe. No knowledge files change.

- Failures attribute cleanly to one layer. Was the knowledge wrong, or the procedure?

Recipes compose into pipelines, much like a head chef’s master recipe for a dinner service delegates to sub-recipes for each component (the sauce, the protein, the garnish) without containing those details itself. Our top-level routing recipe examines the input and selects which downstream recipes to invoke, each handling one analytical phase.

This is also what enables progressive disclosure. Rather than front-loading a monolithic instruction set covering every possible scenario, each recipe step carries only the instructions and knowledge relevant to that phase. Early versions used a single flat instruction file and loaded all sources via semantic search, pulling a large volume of mixed-relevance files into the context window on every run. After restructuring into recipe-driven stages, each query touches only a small, targeted subset, cutting tokens consumed per turn by around 80%. Context windows are finite and attention degrades with volume, so delivering the right instructions at the right time directly improves reasoning quality.

Keeping Humans in Control

03

Expert corrections treated as a compile problem

Checkpoints keep people at the wheel. Each correction is traced to a root cause, compiled into the smallest verified edit, replayed and regression-tested blind, reviewed adversarially, then added to the regression suite for good.

Human experts stay in control of this system throughout. The agent accelerates and structures their work; it does not replace their judgment or their authority over the outcome.

We enforce this through two mechanisms:

Checkpoints are defined points in the analysis where the agent surfaces its intermediate reasoning for expert review before proceeding, and the expert can confirm, correct, or redirect.

Escalations trigger when the agent hits genuine ambiguity, whether from underspecified inputs or evidence that supports more than one defensible reading. Rather than forcing a resolution, it hands the question to the expert, whose choice determines the path the analysis takes.

Checkpoints and escalations serve three purposes simultaneously:

- Quality and direction control: Experts catch errors before they compound downstream, and keep the analysis on the path they consider most relevant.

- Training signal: Every correction and every escalation becomes input for the self-improvement loop.

- Trust calibration: Experts build confidence incrementally by observing the agent’s reasoning rather than just its final output, and by seeing it flag uncertainty instead of masking it.

The more consequential the decision, the more this matters, which is why we recommend keeping a human in the loop by default across domains like compliance, financial risk assessment, security review, and engineering safety.

The Self-Improvement Flywheel

We consider the self-improvement flywheel to be the most distinctive part of this system. While the structured knowledge system and composable recipes produce a system that is legible by humans and agents, testable, and modular, the number of interdependent files make manual maintenance impossible to scale. When domain experts provide feedback to an agent that feedback has to be translated into precise file edits. This process can take weeks because it requires understanding the full dependency graph, verifying nothing else breaks, and validating the fix actually works.

There has been a large body of work – from RAG memory systems to model-weight knowledge editing – devoted to addressing how agents store, retrieve, and update knowledge. But much less attention has gone toward keeping a document-based institutional knowledge base correct as it grows and as expert positions evolve. Agents that auto-draft their own fixes are increasingly common; but we haven’t seen this level of validation rigor applied to a structured knowledge base without model retraining.

We treat that maintenance as a compilation problem and automate it. Every expert correction moves through four phases:

- Diagnose expert feedback into actionable issues with their root cause.

- Compile issues into minimal verified edits.

- Validate that fixes work without regressions.

- Have domain experts review them.

Once the loop completes, the regression test suite is enriched with the issue that was just fixed so that future updates preserve this behavior.

The self-improvement loop. Expert corrections are diagnosed to their root cause, compiled into minimal verified edits, and evaluated against replay and regression tests before they are reviewed and landed. Each fix is then folded back into the regression suite, so the gain is permanent.

Diagnosis: Attributing Each Correction to a Root Cause

Raw expert feedback comes from conversation traces where domain SMEs interacted with the agent and provided corrections. The diagnosis phase extracts structured signals from these conversations.

Our first approach classified feedback by conversational form. If the expert provided information, it must be a knowledge gap; if they redirected the agent, it must be a procedure problem. This heuristic failed because conversational form is a poor proxy for root cause. An expert correcting a conclusion might be exposing a knowledge gap, a recipe flaw, or a genuine ambiguity.

The working approach separates extraction from classification. First, extract every substantive signal from the expert alongside the agent’s full knowledge manifest (every file loaded, when, and how used). Second, read the actual knowledge files and apply a single attribution test: Could the agent have reached the correct conclusion from its source materials?

- If the materials contained the right answer but the agent still erred: recipe problem.

- If the materials did not contain the right answer: knowledge gap.

- If experts themselves disagree on the right answer: ambiguity, flagged for human discussion.

Compilation: Surgical Multi-Agent Edits

The compiler translates each diagnosed issue into minimal file edits. Sub-agents analyze impact in parallel, examining cross-references, conflicts with existing positions, token budget impact, test coverage, and duplication risk.

Two design choices make this trustworthy:

Independent adversarial review. A separate agent, running in a fresh context with no knowledge of the improvement rationale, receives only the proposed diffs to the knowledge base. Its job is to find problems such as contradictions introduced, edge cases broken, or positions undermined. Because it shares no context with the proposing agents, it cannot inherit their blind spots.

Deterministic structural validation. A linter catches issues programmatically, dangling cross-references, file size budget violations, identifier collisions, and dependency cycles. This layer is not probabilistic. It passes or fails.

Evaluation: Proving the Fix Works

Every proposed change goes through a two-stage validation:

Targeted replay runs the agent on the original scenario that triggered the feedback. The agent does not know it is being tested. A separate judge evaluates the new output against the original expert feedback without knowing what was changed. This deliberately blind design prevents confirmation bias. If targeted replay fails, compilation is retried.

Regression testing runs multiple benchmarks for that domain, which are usually structured test suites of Q&A pairs. For analytical domains whether there might be multiple correct answers, an independent LLM judge gives each test case a pass/fail based on certain criteria. The agent is run in parallel, independent sessions against the benchmark questions, and regressions in performance are detected. If regression testing fails, compilation is retried with an updated prompt describing where the agent regressed, along with the original issue and attempted fix.

Landing and Enrichment: Compounding Returns

The output of the pipeline is a pull request (diff) with a complete audit trail. A human expert reviews a proven fix rather than debugging a raw failure. Once approved and landed (the knowledge file or recipe is updated), the original failing scenario and its validated correct answer are automatically added to the regression test suite. This means every fix permanently raises the bar and future changes to the knowledge system must preserve the behavior that was just corrected.

04

Organizational memory lives in text, not weights

Six weeks, three sprints: experts rate the output “almost always useful”, a single review drops from days to minutes, zero regressions. Every improvement is a text edit an expert can review in 30 seconds and roll back.

Results

After three development sprints spanning six weeks, the system achieved:

- Domain SMEs rated agent outputs useful almost all the time, a significant improvement from early versions where outputs frequently required substantial rework.

- Days to minutes reduction in individual assessment time.

- Automated self-improvement producing validated knowledge edits at a rate that previously required full engineering sprints.

- Zero regressions across improvement cycles, with every fix automatically strengthening the regression suite.

- Domain experts consistently reported the agent handles the vast majority of the analytical work, allowing them to focus on the genuinely ambiguous cases that require human judgment.

Applying This Architecture

The specific domain we built this for required synthesizing dozens of sources (both internal positions and external material) into risk-weighted assessments. But the architecture is domain-independent. It applies wherever:

- Specialist knowledge lives as tribal knowledge in experts’ heads.

- Consistency across assessments matters.

- The volume of work exceeds available expert capacity.

- Off-the-shelf LLMs produce inadequate analysis.

Concrete domains where this pattern fits include regulatory compliance, protocol adherence, financial risk assessment, security review, engineering standards compliance, and procurement evaluation. The common thread is that organizations need AI systems with genuine institutional expertise, not just general knowledge.

The requirements for adopting this architecture are:

- A structured knowledge system with explicit file boundaries, cross-references, and a dependency graph (the organizational second brain for the domain).

- A procedural layer that separates domain knowledge from analytical methodology (recipes).

- An automated evaluation suite that grows with each improvement cycle.

- Human-in-the-loop checkpoints calibrated to the domain’s risk tolerance.

The deeper principle is straightforward: keep the complexity in text files that are readable by both humans and agents, rather than fine-tuned model weights. Every improvement is a text edit that a domain expert can review in 30 seconds. Every change is version-controlled, diffable, and reversible. The compilation pipeline is sophisticated, but its outputs are always transparent.

The goal is a system where expert effort compounds permanently. Every expert interaction makes the system better. Every correction persists as a verified improvement. The organization’s collective knowledge stops being trapped in individuals and starts being available, consistently and at scale, to everyone who needs it.

Where Indigo landsFurther

Indigo's conclusion

For high-stakes institutional knowledge that must be auditable and reversible, the end state is clearly text, not weights: a governance choice, not a capability compromise. The flywheel turns “verification can't be skipped” into the product itself: verification isn't a side step, it is the whole difficulty and the whole value.

What to remember

  1. Knowledge and reasoning fully separated: no procedures in knowledge files, no facts in recipes; failures trace cleanly, and loading on demand saves about 80% of tokens.
  2. Verification as compilation is the core: diagnose the root cause, compile the smallest edit, blind evaluation plus adversarial review, then into the regression suite, so gains are kept for good.
  3. Another case of value moving to what can't be rented: the base model is a commodity; the moat is private structured knowledge, recipes and evaluation suites.

Back on the long-running theses

conflicts

Continual learning ends in the weights: intuition over memory Meta keeps institutional knowledge out of the weights, in auditable text: weights are the end state for model intuition; retrievable, reversible organizational knowledge ends in text.

confirms

Verification can't be compressed: once generation is free, verification is the bottleneck The flywheel is verification made into a product: blind evaluation, adversarial review, regression tests. The most engineered implementation yet where no cheap judge exists.

confirms

You don't own the model: value moves to what can't be rented Off-the-shelf models lack institutional context; what's valuable here is private structured knowledge, recipes and evaluation suites.

confirms

Ashwin Gopinath: engines are commodities, memory is the moat, and memory is a compiler, not a database Meta treats knowledge maintenance as a compile problem: “memory is a compiler, not a database” made literal.

adds to

Tara Seshan: knowledge work has a verification problem coding doesn't Tara says knowledge work has no machine judge; Meta's answer is to build one and write people-in-the-loop into the process.

adds to

Ivan Zhao (Notion CEO): Steel, Steam and Infinite Minds Ivan's two bottlenecks are scattered context and verifiability; Meta's structured knowledge plus verification flywheel is one in-house implementation.

What would change my mind

frontier labs make weight-level continual learning precisely reversible and audit-ready, or this architecture regresses badly when moved to other domains.

Finished. Indigo's take on this piece is in two places: