Mind · In / Out · In · 文章

揭秘 Jev 的架构

Jev’s Architecture Unmasked

Archer Hume · archerhume.com · 2026-09-17

把「让大模型自报 90% 确定」这个行业烂习惯,换成用真实结果训练出的概率;这篇逆向本身也是一堂校准课。

Indigo 的结论

Jev 是「验证省不掉」的商业化答案:不压掉验证,而是把对着真实结果校准过的决策信号做成产品来卖。但它只在有真实结果可以回训的快决策领域成立,没有终审的那一半仍然空着。

怎么读这篇 独立的逆向工程长文,作者和 TypeSafe 没有明显利益关系。证据强度要分层看:他自己的探测实验是一手的、可验证,证据包也公开了;对 TypeSafe 真实架构的结论是从黑箱外推的,他本人反复标为推测;TypeSafe 的产品说法是厂商自报。

需要记住的几件事

  1. 它解决两件事:决策信号不可靠,加上生成这个信号的算力白白浪费。
  2. 读出的概率不等于描述概率的文字:生成的「91%」是一串字,分类器的 0.91 是预测分布里的一项;两者都可能没校准。
  3. 架构逆推(作者标为推测):因果 transformer,很可能是稀疏混合专家,共享内容的缓存加上互相隔离的问题。
  4. RLCD 是为校准决策做的强化学习,目标是给出诚实的概率;MMLU 1200 题的校准误差 0.031。

拆解 · 7 步

  1. 01

    生成的「90% 确定」不是概率

    欺诈筛查、审核、路由、风控都建在这种做法上:既为逐字生成付钱,又把没验证过的置信声明,当成软件能照着行动的概率。 读这一段原文 →

  2. 02

    有了概率,下游规则才算得清账

    误升级成本 1,漏掉紧急事件成本 9,规则就是紧急概率超过 0.1 时升级;只有概率可靠,这笔账才算得成。 读这一段原文 →

  3. 03

    推理靠读出结束,不靠逐字生成

    output_tokens 听着像生成记录,其实是计费数字:200 个选项和 2 个选项返回一样快,别拿它除以时长当生成速度。 读这一段原文 →

  4. 04

    共享内容只算一次,问题之间互不相见

    暗号放在另一个问题里,这个问题问它得 0.00;放进共享内容,升到 0.90 到 0.92;跨接口挪动一句话,作用就变了。 读这一段原文 →

  5. 05

    选项被当成一张列表一起读

    加一个无关的「坏天气」选项,原来两项的对数几率从 0.38 掉到 0.11,十组随机实验每组都降,「每项独立打分」被证伪。 读这一段原文 →

  6. 06

    便宜的概率,也可能是坏概率

    校准问的是标 0.8 的那批里是否真有约 80% 答对;MMLU 1200 题校准误差 0.0313,但 990 题挤在 0.9 到 1.0 区间。 读这一段原文 →

  7. 07

    每条推断都标明证据强度

    问题分支被当成独立任务成批处理;结尾把直接输出概率(公开)、问题隔离(可观察)、共享缓存和稀疏专家(推断)分成三层。 读这一段原文 →

什么会让我改口

在没有真实结果可训练的领域,这种校准概率也能做得可信;那就不是把「验证省不掉」切成两半,而是把它压掉了。

怎么读这篇

独立的逆向工程长文,作者和 TypeSafe 没有明显利益关系。证据强度要分层看:他自己的探测实验是一手的、可验证,证据包也公开了;对 TypeSafe 真实架构的结论是从黑箱外推的,他本人反复标为推测;TypeSafe 的产品说法是厂商自报。

拆解 · 7 步
  1. 生成的「90% 确定」不是概率
  2. 有了概率,下游规则才算得清账
  3. 推理靠读出结束,不靠逐字生成
  4. 共享内容只算一次,问题之间互不相见
  5. 选项被当成一张列表一起读
  6. 便宜的概率,也可能是坏概率
  7. 每条推断都标明证据强度
01

生成的「90% 确定」不是概率

欺诈筛查、审核、路由、风控都建在这种做法上:既为逐字生成付钱,又把没验证过的置信声明,当成软件能照着行动的概率。

X 上充斥着关于 Jev 发布的热辣点评,而绝大多数完全没有击中要害:“一个 JSON 分类器能有 1200 万浏览量?好吧,我们确实处于泡沫之中。”普通 LLM 只是将“90% 把握”作为文本生成出来;它产生这几个词的概率,并不等于它有 90% 的概率是正确的。然而,我们构建欺诈筛查、内容审核、工单路由和风险评估时,却正是围绕着这种模式进行的:按 Token 逐字生成的模式付费,然后将未经验证的置信度声明当作软件可以据以操作的概率。

Jev 的主张是保留预训练 LLM 的知识,同时用直接从其内部表示中读取的决策概率来替代生成的置信度声明。这些概率是针对实际结果训练出来的。只需向它提供共享状态、问题和允许的答案;它就会并行返回概率分布,而不生成任何文本。1 对于这整类应用而言,这同时解决了两个问题:决策信号的可靠性,以及用于生成该信号的无谓计算成本。

唯一的麻烦在于,它并非开源权重,且 TypeSafe 拒绝分享他们的研究……因此,我将(尽我所能地)来做这件事。

种种证据指向一个重新用于决策的因果 Transformer(可能采用了稀疏 MoE):共享状态编码、隔离的问题分支,以及直接的概率读取而非文本生成。在对 TypeSafe API 进行探针测试(寻找不同上下文长度、问题重排序等条件下的延迟缩放特征)、使用 Astra 翻遍所有公开文档和研究,并寻找先验技术之后,我认为自己对它的运行机制和架构有了相当准确的了解。

稀疏主干网络是确定性最低的部分,但它对该领域尤为有利,且相比自回归 LLM 缺陷更少,因此如果不是这样反而会令人意外。共享计算和直接概率输出显然有强得多的证据支持。这一切显然带有相当大的推测性质,因此我会尽量明确区分哪些证据是由 TypeSafe 公布的、哪些是在实验中观察到的,以及哪些是从中推断出来的。黑盒 API 让人能够以惊人的轻松程度描绘出隐藏事物的轮廓,并掌握其架构的大致形状。

原文此处有图 · 去原文看图
图 1。设想的决策模型计算过程。消息被一次性编码至绿色网格中,代表在每个 Transformer 层保留的状态信息。每个彩色问题网格并行构建,将其自身文本和允许的答案与对共享状态的注意力结合起来。问题之间无法互相注意力。针对实际结果训练的读取模块将它们的最终表示直接转化为答案概率,而不生成文本。移动的单元格仅展示信息流向,而非字面上的复制;颜色、注意力路径和概率均为示意图,非 Jev 的实际测量值。
02

有了概率,下游规则才算得清账

误升级成本 1,漏掉紧急事件成本 9,规则就是紧急概率超过 0.1 时升级;只有概率可靠,这笔账才算得成。

为什么这种设计很实用¶

考虑一个示意性的客服工单路由请求。这展示了 API 的结构;以下示例概率均为虚构。

一个有用的答案可能会赋予 payments(支付问题)0.91 的概率,而赋予 urgent escalation(紧急升级)仅 0.42 的概率。这是不同的不确定性。软件可以自动路由该工单,同时将升级交给独立的策略来处理。

因果 Transformer 已经懂得如何从左到右构建文本的表示。在普通的语言模型推理过程中,它处理提示词,预测一个 Token,将该 Token 喂回输入,然后重复此过程。这建立在原始 Transformer 中引入的解码器注意力和输出投影的基础之上。3 但提示词处理阶段已经输出了丰富的表示。如果任务是在三个队列中做出选择,我们可以附加一个小型函数,将该表示直接映射为三个数值。

这里的一个细节化解了关于并行答案的许多困惑:因果注意力描述的是哪些位置可以使用哪些信息,而不是输入 Token 必须被执行的顺序。在提示词处理(或预填充/prefill)阶段,每一个输入 Token 都是已知的。模型可以在同一层内共同处理它们的位置,同时注意力掩码会阻止访问后面的位置;各层依然按顺序运行。自回归解码则增加了另一种依赖:在选中前一个预测之前,下一个 Token 并不存在。我们设想的模型在 prefill 和 readout 之后即告结束,因此避免了这种逐 Token 的依赖。

这改变了任务的计算形态。输出不再需要针对 "payments": 0.91 进行一连串的拼写决策。JSON 格式化发生在普通的应用程序代码中,神经网络只需提供概率。

现在假设状态是一份冗长的事件报告,并且有 50 个问题。大部分输入是共享的。Transformer 会将其处理过的 Token 的中间信息存储在其键值缓存(通常简称为 KV cache)中。在设想的设计中,每个问题都会读取相同的状态缓存。每个分支仅添加其自身的指令和答案选项。

对于包含 SSS 个 Token 的状态和 QQQ 个问题,单独的请求将大约对该状态处理 QQQ 次。共享机制将重复的状态 Token 处理过程从 QSQSQS 减少至 SSS。问题依然需要对状态进行注意力计算;这部分工作并没有消失。但模型无需重复重构状态的表示。

隔离还赋予了该接口实用的含义。询问客户是否愤怒不应改变哪个队列接收该工单。两个问题都可以检查相同的证据,而无需读取彼此的指令。各分支之间在计算上没有任何依赖关系,即使它们的答案在统计上存在关联。

最后,概率使得下游策略变得清晰明确。如果一次不必要的升级带来 1 个单位的成本,而漏掉一次紧急情况带来 9 个单位的成本,那么简化后的决策规则会在 p(urgent)>0.1p(\text{urgent}) > 0.1p(urgent)>0.1 时进行升级。只有当概率对于该工作流而言足够可靠时,这种计算才有意义。因此,对概率分布进行训练和评估构成了产品本身的一部分,而非仅仅是一个装饰性的置信度字段。

这一切都不需要扩散模型(diffusion)。并行分类技术已经存在了几十年。真正有趣的组合在于:通用能力强大的 Transformer、共享的上下文计算、强类型的输出接口,以及奖励有用不确定性的训练方式。

03

推理靠读出结束,不靠逐字生成

output_tokens 听着像生成记录,其实是计费数字:200 个选项和 2 个选项返回一样快,别拿它除以时长当生成速度。

1. 以读取模块(readout)结束推理¶

第一个组件最简单:用预测头(prediction head)替代解码循环。

已公布的证据。TypeSafe 的发布公告指出:“Jev 会并行输出所有概率,而非按 Token 进行自回归生成。”其文档开放了有限选项、是/否决策以及有序评分。这些自然可以用固定的数值输出表示。12

观察到的证据。API 依然会返回一个 output_tokens 字段,听起来像是生成记录。但事实并非如此。对于“是/否”类问题,该计数完全吻合:4 个共享 Token,加上每个答案 15 个 Token,再加每个问题标识符(identifier)的 Token 长度。TypeSafe 的文档指出,该标识符“不会被发送至底层的模型,也不用于推理”。一个随着模型从未见过的文本而变化的计数,是在推理之后根据序列化响应计算出来的。返回的数值也不会影响它:0.0 的答案与 0.01 的费用相同,尽管按通常规则每个数字都会算作一个 Token。224

该计数背后的 Tokenizer 与我们测试的 192 个公开 Tokenizer 中的任何一个都不匹配。它确实与 Jev 处理普通文本时的自身输入计数器相符,仅在连续较长的空格和标点符号上有所不同。output_tokens 只是一个计费数字。它完全没有告诉我们 Jev 是否生成了文本,即使真的生成了文本,它也无法对其进行测量。延迟同样不受该数字的影响:包含 200 个选项的问题(1,911 个输出 Token)返回的速度与包含 2 个选项的问题一样快,且服务器耗时仅随输入长度而增长。2425

一个包含 255 个选项的响应报告了 2,714 个输出 Token。4 如果将该数字除以请求耗时,并把结果称为模型的解码速度,那就错了。服务器可以在单次模型评估之后序列化数千个字符。这个计费字段并没有告诉我们发生了多少次神经解码步骤。

设想的 readout 接收最终的隐藏向量 hhh 并生成 Logit:

这里 KKK 是允许答案的数量。矩阵 WWW 将表示转化为答案得分;softmax 将这些得分转化为概率分布。对于“是/否”决策,一个标量和一个 sigmoid 就足够了。

这些类别无需是像 “payments” 这样的固定概念。它们可以是选项槽位(option slot):第一个选项、第二个选项、第三个选项。分支提供每个槽位的含义;应用程序代码将该概率映射回调用者的选项键(option key)。有序的 Score 可以类似地预测各等级上的概率,并返回它们的概率加权平均值。这支持了新的决策,而无需为每个客户的标签训练新的分类头。第 4 节中对比的指针式打分器(pointer-style scorer)是主要替代方案:它对每个选项自身的表示进行打分,而非编号槽位。

这并不能证明 Jev 拥有一个独立命名的分类器模块。语言模型的词表头(vocabulary head)同样也是一个后接 softmax 的矩阵。从该矩阵中选择 KKK 个预留标签行,可以实现与专用 KKK 类分类头相同的计算。这些行可能与输入嵌入绑定,也可能是独立训练的;我们无法在此区分这些安排。

重要的区别在于“读取概率”与“生成描述概率的文本”。生成的“91%”是一个 Token 序列。而分类器的 0.91 则属于其预测分布中的一项。两者都有可能校准不佳。没有任何一种仅因其格式就值得信任。

受限文本解码(Constrained text decoding)依然是构建类似接口的一种可能方式,但 TypeSafe 明确描述了一种不同的输出路径。其官方声明比延迟论据是更有力的证据。证据指向直接数值读取,即第 4 节对比的两种设计之一。预留标签 Token 依然有可能,尽管那里的伪造选项(fake-option)测试对这种可能性构成了不利证据。

04

共享内容只算一次,问题之间互不相见

暗号放在另一个问题里,这个问题问它得 0.00;放进共享内容,升到 0.90 到 0.92;跨接口挪动一句话,作用就变了。

2. 共享状态,隔离问题¶

下一个决策涉及在何处复用计算。

观察到的证据。在小型受控示例中,Token 计费完全具有可加性。一个极简的“是/否”问题消耗了 268 个输入 Token;两个消耗了 276 个。一个包含一个“是/否”问题、一个双选项 Choice 以及一个双层级 Score 的请求消耗了 318 个,与它们在共享开销之上的实测贡献总和相符。这契合公共前缀加上问题后缀的形式,尽管单凭计费本身尚无法确定完整的计算图。4

一个更有启发性的实验在这些区域之间移动证据。状态最初写道:

一个兄弟问题包含:

探针询问另一个问题提到了哪个代码,选项包括 ZEBRA-7741、两个干扰项以及 none(无)。当保密信息位于兄弟问题中时,其报告的概率为 0.00。移除该兄弟问题产生了相同的结局。改为将声明放在状态中则将其提升至 0.90–0.92。每种条件各重复了 5 次(探针记录中的 visibility 项)。5

这是一项有价值的干预测试:将声明跨 API 边界移动改变了其效果。它支持了问题之间的行为隔离以及对共享状态的访问。这并没有暴露具体的注意力掩码。独立的模型调用、树形掩码(tree mask)或其他限制信息流的机制都能产生相同的结局。探针的措辞即使在状态条件中也询问了“另一个问题”,因此它并不是对字面指令遵循的完美测试。

服务测量数据补充了拼图的另一角。在大约 100 个问题以内,服务器时间几乎没有变化。超过该数量后,时间稳步上升,且按 Token 计算,问题文本的开销大约是状态文本的两倍。这与对状态进行一次性计算并对问题任务进行 Batch 批处理相吻合。6

原文此处有图 · 去原文看图
图 2。服务器时间随着请求以两种方式增长:带有一个问题的更长状态(绿色),或带有一个短状态的 1 到 1,500 个问题(紫色)。两个面板使用相同的时间尺度。每种规模按打乱的顺序各请求了 8 次,一次一个请求。灰色叉号代表单个请求;实线连接每个规模的中位数,虚线连接最快的请求。两者都随工作量增长,但 1,500 个问题依然能在几百毫秒内返回。时间来自 API 的上游服务响应头,其中包含了共享服务开销;它们并非硬件基准测试。方法 ↗

这些是服务器报告的上游耗时,而非本地笔记本电脑上的计时。它们包含了上游服务所涉及的所有工作与等待时间,且该服务当时与其他用户共享。

Jev 实施了两项限制。每个分支(状态加上一个问题)上限约为 32,768 个 Token,整个请求上限约为 65,536 个。请求限制只对状态计数一次:包含 2.3 万 Token 状态和 5,000 个问题的请求均在该限制之内。如果每个问题都各自处理一份状态副本,该请求将超过 1 亿个 Token。这一对设置契合最高 2¹⁶ 个 Token 的单条打包序列,其中容纳一次状态以及其后的每个问题,且每个分支被限制在 2¹⁵ 的上下文窗口内。22

带有独立因果后缀的前缀 KV cache(prefix KV cache)是自然的实现方式。Hydragen 描述了共享前缀序列的高效注意力机制;DeFT 开发了用于树状结构推理的注意力机制。这些证明了该服务模式是可行的。它们是现有技术,并非 TypeSafe 使用了其中某个库的证据。78

这种设计还澄清了一个表面上的矛盾:隔离的问题依然可以在同一个加速器上共同评估。“并行”描述的是它们的调度方式以及缺乏答案依赖性,并不一定意味着每个问题分配一块 GPU。

05

选项被当成一张列表一起读

加一个无关的「坏天气」选项,原来两项的对数几率从 0.38 掉到 0.11,十组随机实验每组都降,「每项独立打分」被证伪。

3. 因果主干网络¶

实验无法将因果解码器(causal decoder)与双向编码器(bidirectional encoder)区分开来:在这两者中,最终决策都可以读取整个输入。不过我依然假设它是因果解码器,并且理由充分。Jev 广泛的知识面(在 MMLU-Pro 上达到 84.6%)需要前沿级别的预训练,而该规模下的每个模型都是因果解码器,且 TypeSafe 将 RLCD 描述为对预训练语言模型进行后训练(post-training)。一个双向的 Jev 将意味着要么基座模型弱得多,要么需要支付额外成本转换解码器,同时放弃因果服务所提供的共享前缀缓存。那将令人吃惊,但从外部看无法完全排除这种可能性。1215

具体使用的是哪款预训练模型尚不得而知,Tokenizer 也未能揭示。在跨越 415 次探针测试中,Jev 的 Token 计数与我们测试的 192 个公开 Tokenizer 中的任何一个都不匹配。它会对每个数字单独拆分,并在合并前查找整块内容:8 个 a 算作 1 个 Token,但 16 个则算作 4 个。它的词表紧密贴合 OpenAI 的 o200k,因为被 Jev 算作单个 Token 的每个字符串同样也是单个 o200k Token,然而数字拆分和若干合并规则排除了 o200k 本身。最接近的公开匹配者 Qwen 在 415 个探针中有 348 个相符。这排除了未经修改的公开 Tokenizer,但并未排除公开基座模型:更换词表、继续预训练或蒸馏都可以解释这一点,计费 API 与模型采用不同的 Token 计数方式亦同理。18

实验确实展现了决策可以读取的内容。我在一个问题的选项中放入了一张参考卡片(reference card),并要求 Jev 选择满足卡片条件的那一个选项。以下是其中一个确切的选项集:

指令是:“阅读参考卡片,并选择满足条件的唯一选项。”将 reference 值改为 indigo 会切换正确答案,同时保持可选选项不变。

我测试了两个数值、全部 6 种选项排列组合,以及使用 route = east/west 的第二个模板,各重复 2 次。配对对照组则将 reference 放在共享状态中。这产生了 48 次选项-参考卡片试验和 48 次状态-参考卡片对照组。20

Jev 可以利用置于候选描述之后的信号信息。当参考卡片放在最后时,它在每次试验中都选出了正确选项,正确答案的平均概率约为 0.88。

原文此处有图 · 去原文看图
图 3。Jev 1.13.0 给出正确选项的概率,按参考卡片放置的位置划分。实心单元格标记每种选项顺序中的卡片 (α);最后一行将其移至共享状态,并汇总所有 6 种顺序。灰色叉号代表单个请求(2 项任务 × 2 个卡片数值 × 每种顺序 2 次重复);彩色刻度标记均值。当卡片在最后时,每个请求均正确。当卡片在最前或中间时,尽管最终决策始终可以读取卡片,概率却在 0.5 附近广泛分散。这是顺序敏感性,而非还原出的注意力掩码。记录于 2026 年 9 月 17 日。请求载荷与答案 ↗

数值切换对照组与位置同样重要。当卡片在最后时,仅将 amber 改为 indigo 就会改变先前哪个选项胜出,尽管那些较早的描述和状态保持完全相同。一个仅凭选项自身文本和状态独立对每个选项打分、然后仅仅对得分进行归一化的模型,没有任何路径能够让这一事实改变早期选项之间的相对排名。结果支持了选项能够影响联合决策的路径。20

这契合在整个列表之后计算的任何 readout,包括第 4 节对比的两种设计,以及独立的选项混合阶段。其余的错误展示了在这两个模板上对位置敏感的处理过程;它们并没有指出唯一的单一原因。

扩散机制对于这种计算是不必要的,且这些实验中没有任何内容需要迭代去噪。可以站得住脚的架构推断范围更窄:答案计算能够访问完整的选项列表。下一个实验测试了它是否真的利用了该联合上下文。

4. 让选项在做出选择前发生交互¶

在问题内部,证据指向了一个不同的信息边界:备选项被作为有序列表一起读取,随后跟随一个决策位置。

为什么要允许这种交互?像“以上皆非”这样的选项取决于其他选项。即使是普通备选项也能对问题起到澄清作用。“payments(支付)”、“account access(账户访问)”和“other(其他)”定义的决策不同于“bank(银行)”、“payment provider(支付提供商)”和“customer(客户)”。列表式(listwise)表示让模型能够在产生概率分布之前对这种区别进行解释。

最强有力的证据是一个带有无关额外选项的实验。

首先设置付款失败的 4 个可能原因:bank、provider、customer 和 unknown。然后追加 weather(天气):Bad weather caused it(坏天气导致的)。如果每个原始选项都接收到一个独立的、未经改变的 Logit,且服务器应用相同的 softmax 温度,那么添加第 5 个选项会改变归一化过程,但无法改变两个现有选项之间的赔率(odds):

公共分母相互抵消。这给了我们一个具体的、可证伪的预测。

原始研究发现数值从大约 +0.49 移到了 +0.08。10 为了验证这一现象在日常请求波动中是否依然存在,我在 10 个随机区块中重复了该实验。每个区块包含 4 选项基线、一个相同的 4 选项对照组、追加了 weather 的 5 选项版本、一个相同的 5 选项对照组,以及追加描述从“Bad weather caused it”改为“Wild birds caused it(野生鸟类导致的)”的 5 选项版本。每个请求包含一个问题。21

选项扩充的结果成功复现。汇总每个区块内每种条件的两个相同请求,平均对数几率(log-odds)从 +0.38 降至 +0.11。每个区块都显示出下降;平均变化量为 −0.28,描述性 95% 配对 t 区间约为 −0.36 至 −0.19。这里的汇总利用对照请求来降低日常请求噪声,而非将重复输出视为独立实验。21

原文此处有图 · 去原文看图
图 4。添加无关选项会改变两个现有选项之间的赔率吗?每行代表一个随机区块。灰色叉号为单个请求(每种列表规模 2 个相同载荷);灰点汇总了 4 选项请求,紫点汇总了追加“bad weather caused it”的 5 选项请求。如果每个选项保持固定得分且 softmax 温度不变,共享分母将相互抵消,两个点将会重合。区间是跨 10 个区块(9 个自由度)的配对 t 区间,来自对单一场景的探索性研究;概率舍入、请求噪声以及依赖于列表的温度依然是可能的贡献因素。它表明选项发生了交互,而非指出交互发生在模型的哪个位置。请求与答案 ↗ · 摘要 ↗

这是反对“固定独立 Logit 后接未改变 softmax”的证据。它并没有唯一确定具体的机制。在保持 5 个选项的同时更改追加的描述,会产生更小且不具决定性的偏移:其配对区间包含了零。除了依赖内容的混合之外,依赖于选项集的温度也是有可能的。

能够看到完整列表的 readout 自然地解释了这一点:添加选项改变了它所读取的上下文。FIRST 列表式重排序方法也是以相同方式工作的,它从首 Token 的 Logit 中提取排名,而非逐 Token 生成。11

两种 readout 设计契合该证据。最终位置分类头(final-position head)从决策 Token 的表示对每个选项槽位进行打分;指针式打分器(pointer-style scorer)则将该表示与每个选项自身的最终隐藏状态进行比较。两者都能让选项相互影响。API 最多接受 255 个选项(2⁸ − 1),这适合固定的 256 槽位分类头,但该限制是由请求验证强制实施的,而非模型本身。在 200 个选项的情况下,一个复制的答案在每个位置上的得分均为 1.00,且错误未溢出至相邻选项,这适合指针式。两个结果均不具决定性。26

注入的伪造选项从未取代真实选项,因此选项边界是以某种文本无法伪造的方式标记的,并且其条件在列表中别处被重复的选项会将其概率输给竞争对手。23

这种权衡在普通任务中也清晰可见:颠倒选项顺序列将技术支持分类的概率从大约 0.84–0.89 改变到了 0.93–0.96。这来自 option_order 探针。9 对于已部署的决策策略而言,这至关重要。接近 0.9 的阈值可能会改变最终操作,即使标签和证据完全相同。排列组合测试(Permutation tests)理应包含在该设计任何实现的评估之中。

06

便宜的概率,也可能是坏概率

校准问的是标 0.8 的那批里是否真有约 80% 答对;MMLU 1200 题校准误差 0.0313,但 990 题挤在 0.9 到 1.0 区间。

5. 训练分布,然后计算置信度¶

第五个组件是训练目标。直接数值输出节省了解码工作,但廉价的概率依然可能是糟糕的概率。

设想一组被赋予 0.8 紧急概率的案例。校准(Calibration)所关心的,是其中是否真的有大约 80% 属于紧急情况。这是跨案例预测的一种属性。我们无法根据某个具体案例的结局好坏,来判断单次预测是否经过了校准。

TypeSafe 将其训练方法称为校准决策强化学习(Reinforcement Learning for Calibrated Decisions,简称 RLCD)。发布公告称其优化目标是“在 System One 任务上给出具有认知诚实概率的答案”;该公司的入门指南将 RLCD 描述为从预训练语言模型演进的后训练路径。112 确切的配方未曾公布。我设想的训练配方利用基于结果的目标,将 Transformer 和 readout 调整至适应强类型决策任务。这给主干网络提供了一个机会去构建对可靠决策有用的表示,而非仅仅是流畅的补全。

一个自然的目标是对数损失(log loss),即针对观察结果 yyy 的 −log⁡p(y)-\log p(y)−logp(y)。另一个是 Brier 损失,即预测分布与观察到的单热(one-hot)结果之间的平方距离。两者都是严格的对规评分规则(proper scoring rules):在期望层面,报告真实的条件分布能够使损失最小化。Gneiting 和 Raftery 给出了形式化定义与理论。13 这解释了此类训练试图实现的目标。它并不能确定 TypeSafe 使用了哪种损失、其流程在狭义算法层面是否属于强化学习,或者是否更新了每一个主干网络权重。

对规性(Properness)同样不是部署时的铁板保证。有限的数据、模型的局限性、优化误差以及分布偏移(distribution shift)都可能导致校准不够完美。Guo 等人展示了现代神经网络的校准问题以及事后调整(post-hoc adjustments)的实用性。训练与事后校准是兼容的机制;API 无法区分它们各自的贡献。14

观察到的证据。基准测试记录使我们能够比较预测概率与观察到的准确率,不论是在总体上还是在概率分箱(probability bins)内部。图表展示了这些检验。仅靠平均值的一致性是比分箱内一致性更弱的证据:一个组的过度自信可能会抵消另一个组的自信不足。在包含 1,200 个条目的 MMLU 样本上,十箱预期校准误差(expected calibration error,ECE)为 0.0313(分箱定义与条目级预测)。大多数预测集中在接近确定的区间:有 990 个落入 0.9–1.0 分箱中。15

原文此处有图 · 去原文看图
图 5。给定的概率是否与 Jev 猜对的频率吻合?两个面板均绘制了观察到的准确率与 Jev 给出的所选答案概率的对比,概率系根据记录的输出重新计算,而非来自 API 独立的 confidence 字段;虚线对角线上的点代表完美校准。紫色:分组为概率分箱的 1,200 个 MMLU 条目,标有每个分箱的条目数(概率在分箱前四舍五入保留两位小数)。铁锈色:新生成的数学题,每个族系(family)一个点。垂直线为 95% 威尔逊区间(Wilson intervals),仅覆盖采样噪声,不覆盖基准测试选择或训练暴露。预期校准误差(ECE)按各分箱的条目占比对其与对角线的差距加权。族系平均值一致是比分箱一致更弱的证据,因为在一个族系内部过度自信与自信不足可能会相互抵消;模幂运算(modular exponentiation)是明显的例外,在平均概率为 35% 的情况下正确率达到 56%。总体证据 ↗ · 可靠性数据 ↗

这项小型的新数学题研究增加了有价值的变体。在生成的三位数乘法问题上,准确率为 86.7%,平均最高概率为 0.83。在两步应用题上,准确率降至 32%,平均最高概率降至 0.30。模型在更难的任务上置信度更低(fresh_math_results,包含 30 个乘法条目和 25 个应用题条目)。15 这令人鼓舞,尽管小规模类别级的平均值尚无法确立对每种未见过的问题的校准表现。

这些结果还解释了为什么公开基准测试分数是衡量模型所知知识的不完美指标。MMLU-Pro 准确率为 84.6%;而新生成的应用题则难得多。15 任务结构、干扰项、难度以及训练暴露方面的差异都可能是原因。这一差距并不能证明基准测试受到了污染。全新的措辞同样也不会让底层的数学技能或事实性知识变成未见过的状态。

关于 API 中名为 confidence 的字段,存在一项独立且异常清晰的发现。官方适配器根据归一化分布计算 Choice 置信度(当 K>1K > 1K>1 时),公式为:

对于最高概率为 0.8 的 3 个选项,该公式给出 0.7。适配器单独处理单选项的情况,返回 1。它测量的是领先答案高出均匀分布的程度。它并不是对答案正确性的另一种学习到的估计。Score 类型则使用不同的公式,反映与众数层级(modal level)的距离。16

在设想的系统中,训练产生了预测分布;而普通的算术运算输出了这个总结性字段。将这两个对象区分开来可以避免常见的概念性错误:一个高度集中的分布依然可能会“自信地犯错”。

6. 稀疏容量¶

我预计 Jev 使用了稀疏混合专家(sparse mixture-of-experts,MoE)Transformer。在特定层,路由模块(router)将每个 Token 发送通过前馈网络的一小部分子集,因此模型能够存储大量参数,同时对每个 Token 仅激活其中一部分参数:这即是 Shazeer 等人通过稀疏门控 MoE 层展示的条件计算(conditional-computation)思想。17

无法从外部直接观察到稀疏专家,但它们是极有可能的选择。一个纯 prefill 的模型受限于计算量,而这恰恰是稀疏路由所能节省的。MoE 通常的服务端成本大部分消失了:不存在逐 Token 解码(在解码阶段内存带宽占主导,且无论如何大多数专家最终都会被激活),也不存在长期存留的 KV cache 与专家权重争夺显存的情况。测量数据也指向了同一方向。Jev 在大约 160 毫秒内处理了约 3 万个 Token;在 8×H100 节点上的稠密 70B 模型需要约 1 秒,而带有约 10B 激活参数的 MoE 则正好契合。此外,近期最强的大多数基座模型(DeepSeek-V3、Qwen3、GLM-4.5、Kimi K2、gpt-oss)均为 MoE。专用硬件可能会让稠密模型追上这一速度,且基准测试分数可能会高估模型掌握的知识量,因此这依然是一项推断,而非实测结果。615

重构过程中的其他部分概不依赖于此。即使替换为稠密 Transformer,接口、共享状态、隔离的分支以及 readout 依然会与所述完全一致。

07

每条推断都标明证据强度

问题分支被当成独立任务成批处理;结尾把直接输出概率(公开)、问题隔离(可观察)、共享缓存和稀疏专家(推断)分成三层。

7. 将分支作为 Batch 调度,而非对话¶

最后一个组件是一个服务引擎,它将问题分支视为独立的工作项。它们的后缀可以在读取共享状态表示的同时被打包进 Batch。然后,应用程序代码将数值输出与问题标识符关联起来,并序列化响应。

测量数据揭示了重复相同答案之间的微小差异,包括单个请求内部重复问题之间的差异。这意味着不应假定 API 级别的确定性(noise、dup 和 determinism)。19 这并不意味着模型在生成或采样文本:数值算子(numerical kernels)、动态 Batch(dynamic batching)、路由或故意引入的随机性都会影响直接读取。

响应键(key)的顺序也在少数反复出现的模式中有所变化。19 具有不同哈希顺序的多个 Worker 是一个合理的解释。然而,这一侧信道并没有确定 Worker 的数量、确立 KV cache 的位置,或者告诉我们使用了哪种数值精度。这些是现有观察手段无法解决的实现细节。

对设想架构而言,关键在于答案之间不存在依赖链。模型无需在开始紧急程度估计之前先完成队列分类的写入。两者都依赖于状态;任何一方都不消耗另一方生成的答案。

依赖性限制依然存在。如果后面的问题真正需要前面的答案,应用程序必须引入另一个决策阶段,或者在一个问题中表达联合决策。共享上下文并不能消除工作流的逻辑结构。

哪些因素会改变我的看法?¶

本重构做出了不同层面的推断。直接概率输出是公开描述的。问题隔离与选项顺序效应是可观察的行为。而 KV 共享、因果注意力、最终位置或指针式 readout,以及稀疏专家,则是层层递进的具体解释。

参考卡片实验解决了一个问题:决策可以使用放在最后的选项。伪造选项测试表明,输入格式中的小技巧无法伪造选项边界。更广泛的关联任务可以进一步约束表示,尽管单凭行为上的成功依然无法唯一确定注意力掩码。

在选项处理方面,随机化后续实验复现了选项集(choice-set)效应,但固定大小描述的干预依然缺乏定论。更多的模板和独立的请求区块有助于区分共享的温度变化与依赖内容的交互。包含 200 个选项的中等难度任务可以将槽位分类头(slot head)与指针打分器区分开来。对于校准而言,保留的工作流数据和在偏移条件下的重复评估,比另一个总体基准测试分数更具参考价值。证实稀疏专家可能需要该 API 之外的信息披露或证据。

我对 Jev 最好的重构依然是开篇图示中的模型:一个带有共享状态前缀、隔离问题后缀、列表式选项处理、强类型数值 readout,以及面向预测分布进行训练的因果 Transformer。稀疏专家是极有可能的主干网络,尽管设计中的其他部分都不依赖于它们。

它的实用性来自于计算图与具体任务的契合。决策服务需要读取证据、比较允许的结果,并揭示不确定性。Transformer 可以做到这一点,而无需先将每个决策都转化为一句话。

方法¶

本文基于 2026 年 9 月 17 日对 jev-1.13.0 的调查,使用了一个早期访问账号和一个观察到的服务区域。源研究包含 1,029 条带仪表的探针记录(包括 190 个生成的数学题条目)、6,800 条基准测试记录以及独立的事实核查。后续研究补充了 146 个关联与选项交互请求(trials、summary)、311 个 Token 计费请求、445 个 Tokenizer 指纹请求、192 个延迟请求、148 个选项数量延迟请求、181 个选项位置请求、105 个伪造选项请求以及 35 个上下文限制请求。每一项均在参考文献中给出链接,并附有确切的请求与脱敏后的响应。重复的基准测试配置共享底层条目;这些计数并非独立问题的计数。

可下载的证据包记录了本文所使用的观察数据。开篇章节中的 API 示例为示意图。引用的 visibility 和 reference-card 提示词来自探针脚本和保存的后续请求。数值观察数据仅针对该模型版本和测试活动。

延迟数字来自 x-envoy-upstream-service-time 响应头。它们是上游服务的耗时,具有未知的排队和执行边界,而非独立的模型计时。延迟图表的扫描测试按打乱的顺序一次运行一个请求;所有研究均未控制服务器负载。本地墙上时钟(wall-clock)测量值不作为架构证据。

概率通常以两位小数精度返回。单次请求中的重复问题共享条件,并可能具有相关误差。MMLU 校准图表使用了 10 个等宽分箱:[0, 0.1)、[0.1, 0.2) 等,其中 1.0 包含在最后一个分箱中。预期校准误差(ECE)是每个分箱内准确率与平均最高概率之差的样本加权绝对值。估计值取决于样本选择、分箱和响应舍入。证据支持关于测试分布的断言,而非保证在未来客户工作流中的校准表现。

来源与相关工作¶

实验参考文献指出了原始脚本标签,以便在证据包中定位每一项观察结果。论文引用确立了设想机制及其前身;它们并不能证明 Jev 使用了它们。

来源

- [1] TypeSafe (2026). Introducing System One Models and Jev. 并行输出主张及 RLCD 声明主要来源。

- [2] TypeSafe. 完整 API 文档,访问于 2026 年 9 月 17 日。强类型问题、响应分布及 API 契约。

- [3] Vaswani et al. (2017). Attention Is All You Need. 解码器掩码、注意力及线性/softmax 输出层。

- [4] API 实验:type_preamble 与 outputs。Token 计费可加性、标识符变更及 255 选项响应。

- [5] API 实验:visibility。对保密信息位于兄弟问题中、缺失于兄弟问题中以及位于状态中分别进行 5 次重复。

- [6] 延迟扫描(192 个顺序请求)。状态长度与问题数量,打乱重复 8 次;服务器报告的上游耗时。

- [7] Juravsky et al. (2024). Hydragen: High-Throughput LLM Inference with Shared Prefixes.

- [8] Yao et al. (2024). DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference.

- [9] API 实验:option_order。普通工单顺序敏感性。

- [10] API 实验:iia。每种条件 3 个请求,每个请求包含 40 个重复问题;原始、追加和前置选项集。

- [11] Reddy et al. (2024). FIRST: Faster Improved Listwise Reranking with Single Token Decoding.

- [12] TypeSafe. 机器学习入门指南。RLCD 后训练路径及校准契约的主要描述。

- [13] Gneiting and Raftery (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. Journal of the American Statistical Association 102(477):359–378.

- [14] Guo et al. (2017). On Calibration of Modern Neural Networks.

- [15] 基准测试与生成数学题记录:准确率及平均预测概率。; MMLU 可靠性分析:分箱定义、ECE、威尔逊区间及 1,200 个条目级预测。

- [16] TypeSafe. 官方 Python 适配器,confidence_metrics.py,版本 fb52b103。Choice 与 Score 置信度公式(查阅于 2026 年 9 月 17 日)。

- [17] Shazeer et al. (2017). Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

- [18] Tokenizer 指纹实验(445 个请求)。游程长度(Run-length)、词表及预 Token 化探针,对比 192 个公开 Tokenizer。

- [19] API 实验:noise、dup 与 determinism。重复概率及响应键顺序。

- [20] 后续关联实验(96 个请求)。2 个模板、2 个参考值、6 种排列组合、2 个位置及 2 次重复;确切请求与脱敏响应。

- [21] 后续选项交互实验(50 个请求)。10 个随机区块包含 base4、null4、append5、replace5 和 null5;配对变化量与标准误差。

- [22] 上下文限制实验(35 个顺序请求)。单分支与全请求 Token 限制,包含接受与拒绝的边界用例。

- [23] 伪造选项注入实验(105 个请求)。7 种分隔符格式、饱和与模糊的基础任务以及完整概率向量。

- [24] Token 计费实验(311 个请求)。问题 ID 长度、Batch 规模、状态难度、单词 ID 以及匹配的状态与 ID 字符串。

- [25] 选项数量延迟实验(148 个请求)。带 2–200 个选项的 1 个或 20 个问题,长短标签,以及输入/输出解耦对照组。

- [26] 选项位置实验(181 个请求)。选项数量上限,正确答案在 10、50、200 和 255 选项列表中移动。

判断收口延伸

Indigo 的结论

Jev 是「验证省不掉」的商业化答案:不压掉验证,而是把对着真实结果校准过的决策信号做成产品来卖。但它只在有真实结果可以回训的快决策领域成立,没有终审的那一半仍然空着。

需要记住的几件事

  1. 它解决两件事:决策信号不可靠,加上生成这个信号的算力白白浪费。
  2. 读出的概率不等于描述概率的文字:生成的「91%」是一串字,分类器的 0.91 是预测分布里的一项;两者都可能没校准。
  3. 架构逆推(作者标为推测):因果 transformer,很可能是稀疏混合专家,共享内容的缓存加上互相隔离的问题。
  4. RLCD 是为校准决策做的强化学习,目标是给出诚实的概率;MMLU 1200 题的校准误差 0.031。

可回查的判断

判断谁说的何时见分晓证据多硬
Jev 不逐字生成,只读入一次,直接读出概率Archer(据 TypeSafe 说法和延迟实验)现状TypeSafe 一手说法,加一手实验佐证
底座是因果 transformer,很可能是稀疏混合专家Archer现状黑箱推断,作者标为推测
同一个问题里的选项会互相影响(不是各自独立打分)Archer(加无关选项实验)现状一手实验,十组复现,平均 −0.28
Jev 在测试分布上校准尚可(MMLU 校准误差 0.031),但不保证适用于未来的客户流程Archer现状一手实验,作者明确不外推
底座具体是哪个预训练模型未知分词结果与 192 个公开分词器都不匹配,无法定位

放回主线

证实

验证不可压缩:瓶颈·护城河·断点 Jev 是这条判断的商业化答案:不压掉验证,把对着结果校准的决策信号本身做成产品卖。

补充

harness:吃掉的是招式,留下的是接口 Jev 剥掉「生成文字」这个招式,只留带类型的决策接口;护城河在评估,不在底座模型。

证实

Ivan Zhao(Notion CEO)钢铁蒸汽与无限心智 Ivan 描述没有终审的那一半,Jev 拿下能训练的那一半:同一条判断的两侧。

补充

Google AI-in-Science 首份大样本实证 省下时间的人大多把时间花在检查 AI 的产出上;Jev 是供给侧针对这件事做出的产品。

补充

Noam Brown(OpenAI)内部人复盘 OAI-HF 一个讲监控模型自报置信有多脆弱,一个讲绕开生成、直接训练可信的概率。

什么会让我改口

在没有真实结果可训练的领域,这种校准概率也能做得可信;那就不是把「验证省不掉」切成两半,而是把它压掉了。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Essay

Jev’s Architecture Unmasked

Archer Hume · archerhume.com · 2026-09-17

Replaces the industry's bad habit of letting a model announce “90% sure” with probabilities trained on real outcomes; the reverse-engineering itself is a lesson in calibration.

Indigo's conclusion

Jev is the commercial answer to “verification can't be skipped”: rather than skip verification, it sells decision signals calibrated against real outcomes. But it holds only for fast decisions with real outcomes to train on; the half without a judge is still exposed.

How to read this An independent reverse-engineering essay; the author has no evident stake in TypeSafe. Weigh the evidence in layers: his own probing experiments are first-hand and checkable, with the evidence published; his conclusions about TypeSafe's real architecture are extrapolated from a black box, and he marks them as speculation again and again; TypeSafe's product claims are the vendor's own.

What to remember

  1. It solves two things at once: unreliable decision signals, and compute wasted generating them.
  2. A readout probability isn't text describing a probability: a generated “91%” is a string, a classifier's 0.91 is an entry in its predicted distribution; either can be miscalibrated.
  3. The reconstructed architecture (the author calls it speculation): a causal transformer, likely a sparse mixture of experts, with a shared-state cache and isolated questions.
  4. RLCD is reinforcement learning for calibrated decisions, aimed at honest probabilities; calibration error on 1,200 MMLU questions is 0.031.

Breakdown · 7 steps

  1. 01

    A generated “90% sure” is not a probability

    Fraud screening, moderation, routing and risk scoring are built on this: paying for token-by-token generation, then treating an unverified confidence claim as a probability software can act on. Read this part →

  2. 02

    Probabilities let downstream rules do the math

    If a wrong escalation costs 1 and a missed urgent case costs 9, the rule is to escalate when the urgency probability exceeds 0.1. The math only works if the probability is reliable. Read this part →

  3. 03

    Inference ends with a readout, not decoding

    output_tokens sounds like a record of generation but is a billing number: 200 options return as fast as 2. Don't divide it by time to get a decoding speed. Read this part →

  4. 04

    Shared state computed once; questions can't see each other

    Put a code word in a sibling question and this question reports 0.00 for it; put it in the shared state and it rises to 0.90–0.92. Moving one sentence across the API boundary changes what it does. Read this part →

  5. 05

    Options are read together as a list

    Adding an irrelevant “bad weather” option cut the log-odds between two existing options from 0.38 to 0.11, in every one of ten random blocks: fixed independent scores are ruled out. Read this part →

  6. 06

    Cheap probabilities can still be bad ones

    Calibration asks whether cases labeled 0.8 are right about 80% of the time. On 1,200 MMLU questions the calibration error is 0.0313, but 990 of them sit in the 0.9–1.0 bin. Read this part →

  7. 07

    Every inference labeled with its strength of evidence

    Question branches are batched as independent jobs. The close sorts claims into three layers: direct probability output (published), question isolation (observable), shared caching and sparse experts (inferred). Read this part →

What would change my mind

calibrated probabilities that are trustworthy even where no real outcomes exist to train on. Then verification wouldn't be split in two; it would be skipped.

How to read this

An independent reverse-engineering essay; the author has no evident stake in TypeSafe. Weigh the evidence in layers: his own probing experiments are first-hand and checkable, with the evidence published; his conclusions about TypeSafe's real architecture are extrapolated from a black box, and he marks them as speculation again and again; TypeSafe's product claims are the vendor's own.

Breakdown · 7 steps
  1. A generated “90% sure” is not a probability
  2. Probabilities let downstream rules do the math
  3. Inference ends with a readout, not decoding
  4. Shared state computed once; questions can't see each other
  5. Options are read together as a list
  6. Cheap probabilities can still be bad ones
  7. Every inference labeled with its strength of evidence
01

A generated “90% sure” is not a probability

Fraud screening, moderation, routing and risk scoring are built on this: paying for token-by-token generation, then treating an unverified confidence claim as a probability software can act on.

X is full of hot takes about Jev’s launch, and most miss the point entirely: “12 million views for a JSON classifier? Yeah, we’re in a bubble.” An ordinary LLM generates “90% confident” as text; its probability of producing those words does not establish a 90% probability of being right. Yet we build fraud screening, moderation, routing and risk assessment around precisely this pattern: paying for token-by-token generation, then treating an unvalidated confidence claim as a probability our software can act on.

Jev’s proposition is to retain the knowledge of a pretrained LLM while replacing generated confidence claims with decision probabilities read directly from its internal representations. Those probabilities are trained against outcomes. Give it shared state, questions, and allowed answers; it returns the distributions in parallel, without generating text.1 For this whole class of applications, that addresses both problems: the reliability of the decision signal and the unnecessary computation spent producing it.

Theres just one problem, its not open weight, and TypeSafe refuses to share their research… So I will (try my best).

The evidence points toward a causal transformer (likely using sparse MoE) repurposed for decisions: shared-state encoding, isolated question branches, and direct probability readouts instead of text generation. After probing the TypeSafe API (looking for signatures in latency scaling under different context lengths, question reordering, etc.), scouring any public documentation and research with Astra, and looking for prior art, I think I have a fairly accurate model of how it works and its architecture.

The sparse backbone is the least certain part, but its particularly beneficial for this domain with less downsides than AR LLMs, so would be weird if it wasn’t . Shared computation and direct probability outputs are much better supported obviously. This is clearly all quite speculative, so I’ll try and be as clear as possible around what evidence was published by TypeSafe, what was observed in experiments, and whats inferred from them. Black box APIs make it shockingly easy to throw a blanket over the ghost and get a rough shape of what the architecture looks like.

There is a figure here in the original · See it in the original
Figure 1. Proposed decision-model computation. The message is encoded once into the green grid, representing state information retained at each transformer layer. Each coloured question grid builds in parallel, combining its own text and allowed answers with attention to the shared state. Questions cannot attend to one another. Outcome-trained readouts turn their final representations directly into answer probabilities, without generating text. Travelling cells illustrate information flow, not literal copying; colours, attention paths and probabilities are schematic, not measurements of Jev.
02

Probabilities let downstream rules do the math

If a wrong escalation costs 1 and a missed urgent case costs 9, the rule is to escalate when the urgency probability exceeds 0.1. The math only works if the probability is reliable.

Why this design is useful¶

Consider a schematic support-routing request. This illustrates the API’s structure; the example probabilities below are invented.

A useful answer might assign payments 0.91 probability while assigning urgent escalation only 0.42. Those are different uncertainties. Software can route the ticket automatically while leaving escalation to a separate policy.

A causal transformer already knows how to construct a representation of text from left to right. During ordinary language-model inference, it processes the prompt, predicts one token, feeds that token back in, and repeats. This builds on the decoder attention and output projection introduced in the original Transformer.3 But the prompt-processing stage has already produced a rich representation. If the task is to choose among three queues, we can attach a small function that maps that representation directly to three numbers.

A detail here resolves much of the confusion about parallel answers: causal attention describes which positions can use which information, not the order in which input tokens must be executed. During prompt processing, or prefill, every input token is already known. The model can process their positions together within a layer while the attention mask blocks access to later positions; the layers still run sequentially. Autoregressive decoding adds another dependency: the next token does not exist until the previous prediction has been chosen. Our proposed model ends after prefill and the readout, so it avoids that token-by-token dependency.

This changes the computational shape of the task. The output no longer needs a sequence of spelling decisions for "payments": 0.91. JSON formatting happens in ordinary application code. The neural network supplies the probabilities.

Now suppose the state is a lengthy incident report, and there are fifty questions. Most of the input is shared. A transformer stores intermediate information about processed tokens in its key–value cache, usually shortened to KV cache. In the proposed design, every question reads the same state cache. Each branch adds only its own instructions and answer options.

For a state of SSS tokens and QQQ questions, separate requests would process the state roughly QQQ times. Sharing reduces the repeated state-token processing from QSQSQS to SSS. The questions still have to attend to the state; that work does not disappear. But the model need not repeatedly reconstruct the state’s representations.

Isolation also gives the interface a useful meaning. Asking whether the customer is angry should not change which queue receives the ticket. Both questions can inspect the same evidence without reading each other’s instructions. The branches have no computational dependency on one another, even when their answers are statistically related.

Finally, probabilities make downstream policy explicit. If an unnecessary escalation costs one unit and a missed urgent case costs nine, a simplified decision rule escalates when p(urgent)>0.1p(\text{urgent}) > 0.1p(urgent)>0.1. That calculation is meaningful only to the extent the probabilities are reliable for this workflow. Training and evaluating the probability distribution therefore becomes part of the product, rather than a cosmetic confidence field.

None of this requires diffusion. Parallel classification has existed for decades. The interesting combination is a broadly capable transformer, shared contextual computation, a typed output interface, and training that rewards useful uncertainty.

03

Inference ends with a readout, not decoding

output_tokens sounds like a record of generation but is a billing number: 200 options return as fast as 2. Don't divide it by time to get a decoding speed.

1. End inference with a readout¶

The first component is the simplest: a prediction head instead of a decode loop.

Published evidence. TypeSafe’s launch announcement says: “Jev outputs all probabilities in parallel instead of autoregressively generating by token.” Its documentation exposes finite choices, yes/no decisions, and ordered scores. These are naturally represented by fixed numerical outputs.12

Observed evidence. The API still reports an output_tokens field, which sounds like a record of generation. It isn’t one. For yes/no questions, the count fits exactly: 4 shared tokens, plus 15 per answer, plus the token length of each question’s identifier. TypeSafe’s documentation says that identifier “is not sent to the underlying model and is not used in inference.” A count that changes with text the model never sees is calculated after inference, from the serialised response. The returned values don’t affect it either: an answer of 0.0 costs the same as 0.01, although each digit otherwise counts as a token.224

The tokenizer behind the count doesn’t match any of the 192 public tokenizers we tested. It does match Jev’s own input counter for ordinary text, differing only on long runs of whitespace and punctuation. output_tokens is a billing figure. It tells us nothing about whether Jev generates text, and wouldn’t measure that text even if it did. Latency doesn’t track that figure either: a question with 200 options (1,911 output tokens) returned as quickly as one with two, and server time grew only with input length.2425

A 255-option response reported 2,714 output tokens.4 It would be a mistake to divide that number by request duration and call the result the model’s decoding speed. A server can serialize thousands of characters after a single model evaluation. The accounting field does not tell us how many neural decoding steps occurred.

The proposed readout takes a final hidden vector hhh and produces logits:

Here KKK is the number of allowed answers. The matrix WWW converts a representation into answer scores; softmax turns those scores into a distribution. For a yes/no decision, one scalar and a sigmoid would suffice.

The classes need not be fixed concepts such as “payments”. They can be option slots: first option, second option, third option. The branch supplies each slot’s meaning; application code maps its probability back to the caller’s option key. An ordered Score can similarly predict probabilities over levels and return their probability-weighted average. This supports new decisions without training a new head for each customer’s labels. A pointer-style scorer, compared in section 4, is the main alternative: it scores each option’s own representation instead of a numbered slot.

This does not establish that Jev has a separately named classifier module. A language model’s vocabulary head is also a matrix followed by softmax. Selecting KKK reserved label rows from that matrix can implement the same computation as a dedicated KKK-class head. The rows might be tied to input embeddings or trained independently; we cannot distinguish those arrangements here.

The important distinction is between reading out probabilities and generating text that describes probabilities. A generated “91%” is a token sequence. A classifier’s 0.91 is an entry in its predictive distribution. Either can be miscalibrated. Neither becomes trustworthy solely because of its format.

Constrained text decoding remains a possible way to build a similar interface, but TypeSafe explicitly describes a different output path. Its statement is stronger evidence than a latency argument. The evidence points to a direct numerical readout, one of the two designs compared in section 4. Reserved label tokens remain possible, although the fake-option test there weighs against them.

04

Shared state computed once; questions can't see each other

Put a code word in a sibling question and this question reports 0.00 for it; put it in the shared state and it rises to 0.90–0.92. Moving one sentence across the API boundary changes what it does.

2. Share the state, isolate the questions¶

The next decision concerns where computation is reused.

Observed evidence. Token accounting is exactly additive in the small controlled examples. One minimal yes/no question used 268 input tokens; two used 276. A request containing one yes/no question, one two-option Choice, and one two-level Score used 318, matching the sum of their measured contributions above the shared overhead. This fits a common prefix plus question suffixes, though accounting alone does not identify a computational graph.4

A more informative experiment moves evidence between those regions. The state initially said:

A sibling question contained:

The probe asked which code another question mentioned, with ZEBRA-7741, two distractors, and none as choices. With the secret in the sibling question, its reported probability was 0.00. Removing that sibling produced the same result. Putting the declaration in the state instead raised it to 0.90–0.92. These were five repeats per condition (visibility in the probe records).5

This is a useful intervention: moving the declaration across an API boundary changes its effect. It supports behavioural isolation between questions and access to the shared state. It does not expose the exact attention mask. Separate model calls, a tree mask, or another mechanism that restricts information flow could produce the same result. The probe’s wording also asks about “another question” even in the state condition, so it is not a pristine test of literal instruction following.

The serving measurements add another piece. Up to about 100 questions, server time barely changed. Beyond that it rose steadily, and token for token, question text cost roughly twice as much as state. That is consistent with computing the state once and batching the question work.6

There is a figure here in the original · See it in the original
Figure 2. Server time as a request grows in two ways: a longer state with one question (green), or 1 to 1,500 questions with a short state (purple). Both panels use the same time scale. Each size was requested 8 times, one request at a time, in shuffled order. Grey crosses are individual requests; the solid line joins the median at each size and the dashed line the fastest request. Both grow with the work, but 1,500 questions still return in a few hundred milliseconds. Times come from the API’s upstream service header, which includes shared-service overhead; they are not a hardware benchmark. Methods ↗

These are server-reported upstream durations, not local laptop timings. They include whatever work and waiting the upstream service includes, and the service was shared with other users.

Jev enforces two limits. Each branch (the state plus one question) is capped at roughly 32,768 tokens, and the whole request at roughly 65,536. The request limit counts the state once: a 23k-token state with 5,000 questions fits within it. If each question processed its own copy of the state, that request would be over 100 million tokens. The pair fits a single packed sequence of up to 2¹⁶ tokens, holding the state once and every question after it, with each branch limited to a 2¹⁵ context window.22

A prefix KV cache with separate causal suffixes is the natural implementation. Hydragen describes efficient attention for sequences sharing a prefix; DeFT develops attention for tree-structured inference. These establish that the serving pattern is practical. They are prior art, not evidence that TypeSafe uses either library.78

This design also clarifies an apparent contradiction: isolated questions can still be evaluated together on the same accelerator. “Parallel” describes their scheduling and lack of answer dependencies. It need not mean one GPU per question.

05

Options are read together as a list

Adding an irrelevant “bad weather” option cut the log-odds between two existing options from 0.38 to 0.11, in every one of ten random blocks: fixed independent scores are ruled out.

3. A causal backbone¶

The experiments can’t tell a causal decoder from a bidirectional encoder: in both, the final decision can read the whole input. I assume a causal decoder anyway, for good reason. Jev’s breadth of knowledge (84.6% on MMLU-Pro) requires frontier-scale pretraining, every model at that scale is a causal decoder, and TypeSafe describes RLCD as post-training a pretrained language model. A bidirectional Jev would mean either a far weaker base or converting a decoder at extra cost, while giving up the shared-prefix caching that causal serving provides. That would be surprising, but it can’t be ruled out from the outside.1215

Which pretrained model is unknown, and the tokenizer doesn’t reveal one. Jev’s token counts match none of the 192 public tokenizers we tested across 415 probes. It splits every digit individually and looks up whole chunks before merging: 8 as count as one token, but 16 count as four. Its vocabulary tracks OpenAI’s o200k closely, since every string Jev counts as a single token is also a single o200k token, yet digit splitting and several merges rule o200k itself out. The closest public match, Qwen, agrees on 348 of 415 probes. That rules out an unchanged public tokenizer, not a public base model: a replaced vocabulary, continued pretraining or distillation could each explain it, as could an API that counts tokens differently from the model.18

The experiments do show what the decision can read. I placed a reference card among a question’s options and asked Jev to pick the option whose condition the card satisfies. Here is one exact option set:

The instruction was: “Read the reference card and select the one option whose condition is satisfied.” Changing the reference value to indigo switches the correct answer while leaving the selectable options unchanged.

I tested both values, all six option permutations, and a second template using route = east/west, with two repeats. A paired control put the reference in the shared state instead. That gave 48 option-reference trials and 48 state-reference controls.20

Jev can use information placed after the candidate descriptions. With the reference last, it selected the correct option in every trial, with mean correct-answer probability about 0.88.

There is a figure here in the original · See it in the original
Figure 3. Probability Jev 1.13.0 gave the correct option, by where the reference card sat. Filled cells mark the card (α) in each option order; the last row moves it into the shared state and pools all six orders. Grey crosses are individual requests (two tasks × two card values × two repeats per order); coloured ticks mark the mean. With the card last, every request was correct. With the card first or in the middle, probabilities scattered widely around 0.5, although the final decision could always read the card. That is order sensitivity, not a recovered attention mask. Recorded 17 September 2026. Request payloads and answers ↗

The value-switch control matters as much as the position. With the card last, changing only amber to indigo changes which earlier option wins, although those earlier descriptions and the state stay identical. A model that independently scores each option from its own text and the state, then merely normalises the scores, has no route for that fact to change the earlier options’ relative ranking. The results support a path through which options influence the joint decision.20

This fits any readout computed after the whole list, including both designs compared in section 4, as well as a separate option-mixing stage. The remaining errors show position-sensitive processing on these two templates; they do not identify a unique cause.

Diffusion is unnecessary for this computation, and nothing in these experiments requires iterative denoising. The defensible architectural inference is narrower: the answer computation has access to the full option list. The next experiment tests whether it actually uses that joint context.

4. Let the options interact before choosing¶

Within a question, the evidence points to a different information boundary: the alternatives are read together as an ordered list, followed by one decision position.

Why allow that interaction? Options such as “none of the above” depend on the other choices. Even ordinary alternatives can clarify a question. “Payments”, “account access”, and “other” define a different decision from “bank”, “payment provider”, and “customer”. A listwise representation lets the model interpret that distinction before producing the distribution.

The strongest evidence is an experiment with an irrelevant extra option.

Start with four possible causes of a payout failure: bank, provider, customer, and unknown. Then append weather: Bad weather caused it. If every original option receives an independent, unchanged logit and the server applies the same softmax temperature, adding a fifth option changes the normalisation but cannot change the odds between two existing options:

The common denominator cancels. This gives us a specific, falsifiable prediction.

The original study found a shift from approximately +0.49 to +0.08.10 To check whether this survived ordinary request variability, I repeated the experiment in ten randomised blocks. Each block included the four-option baseline, an identical four-option control, a five-option version with weather appended, an identical five-option control, and a five-option version whose added description changed from “Bad weather caused it” to “Wild birds caused it”. Each request contained one question.21

The expansion result replicated. Pooling the two identical requests for each condition within each block, mean log-odds fell from +0.38 to +0.11. Every block showed a decrease; the average change was −0.28, with a descriptive 95% paired t interval of approximately −0.36 to −0.19. The pooling uses the control requests to reduce ordinary request noise rather than treating duplicate outputs as independent experiments.21

There is a figure here in the original · See it in the original
Figure 4. Does adding an irrelevant option change the odds between two existing ones? Each row is one randomised block. Grey crosses are individual requests (two identical payloads per list size); the grey dot pools the four-option requests and the purple dot the five-option requests with “bad weather caused it” appended. If each option kept a fixed score and the softmax temperature stayed the same, the shared denominator would cancel and the two dots would coincide. The interval is a paired t interval across the ten blocks (9 degrees of freedom), from an exploratory study of one scenario; probability rounding, request noise and a list-dependent temperature remain possible contributors. It shows the options interact, not where in the model. Requests and answers ↗ · Summary ↗

This is evidence against fixed independent logits followed by an unchanged softmax. It does not uniquely identify the mechanism. Changing the added description while keeping five options gave a smaller, inconclusive shift: its paired interval included zero. A set-dependent temperature remains possible, alongside content-dependent mixing.

A readout that sees the complete list explains this naturally: adding an option changes the context it reads. The FIRST listwise ranking method works the same way, extracting a ranking from first-token logits instead of generating it token by token.11

Two readouts fit the evidence. A final-position head scores each option slot from the decision token’s representation; a pointer-style scorer compares that representation with each option’s own final hidden state. Both let options influence one another. The API accepts at most 255 options (2⁸ − 1), which suits a fixed 256-slot head, but that limit is enforced by request validation, not the model. With 200 options, a copied answer scored 1.00 at every position and errors didn’t spill onto neighbouring options, which suits a pointer. Neither result is decisive.26

Injected fake options never displaced the real ones, so option boundaries are marked in a way text cannot forge, and an option whose condition is duplicated elsewhere in the list loses probability to its rivals.23

The trade-off is visible in ordinary tasks too: reversing options shifted the probability of a technical-support classification from roughly 0.84–0.89 to 0.93–0.96. This came from the option_order probes.9 For a deployed decision policy, that matters. A threshold near 0.9 could change the action even though the labels and evidence are identical. Permutation tests belong in the evaluation of any implementation of this design.

06

Cheap probabilities can still be bad ones

Calibration asks whether cases labeled 0.8 are right about 80% of the time. On 1,200 MMLU questions the calibration error is 0.0313, but 990 of them sit in the 0.9–1.0 bin.

5. Train the distribution, then calculate confidence¶

The fifth component is the training objective. Direct numerical outputs save decoding work, but a cheap probability can still be a bad probability.

Imagine a collection of cases assigned 0.8 probability of being urgent. Calibration asks whether about 80% really are urgent. It is a property of predictions across cases. We cannot determine whether one prediction is calibrated from whether that particular case turns out well.

TypeSafe calls its training method Reinforcement Learning for Calibrated Decisions, or RLCD. The launch says it optimises for “answers with epistemically honest probabilities on System One tasks”; the company’s primer presents RLCD as a post-training path from pretrained language models.112 The exact recipe is unpublished. My proposed training recipe adapts the transformer and readout to typed decision tasks using an outcome-based objective. That gives the backbone an opportunity to construct representations useful for reliable decisions, not merely fluent completions.

A natural objective is log loss, −log⁡p(y)-\log p(y)−logp(y) for the observed outcome yyy. Another is Brier loss, the squared distance between the predicted distribution and the observed one-hot outcome. Both are proper scoring rules: in expectation, reporting the true conditional distribution minimises the loss. Gneiting and Raftery give the formal definition and theory.13 This explains what such training tries to achieve. It does not establish which loss TypeSafe uses, whether its pipeline is reinforcement learning in a narrow algorithmic sense, or whether every backbone weight is updated.

Properness is also not a deployment guarantee. Finite data, model limitations, optimisation error, and distribution shift can all leave calibration imperfect. Guo et al. show both the calibration problems of modern neural networks and the usefulness of post-hoc adjustments. Training and post-hoc calibration are compatible mechanisms; the API cannot separate their contributions.14

Observed evidence. The benchmark records let us compare predicted probability with observed accuracy, both in aggregate and within probability bins. The chart shows those checks. Agreement of the averages alone is weaker evidence than agreement within bins: overconfidence in one group can cancel underconfidence in another. On the 1,200-item MMLU sample, ten-bin expected calibration error was 0.0313 (bin definitions and item-level predictions). Most predictions were concentrated near certainty: 990 fell in the 0.9–1.0 bin.15

There is a figure here in the original · See it in the original
Figure 5. Does a given probability match how often Jev is right? Both panels plot observed accuracy against the probability Jev gave its chosen answer, recomputed from recorded outputs rather than the API’s separate confidence field; points on the dashed diagonal are perfectly calibrated. Purple: 1,200 MMLU items grouped into probability bins, labelled with each bin’s item count (probabilities rounded to two decimals before binning). Rust: newly generated maths problems, one point per family. Vertical lines are 95% Wilson intervals, which cover sampling noise only, not benchmark selection or training exposure. Expected calibration error (ECE) weights each bin’s gap from the diagonal by its share of items. Family averages agreeing is weaker evidence than bins agreeing, since over- and underconfidence can cancel within a family; modular exponentiation is the clear exception, right 56% of the time at a mean probability of 35%. Aggregate evidence ↗ · Reliability data ↗

The small fresh-math study adds useful variation. On generated three-digit multiplication problems, accuracy was 86.7% and average top probability 0.83. On two-step word problems, accuracy fell to 32% and average top probability to 0.30. The model was less confident on the harder task (fresh_math_results, with 30 multiplication and 25 word-problem items).15 That is encouraging, although small category-level averages cannot establish calibration for every kind of unseen problem.

Those results also show why a public benchmark score is an imperfect measure of what the model knows. MMLU-Pro accuracy was 84.6%; newly generated word problems were much harder.15 Differences in task structure, distractors, difficulty, and training exposure could all contribute. That gap does not establish benchmark contamination. Fresh wording also does not make the underlying mathematical skill or factual knowledge unseen.

There is a separate, unusually clear finding about the API’s field named confidence. The official adapter computes Choice confidence from a normalised distribution, for K>1K > 1K>1, as:

For three options with a maximum probability of 0.8, this gives 0.7. The adapter handles the one-option case separately, returning 1. It measures how far the leading answer stands above a uniform distribution. It is not another learned estimate that the answer is correct. The Score type uses a different formula reflecting distance from the modal level.16

In the proposed system, training produces the predictive distribution; ordinary arithmetic produces this summary field. Keeping those two objects separate prevents a common conceptual mistake: a concentrated distribution can still be confidently wrong.

6. Sparse capacity¶

I expect Jev to use a sparse mixture-of-experts transformer. At selected layers, a router sends each token through a small subset of feed-forward networks, so the model can store many parameters while activating only some of them for each token: the conditional-computation idea demonstrated by Shazeer et al.’s sparsely gated MoE layers.17

Sparse experts can’t be observed from outside, but they are the likely choice. A prefill-only model is limited by compute, which is exactly what sparse routing saves. MoE’s usual serving costs mostly disappear: there is no token-by-token decoding, where memory bandwidth dominates and most experts end up active anyway, and no long-lived KV cache competing with expert weights for memory. The measurements point the same way. Jev processed about 30k tokens in roughly 160 ms; a dense 70B model on an 8×H100 node would need around a second, while a MoE with about 10B active parameters fits. And most of the strongest recent base models (DeepSeek-V3, Qwen3, GLM-4.5, Kimi K2, gpt-oss) are MoE. Specialised hardware could let a dense model match the speed, and benchmark scores may overstate how much knowledge the model holds, so this remains an inference, not a measurement.615

Nothing else in the reconstruction depends on it. Swapping in a dense transformer would leave the interface, shared state, isolated branches and readout exactly as described.

07

Every inference labeled with its strength of evidence

Question branches are batched as independent jobs. The close sorts claims into three layers: direct probability output (published), question isolation (observable), shared caching and sparse experts (inferred).

7. Schedule branches as a batch, not a conversation¶

The final component is a serving engine that treats question branches as independent work items. Their suffixes can be packed into batches while reading shared state representations. Application code then associates the numerical outputs with question identifiers and serializes the response.

The measurements reveal small differences between repeated identical answers, including between duplicate questions within one request. This means API-level determinism should not be assumed (noise, dup, and determinism).19 It does not imply that the model generates or samples text: numerical kernels, dynamic batching, routing, or deliberate randomness can all affect a direct readout.

Response key orders also varied in a small number of recurring patterns.19 Multiple workers with different hash ordering are a plausible explanation. However, this side channel does not identify the worker count, establish where the KV cache lives, or tell us which numerical precision is used. Those are implementation details the available observations cannot resolve.

What matters to the proposed architecture is the absence of a dependency chain between answers. The model does not need to finish writing the queue classification before beginning the urgency estimate. Both depend on the state; neither consumes the other’s generated answer.

There is still a dependency limit. If a later question genuinely needs an earlier answer, the application must introduce another decision stage or express the joint decision in one question. Sharing context does not remove the logical structure of the workflow.

What would change my mind?¶

This reconstruction makes different kinds of commitments. Direct probability outputs are publicly described. Question isolation and option-order effects are observable behaviours. KV sharing, causal attention, final-position or pointer-style readouts, and sparse experts are progressively more specific explanations.

The reference-card experiment settles one question: the decision can use options placed last. The fake-option test shows that tricks in the input format cannot forge option boundaries. Broader relational tasks could constrain the representation further, although behavioural success alone would still not uniquely identify an attention mask.

For option handling, the randomised follow-up replicates a choice-set effect, but the fixed-size description intervention remains inconclusive. More templates and independent request blocks could distinguish a shared temperature change from content-dependent interactions. A task of middling difficulty at 200 options could separate the slot head from the pointer scorer. For calibration, held-out workflow data and repeated evaluations under shift would matter more than another aggregate benchmark score. Confirming sparse experts would probably require a disclosure or evidence beyond this API.

My best reconstruction of Jev remains the one in the opening diagram: a causal transformer with a shared state prefix, isolated question suffixes, listwise option processing, typed numerical readouts, and training directed at predictive distributions. Sparse experts are the likely backbone, although nothing else in the design depends on them.

Its usefulness comes from matching the computational graph to the job. A decision service needs to read evidence, compare permitted outcomes, and expose uncertainty. A transformer can do that without turning every decision into a sentence first.

Methods¶

This essay is based on a 17 September 2026 investigation of jev-1.13.0, using one early-access account and one observed service region. The source study contains 1,029 instrumented probe records (including the 190 generated-math items), 6,800 benchmark records, and separate factual checks. Follow-up studies added 146 relational and option-interaction requests (trials, summary), 311 token-accounting requests, 445 tokenizer-fingerprint requests, 192 latency requests, 148 option-count latency requests, 181 option-position requests, 105 fake-option requests and 35 context-limit requests. Each is linked from the references, with exact requests and sanitised responses. Repeated benchmark configurations share underlying items; these counts are not counts of independent problems.

The downloadable evidence bundle records the observations used in this essay. API examples in the opening section are schematic. Quoted visibility and reference-card prompts come from the probe scripts and saved follow-up requests. Numerical observations are specific to this model version and test campaign.

Latency figures come from the x-envoy-upstream-service-time response header. They are upstream service durations, with unknown queueing and execution boundaries, rather than isolated model timings. The latency figure’s sweeps were run one request at a time in shuffled order; none of the studies controlled server load. Local wall-clock measurements are not used as architectural evidence.

Probabilities were generally returned at two-decimal precision. Duplicate questions in one request share conditions and may have correlated errors. The MMLU calibration figure uses ten equal-width bins: [0, 0.1), [0.1, 0.2), and so on, with 1.0 included in the last bin. Expected calibration error is the sample-weighted absolute difference between accuracy and average top probability in each bin. Estimates depend on sample selection, binning, and response rounding. The evidence supports claims about the tested distributions, not guaranteed calibration across future customer workflows.

Sources and related work¶

Experimental references identify the original script tags so that each observation can be located in the evidence bundle. Paper citations establish the proposed mechanisms and their precedents; they do not establish that Jev uses them.

Sources

- [1] TypeSafe (2026). Introducing System One Models and Jev. Primary source for the parallel-output claim and stated RLCD objective.

- [2] TypeSafe. Full API documentation, accessed 17 September 2026. Typed questions, response distributions, and API contract.

- [3] Vaswani et al. (2017). Attention Is All You Need. Decoder masking, attention, and the linear/softmax output layer.

- [4] API experiments: type_preamble and outputs. Token-accounting additivity, identifier changes, and the 255-option response.

- [5] API experiment: visibility. Five repetitions each with the secret in a sibling question, absent from that sibling, and in the state.

- [6] Latency sweeps (192 sequential requests). State length and question count, 8 shuffled repeats each; server-reported upstream durations.

- [7] Juravsky et al. (2024). Hydragen: High-Throughput LLM Inference with Shared Prefixes.

- [8] Yao et al. (2024). DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference.

- [9] API experiment: option_order. Ordinary ticket order sensitivity.

- [10] API experiment: iia. Three requests per condition, each with forty duplicate questions; original, appended, and prepended choice sets.

- [11] Reddy et al. (2024). FIRST: Faster Improved Listwise Reranking with Single Token Decoding.

- [12] TypeSafe. Machine learning primer. Primary description of the RLCD post-training path and calibration contract.

- [13] Gneiting and Raftery (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. Journal of the American Statistical Association 102(477):359–378.

- [14] Guo et al. (2017). On Calibration of Modern Neural Networks.

- [15] Benchmark and generated-math records: accuracy and average predicted probabilities. ; MMLU reliability analysis: bin definitions, ECE, Wilson intervals and 1,200 item-level predictions.

- [16] TypeSafe. Official Python adapter, confidence_metrics.py, revision fb52b103. Choice and Score confidence formulas (read 17 September 2026).

- [17] Shazeer et al. (2017). Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

- [18] Tokenizer fingerprint experiment (445 requests). Run-length, vocabulary and pre-tokenisation probes, compared against 192 public tokenizers.

- [19] API experiments: noise, dup, and determinism. Repeated probabilities and response-key orders.

- [20] Follow-up relational experiment (96 requests). Two templates, two reference values, six permutations, two locations, and two repeats; exact requests and sanitised responses.

- [21] Follow-up option-interaction experiment (50 requests). Ten randomised blocks of base4, null4, append5, replace5, and null5; paired changes and standard errors.

- [22] Context-limit experiment (35 sequential requests). Per-branch and whole-request token limits, with accepted and rejected boundary cases.

- [23] Fake-option injection experiment (105 requests). Seven delimiter formats, saturated and ambiguous base tasks, and full probability vectors.

- [24] Token-accounting experiments (311 requests). Question-ID length, batch size, state difficulty, word IDs, and matched state versus ID strings.

- [25] Option-count latency experiment (148 requests). One or 20 questions with 2–200 options, short and long labels, plus input/output decoupling controls.

- [26] Option-position experiment (181 requests). Option-count cap, correct answer moved across 10-, 50-, 200- and 255-option lists.

Where Indigo landsFurther

Indigo's conclusion

Jev is the commercial answer to “verification can't be skipped”: rather than skip verification, it sells decision signals calibrated against real outcomes. But it holds only for fast decisions with real outcomes to train on; the half without a judge is still exposed.

What to remember

  1. It solves two things at once: unreliable decision signals, and compute wasted generating them.
  2. A readout probability isn't text describing a probability: a generated “91%” is a string, a classifier's 0.91 is an entry in its predicted distribution; either can be miscalibrated.
  3. The reconstructed architecture (the author calls it speculation): a causal transformer, likely a sparse mixture of experts, with a shared-state cache and isolated questions.
  4. RLCD is reinforcement learning for calibrated decisions, aimed at honest probabilities; calibration error on 1,200 MMLU questions is 0.031.

Claims you can check later

ClaimWhoWhen we will knowHow firm
Jev doesn't decode token by token; it reads the input once and reads out probabilities directlyArcher (from TypeSafe's claims and latency tests)CurrentTypeSafe's first-hand claim, backed by first-hand experiments
The base is a causal transformer, likely a sparse mixture of expertsArcherCurrentBlack-box inference; the author calls it speculation
Options within a question influence each other (not independent scores)Archer (irrelevant-option test)CurrentFirst-hand experiment, replicated in ten blocks, average −0.28
Jev is reasonably calibrated on the tested distribution (MMLU calibration error 0.031), with no guarantee for future customer workflowsArcherCurrentFirst-hand experiment; the author explicitly doesn't extrapolate
Which pretrained model it is built onUnknownToken counts match none of 192 public tokenizers; can't be pinned down

Back on the long-running theses

confirms

Verification can't be compressed: bottleneck, moat, breaking point Jev is this view's commercial answer: not skipping verification, but selling outcome-calibrated decision signals as the product.

adds to

Harness: the moves get eaten, the interface stays Jev strips out text generation and keeps a typed decision interface; the moat is evaluation, not the base model.

confirms

Ivan Zhao (Notion CEO): Steel, Steam and Infinite Minds Ivan describes the half without a judge; Jev takes the trainable half: two sides of the same view.

adds to

Google's first large-sample study of AI in science People who save time mostly spend it checking AI output; Jev is a supply-side product answer to that pain.

adds to

Noam Brown (OpenAI) on the OpenAI–Hugging Face incident, from inside One is about how fragile it is to monitor a model's stated confidence; the other about skipping generation and training trustworthy probabilities directly.

What would change my mind

calibrated probabilities that are trustworthy even where no real outcomes exist to train on. Then verification wouldn't be split in two; it would be skipped.

Finished. Indigo's take on this piece is in two places: