Mind · In / Out · In · 文章

为 Apple Silicon 优化端侧推理

Optimizing On-Device Inference for Apple Silicon

Perplexity 博客 · 2026-09-02

MacBook 上,35B 开源模型跑赢 Apple 自家框架;真正的信号是撞上了内存带宽的墙。

Indigo 的结论

快了多少只对这一个模型、这一块芯片成立,搬不走;真正的信号是本地 decode 已经撞上内存带宽的墙,软件榨干之后,再提速得靠硬件,而不是更聪明的代码。

怎么读这篇 这是厂商的工程报告,有展示实力、推销 Hybrid Compute、顺带招人的用意,但技术上非常诚实:做了对照实验,公布了失败的尝试,核对了结果是否一致,也测了硬件极限,可信度高。kernel 细节不必细读,重点是从中提炼出的三个信号。

需要记住的几件事

  1. 专门优化能多换来约 20% 到 35% 的速度,代价是每个模型、每块芯片都要单独做一遍工程。
  2. 「专用胜过通用」正一层层往下走:模型权重、上下文、推理引擎;通用的那一层会变成大路货。
  3. Apple Silicon 不是缩小版的数据中心 GPU,而是完整的本地推理平台。

拆解 · 6 步

  1. 01

    专用引擎在一台 Mac 上跑赢通用框架

    为 Qwen 定制的 kernel,不经 PyTorch 和 MLX;十档测试平均 prefill 快 1.23 倍,decode 快 1.35 倍。 读这一段原文 →

  2. 02

    三种负载形状决定怎么优化

    MoE 的专家分组大小不一,注意力要读越来越长的缓存,Gated DeltaNet 是固定大小的循环;prefill 能复用权重,单用户 decode 不能。 读这一段原文 →

  3. 03

    三条策略:分阶段、少搬数据、看形状

    按推理阶段选矩阵或向量路径,把专家、循环和注意力的工作都变成最少的数据搬运,再按实测的负载形状挑布局。 读这一段原文 →

  4. 04

    prefill:反量化并进矩阵乘,路由不回 CPU

    512 token 的提示上,反量化挪进分组矩阵乘,prefill 提速 77.4%;路由留在 GPU 上再提速 89%;循环状态放进寄存器再提速 5.6%。 读这一段原文 →

  5. 05

    decode:每个 token 少搬字节

    token 交接不回 CPU,互不依赖的 kernel 同时跑,省掉中间写入,GQA 打包把 8 次缓存读取减成 2 次,长上下文换成固定分块布局。 读这一段原文 →

  6. 06

    主要运算已经贴着硬件极限

    投机解码等一批改动没让整体变快;结论是推理引擎要同时贴合模型结构和硬件的计算、内存路径。 读这一段原文 →

对 Rewired Index 意味着什么

端侧和混合推理是一条独立的主题,隐私、成本、延迟三样都指向它。Perplexity 未上市;Apple 是平台底座;另一头是内存带宽的供应链。

什么会让我改口

有人只靠软件,把同一块芯片上单用户 decode 的速度再提高一个量级;或者这套专门优化被证明能直接用到别的模型上。

怎么读这篇

这是厂商的工程报告,有展示实力、推销 Hybrid Compute、顺带招人的用意,但技术上非常诚实:做了对照实验,公布了失败的尝试,核对了结果是否一致,也测了硬件极限,可信度高。kernel 细节不必细读,重点是从中提炼出的三个信号。

拆解 · 6 步
  1. 专用引擎在一台 Mac 上跑赢通用框架
  2. 三种负载形状决定怎么优化
  3. 三条策略:分阶段、少搬数据、看形状
  4. prefill:反量化并进矩阵乘,路由不回 CPU
  5. decode:每个 token 少搬字节
  6. 主要运算已经贴着硬件极限
01

专用引擎在一台 Mac 上跑赢通用框架

为 Qwen 定制的 kernel,不经 PyTorch 和 MLX;十档测试平均 prefill 快 1.23 倍,decode 快 1.35 倍。

Apple Silicon 上的混合计算(Hybrid Compute)把一个任务在云端的前沿智能和 Mac 上的本地模型之间编排。云端模型负责检索和推理,本地模型处理 Mac 上的私有文件和应用。

要让这种分工感觉无缝,本地推理必须跟得上任务的其余部分。这需要一个既能快速处理 prompt、又能维持高 token 生成速率的引擎。

Lily 是我们的轻量本地推理引擎,专为 Apple Silicon 和 Qwen3.6-35B-A3B 而造,对 prefill 和 decode 分别做了优化。一个独立的 demo 已在 GitHub 上公开。

引言

在 Mac 上跑 LLM 的常见做法是用 MLX,也就是 Apple 为 Apple Silicon 开源的机器学习框架。它的配套库 MLX-LM 补上了加载各类语言模型并生成文本所需的组件。MLX 和 MLX-LM 合起来,提供了一套现成的通用本地 LLM 推理栈。

Qwen3.6-35B-A3B 是一个稀疏的混合模型:它用专家混合(MoE)路由,并把固定大小的循环状态和全注意力结合起来。这些架构选择减少了所需的计算量,但也造出了不规则的负载。不同 token 会路由到不同的专家权重,而循环状态本质上是顺序的。

MLX-LM 已经会为不同推理阶段和常见负载形状挑选优化过的 kernel,但它那些可复用的算子必须支持很多种模型架构。一个专门为 Qwen 而造的引擎,可以在模型层和运行时层做专门化,围绕这个模型的固定结构去协调 kernel、数据搬运和调度。

Lily 在单个进程里把这种专门化从头做到尾。一个 Rust 运行时负责加载模型检查点、管理会话状态和生成循环,一个 OpenAI 兼容的 chat-completions API 负责接收请求并流式吐出 token,定制的 Metal kernel 执行 Qwen 专属的运算。执行路径里既没有 PyTorch,也没有 MLX。

FIG. 01 两套推理栈里通用性和专门化各自所处的位置。MLX-LM 把模型描述成可组合的 MLX 数组运算,再由 MLX 通过可复用的 kernel 调度。Lily 则把模型结构、分阶段的执行计划和 kernel 选择,全部放进一个围绕 Qwen 和 Apple Silicon 构建的 Rust 运行时里。我们把 prefill 和 decode 的性能分开测。prefill 吞吐反映引擎处理 prompt 有多快;decode 吞吐反映它生成输出 token 有多快。

我们在一台 MacBook Pro 上给 Qwen3.6-35B-A3B 跑基准测试,机器是 M5 Max、40 核 GPU、128 GB 统一内存。在 prefill 的十档 prompt 长度和 decode 的十档上下文长度上,从 256 到 128K token(K = 1,024),这个引擎平均达到 MLX-LM 的 1.23 倍 prefill 吞吐和 1.35 倍 decode 吞吐。在 4K token prompt 加 4K token decode 上下文时,定制引擎达到每秒 5,749.9 个 prefill token 和每秒 186.6 个 decode token,而 MLX-LM 是 4,737.5 和 140.9。在一次多轮会话里,这些省下来的时间会随每一次额外的模型调用累积。

FIG. 02 从 256 到 128K token 的十个等权长度上的算术平均吞吐。Lily 的 prefill 平均是每秒 4,156 个 token,MLX-LM 是 3,388(1.23 倍);decode 平均是每秒 170.0 个 token,MLX-LM 是 126.4(1.35 倍)。prefill 变的是 prompt 长度,decode 变的是上下文长度。接下来我们解释 Qwen 的架构如何在 Apple Silicon 上创造出模型特有的优化机会。然后我们逐一讲由此产生的 prefill 和 decode 改动。我们也会讲到继续优化从哪里开始不再划算,最后用一次对 MLX-LM 的端到端对比收尾。

02

三种负载形状决定怎么优化

MoE 的专家分组大小不一,注意力要读越来越长的缓存,Gated DeltaNet 是固定大小的循环;prefill 能复用权重,单用户 decode 不能。

Apple Silicon 上 Qwen 特有的优化机会

Qwen 造出三种不同的负载形状

Qwen3.6-35B-A3B 有 350 亿参数,但每个 token 只激活其中约 30 亿。一个路由器给 256 个专家子网打分并选出 8 个,另外还有一个处理所有 token 的共享专家。这种稀疏 MoE 设计减少了计算,却产生了不均匀的工作量:不同专家收到的 token 数量不同,而每个 token 需要的权重来自不同的专家组合。

Qwen 还把 10 层全注意力和 30 层 Gated DeltaNet 混在一起。这两类层保留早前信息的方式不同。

注意力层用的是分组查询注意力(GQA)。Qwen 有 16 个 query 头和 2 个 key–value(KV)头,每 8 个 query 头共享一个 KV 头。共享让 KV 缓存更小,也让缓存下来的数据能在多个 query 头之间复用。缓存仍然要为每个 token 存新的 key 和 value,所以随着上下文变长,每一步 decode 要读的数据都更多。

Gated DeltaNet 则把早前的信息压进一个固定大小的循环状态。一个学出来的门控决定保留多少现有状态,而一次 delta 更新把当前 token 的信息并进去。模型把这些更新定义成递归的,所以每个 token 都依赖前一个 token 产生的状态。不过在 prefill 阶段,引擎可以用两种方式算同一件事。它可以直接顺着 token 扫过去、把状态一路带下去,也可以把更新重组成块,从而暴露出更多矩阵运算和 token 级并行。哪种更快取决于模型维度、负载和硬件。

这些结构合在一起,造出三种计算模式:不均匀的专家分组、在不断增长的缓存上做注意力,以及一个既能直接算也能分块算的固定大小递归。

Apple Silicon 为不同负载提供了不同路径

prefill 一次处理很多 prompt token 的激活行。这里考虑的本地负载通常一次只解码一个请求(batch 1),每一步只处理一个新行。这个差别改变了同一份模型权重的用法。prefill 可以把每一块权重在成百上千行上复用。decode 基本做不到,因为每个新 token 都要再过一遍权重。

Apple Silicon 把 CPU 和 GPU 放在统一内存后面,那是一个两者都能访问的单一物理内存池。这让模型可以常驻,不必再维护一份单独的 GPU 副本,但它并不让数据搬运变成免费。读权重和中间值仍然消耗内存带宽,而寄存器和其他片上存储虽然更快,却小得多。

M5 的 GPU 还提供了不同的计算路径。prefill 的线性层用通用矩阵–矩阵乘法(GEMM),把一个权重矩阵一次作用到很多行上。兼容的 GEMM 可以通过 Metal 4 张量运算,用上每个 GPU 核心里的 Neural Accelerator。batch-1 的 decode 则用通用矩阵–向量乘法(GEMV),把同样的权重作用到一行上。由于几乎没有权重复用,GEMV 主要受内存带宽限制,比起那些为高数据复用的矩阵运算而设计的 Neural Accelerator,它更适合 GPU 的向量算术逻辑单元(ALU)。

这些执行路径不是 Lily 独有的。MLX 也在同一片统一内存上工作,并按负载形状挑选优化过的矩阵和向量 kernel。MLX-LM 的 Qwen 实现已经会把专家工作分组、用一个融合的循环 Metal kernel 算 Gated DeltaNet,并使用感知 GQA 的注意力。这些能力是在 Apple Silicon 上高效跑 Qwen 的共同起点。

03

三条策略:分阶段、少搬数据、看形状

按推理阶段选矩阵或向量路径,把专家、循环和注意力的工作都变成最少的数据搬运,再按实测的负载形状挑布局。

优化策略

Lily 的适用范围更窄,这让它能围绕 Qwen 确切的架构和维度去协调这些共享的执行路径。它使用分阶段的 GPU 路径,把 Qwen 的专家、循环和注意力负载映射成最少的数据搬运,并根据实测的负载形状挑选 kernel 和布局。策略分三部分:

让 GPU 路径匹配推理阶段。当 prefill 能把权重在很多行上复用时,用面向矩阵的执行;当 batch-1 的 decode 每次只处理一行时,用面向向量的执行。

把 Qwen 的结构映射到 GPU 上,同时把数据搬运降到最低。权重在用到之前保持压缩状态,路由到各专家的工作不必回 CPU 就能组织好,Gated DeltaNet 的状态在整个循环扫描期间留在片上,分组查询注意力共享的 KV 数据得到复用。

让 kernel 适配负载形状。在每个阶段内部,根据可用的行数、这些行在各专家之间的分布、运算的维度和当前上下文长度,选择 tile 尺寸、执行布局和注意力路径。

下面几节解释这些选择。对于那些在 M5 Max 上做了配对消融实验的优化,我们通过比较除了被研究的那项优化之外完全相同的引擎配置,来估计它的效果。因为这些实验比的是我们引擎自己的不同版本,它们解释的是机制,而不是把最终结果对 MLX-LM 做分解。

04

prefill:反量化并进矩阵乘,路由不回 CPU

512 token 的提示上,反量化挪进分组矩阵乘,prefill 提速 77.4%;路由留在 GPU 上再提速 89%;循环状态放进寄存器再提速 5.6%。

prefill:复用权重,把路由留在 GPU 上

prefill 一次暴露出很多 token 行,但 Qwen 把这些行不均匀地路由到各个专家,并且沿着序列更新循环状态。它的优化分三组:围绕被路由的行来组织稀疏的专家工作、把 Gated DeltaNet 的扫描留在片上,以及把长 prompt 切成有上限的块。

FIG. 03 一个有上限的 prefill 块穿过一层 Qwen 时的执行与数据驻留情况。注意力扩展 KV 缓存,而 Gated DeltaNet 把它正在用的循环状态放在寄存器里。专家路由的元数据留在 GPU 上,Q4 权重一直保持打包状态,直到在分组 GEMM 内部被反量化,临时激活值只限于当前这个块。

优化稀疏专家计算

在矩阵乘法过程中反量化权重

Qwen3.6-35B-A3B 的检查点用的是分组仿射 4-bit 量化。每个权重存成一个 4-bit 整数码,而每 64 个权重共享一个 bfloat16 的 scale 和 bias,用来重建它们的数值。这把这个 350 亿参数模型从大约 70 GB 的 bfloat16 权重压成一个 19.4 GB 的检查点,使得把模型常驻在 Mac 上变得实际可行。

用于矩阵乘法的 Metal 4 张量运算吃的是 bfloat16 操作数,而不是打包好的 4-bit 表示。在做乘法之前,GPU 必须把权重重建成 bfloat16。Lily 里优化过的分组 GEMM 一次只转换一小块权重 tile,并且只在片上的 threadgroup 内存里保留结果,时间刚好够它和被路由过来的激活行相乘。累加用 32 位浮点,输出写成 bfloat16。完整展开的权重数组从来不会在统一内存里被创建出来。

在消融实验里,反量化是一个单独的运算:它把 4-bit 权重展开成统一内存里的一个 bfloat16 数组,之后矩阵 kernel 再把那个数组读回来。在 512 token 的 prompt 上,把反量化挪进分组 GEMM,通过消掉这一次中间的写和读,让端到端 prefill 吞吐提高了 77.4%。

把专家路由留在 GPU 上

分组 GEMM 需要分配给同一个专家的激活行被存放在一起。在为每个 token 选出 8 个专家之后,一个直方图统计每个专家收到了多少次分配。一次前缀扫描把这些计数变成起始偏移,一个 scatter 步骤把行放进各自的专家分组,而一张块表列出分组 GEMM 必须处理的那些固定大小的矩阵块。

优化过的路径把整个这一串都放在一个命令缓冲区(command buffer,一批有序的 GPU 运算)里,每个 prompt 块一个。消融实验则会停下来,让 CPU 检查路由的中间结果,再提交下一个运算。把直方图和前缀扫描留在 GPU 上多了两个 kernel,却去掉了每个 MoE 层内部的 CPU–GPU 同步。

在 512 token 的 prompt 上,启用 GPU 常驻路由让端到端 prefill 提高了 89%。这也说明为什么光看 kernel 数量会误导人:更快的那条路启动了更多 kernel,却从不在层内部等 CPU。

让 tile 尺寸匹配专家负载

在 2K token 的 prompt 上,每个 token 都路由到 256 个专家中的 8 个,会产生 16,384 次 token–专家分配,平均每个专家 64 个激活行。实际分布并不均匀:有的专家收到很多行,有的收到很少。

分组 GEMM 把每个专家的输出切成 tile,也就是矩阵乘法输出里的小矩形块。每个 tile 分给一个 GPU threadgroup。在 Apple Silicon 的 GPU 上,一个 threadgroup 包含一个或多个 simdgroup,每个 simdgroup 由 32 个步调一致执行指令的线程组成。

更大的 tile 能把启动开销摊到更多行上,也暴露出更多并行工作,但当一个专家只收到很少几行时,大 tile 的一部分就闲着。所以 tile 尺寸和 simdgroup 数量是耦合的。

一次消融把 tile 固定在 16 行。以此为对照,启用 32 行 tile 加 4 个 simdgroup,在 2K token 时让端到端 prefill 提高了 13.2%。

把循环状态留在片上

在 prefill 阶段,每一层 Gated DeltaNet 都按顺序扫过 prompt,同时把循环状态一路带下去。关掉寄存器驻留后,消融实验用的是分块扫描。在 2K token 的 prompt 上,那条路径每层要搬 256 MiB(mebibyte)的状态,而且反复在 barrier 处让协作的线程停下来——barrier 是同步点,所有参与的线程都必须在那里互相等待。

循环状态是一个矩阵。优化过的 kernel 把每一列分给一个 simdgroup。simdgroup 再把这一列分给它的各个线程,把列一次性装进它们的寄存器,然后带着状态走完整个扫描。线程之间通过 simdgroup 运算交换中间结果,而不是通过 threadgroup 内存——那是一个 threadgroup 共享的片上存储。完成的状态只在扫描结束后才写回去。

状态和它的门控用 32 位浮点格式,因为微小的舍入误差会在一连串顺序更新中累积。query 和 key 的激活值仍然是 bfloat16。

在 2K token 的 prompt 上,启用寄存器驻留的扫描让端到端 prefill 提高了 5.6%。专家 GEMM 占了 prefill 时间的大约 90%。顺序扫描暴露不出足够多可复用的矩阵工作,所以沾不到 Neural Accelerator 的光。

用 prompt 分块给临时内存设上限

运行时把一个长 prompt 当作一连串有上限的块来处理,而不是把每个 prompt token 的临时数据同时留在内存里。模型权重仍然常驻在统一内存里,而循环状态和 KV 缓存把上下文从一块带到下一块。更早的上下文一点都不会被丢掉。

不分块的话,临时激活数组会随整个 prompt 一起变大,并且和模型权重、循环状态、KV 缓存争抢统一内存。分块让同一时间只有一段的临时值是活的,然后在处理下一段之前释放或复用那块存储。这就给峰值工作内存封了顶,让引擎能处理更长的 prompt,而不改变模型的输出。

分块 prefill 在很多引擎里都很流行,在这种内存受限的环境里,它对支撑长的多轮轨迹是关键。注意力层的总 prefill 时间仍然是 prompt 长度的二次方,还要加上重复读取更早块的 KV 带来的一点额外开销。

05

decode:每个 token 少搬字节

token 交接不回 CPU,互不依赖的 kernel 同时跑,省掉中间写入,GQA 打包把 8 次缓存读取减成 2 次,长上下文换成固定分块布局。

decode:把每个 token 要搬的字节数降到最低

batch-1 的 decode 一次处理一个新行。由于几乎没有权重复用,它的吞吐主要取决于引擎为每个 token 搬多少字节。decode 的改动分四组:优化单行权重路径、把每一步都留在 GPU 上、减少中间值和状态的流量,以及高效地读注意力缓存。

FIG. 04 一次 batch-1 decode 步骤的数据流,以及两个减少空闲时间和缓存流量的机制。(A) GPU 把 Q4 权重和模型状态串过注意力、Gated DeltaNet、路由和融合后的专家 kernel,然后把选中的 token 直接写进下一步的输入槽,同时发一份副本给 CPU。(B) 感知依赖的调度让互不相干的 kernel 可以重叠。(C) GQA 打包让 4 个 query 头共享同一次 KV 行读取,把 8 次独立请求减成 2 次共享加载。

优化单行权重路径

MLX 已经会把单行的工作派给专门的矩阵–向量 kernel。因为 Lily 不用 MLX,定制运行时必须自己提供同样的基本策略。我们的行并行 GEMV 是为一行激活设计的。一个 simdgroup 协作产出输出,同时并行读取权重矩阵的不同部分。

把每一步 decode 都留在 GPU 上

把 token 交接留在 GPU 上

每一步 decode 以选出下一个 token 结束;下一步以这个 token 作为输入开始。把这个选择送去 CPU 再送回 GPU,等于给每个 token 加一个同步点。我们的运行时改成在两个命令缓冲区和两个 GPU 常驻的 token 槽之间交替。GPU 选出得分最高的 token,把它的 token ID 直接写进下一步 decode 的输入槽,同时 CPU 去准备后续的工作。

让互不相干的 GPU 工作重叠

在一次记录下来的 batch-1 decode 步骤里,生成一个 token 启动了 795 个 GPU kernel。它们的依赖关系构成 555 个顺序阶段,也就是说有些 kernel 本来可以并发跑。但 Metal 的串行执行模式还是把每个 kernel 依次跑了一遍。

优化过的 decode 路径在一个并发的 Metal pass 里记录真实的数据依赖。只要 GPU 资源允许,互不相干的 kernel 启动就能同时进行。只有在后面的工作需要前面的结果时,才插入一个 barrier。

减少中间值和状态的流量

分开的 kernel 常常会把一个中间值落到内存里:一个 kernel 把临时结果写进内存,下一个再把结果读回来。优化过的 decode 路径融合了四条链:两个专家输入投影和它们的门控激活;专家输出投影和它的路由分数以及共享专家的结果;注意力之前的 query 和 key 准备;还有循环更新和它的归一化。每个融合 kernel 把临时值留在寄存器里,而不是让它们过一遍内存。

融合还缩短了依赖图:一次中间写消失了,保护它的消费者的那个 barrier 也就跟着消失了。

高效地读注意力缓存

合并注意力缓存的读取

每一步 decode,注意力都要从 KV 缓存里读 key 和 value。在消融实验里,相邻的 GPU 线程并不总是请求相邻的字节,这迫使内存系统去服务更多次单独的事务。启用合并加载后,相邻线程会请求相邻的字节,硬件就能把它们的读取合并起来。

在 bfloat16 配置下,合并让 key 的带宽从 33.8 GB/s 提高到 47.9 GB/s,让 value 的带宽从 42.0 GB/s 提高到 61.8 GB/s,并在 3,840 token 的上下文上让端到端 decode 提高了 2.1%。

打包 query 头以复用 KV 行

分组查询注意力让 8 个 query 头共享一个 KV 头。在消融实验里,每个 query 头跑在单独的 simdgroup 里,所以 8 个头各自独立地去请求同一行缓存的 KV。优化过的 kernel 把 4 个 query 头打包进一个 threadgroup,这个 threadgroup 把每一行 KV 只加载一次,然后在 4 次注意力计算之间复用。另一个 threadgroup 处理剩下的 4 个头。

这个技巧通常叫 GQA 打包,它做的算术完全相同、产出的输出字节也完全相同,却把 8 次独立的 KV 请求减成 2 次共享加载。对照没打包的消融版本,它在 32K token 的上下文上让端到端 decode 吞吐提高了 23.8%。

长上下文时切换注意力布局

全注意力层的每一步 decode 都要扫一遍现有的 KV 缓存。固定块布局把那份缓存切成相等的几片,让 GPU 能并行处理。缓存小的时候,它多出来的调度不划算;但随着上下文变长,固定块布局能把工作摊得更均匀。

对这个模型,运行时在 32K token 以下保持通用注意力路径,在 32K 及以上用固定块路径。这个切换适用于每个头有 256 个值、且 8 个 query 头共享一个 KV 头的情形;其他形状仍然走通用路径。一次消融关掉这个切换,始终走通用路径。启用固定块路线后,端到端 decode 在 32K 时提高了 7.7%,在 64K 时提高了 27.4%,在 128K 时提高了 40.2%。

06

主要运算已经贴着硬件极限

投机解码等一批改动没让整体变快;结论是推理引擎要同时贴合模型结构和硬件的计算、内存路径。

继续优化的极限

有些改动改善了某个孤立的运算,却没有改善端到端的推理。

投机解码(speculative decoding)用一个更小的模型提出 token、再由完整模型去验证,它让 batch-1 的 decode 慢了 18%。验证一次处理 2 到 5 行,对这套硬件来说是个低效的形状,而且这些行常常选到不同的专家,增加了要读的专家权重数据量。把起草模型的输出词表缩小,让起草模型的吞吐提高了 4.7–5.1%,但并没有让整个投机循环变快。这个结果是工况特异的:我们在 Blackwell 上的批量 Qwen 部署,在不同条件下是用投机解码的。

其他实验还包括减少 GPU 启动次数、让整个阶段重叠、用更大的 prefill tile、做更大范围的融合、给路由器加速,以及把输出投影和 token 选择合并。没有一项改善了完整的推理循环。

对硬件极限的测量也显示,主要的 prefill 和 decode 运算已经没剩多少余量。MoE 的 GEMM 和 GEMV 分别达到了各自访问模式下最快持续读权重速率的 97.9% 和 90.3%。把稀疏 GEMV 里的算术全部拿掉,吞吐只变了 0.2%,这证实了限制性资源是读权重而不是计算。prefill 的矩阵乘法同样在孤立测量下达到理论矩阵极限的 93%,在被测模型内部达到 80–86%。

端到端性能

端到端对比在两个引擎里加载完全相同的 4-bit 检查点字节,在一台 40 核、128 GB 的 M5 Max 上一次跑一个请求。每一轮里两个引擎交替顺序运行,以减少后台负载和芯片温度变化带来的偏差。我们比的是 MLX-LM 最快的直接生成路径,不是它的 server,所以测量聚焦在模型执行上,而不是服务开销。

这次扫描覆盖 prefill 的十档 prompt 长度和 decode 的十档上下文长度,从 256 到 128K token。prefill 吞吐一开始随着引擎把固定的启动开销摊到更多 token 上而上升。prefill 在 4K token prompt 附近见顶,然后回落,因为随着 prompt 变长,那 10 层全注意力要做更多工作。decode 在短上下文上几乎是平的,等到读不断增长的 KV 缓存开始变得显著时才下滑。在每一个记录到的长度上,定制引擎都更快。

因为专门化的执行可能改变浮点运算的顺序,我们还对着 MLX-LM 检查了数值一致性。在一次 teacher-forced 对比里,两个引擎在 192 个位置上各自从同一段参考前缀预测下一个 token,以防早先的差异影响后面的输入。Lily 的困惑度只高了 0.04%,并且在 96.35% 的被测位置上选出了排名第一的同一个 token。

FIG. 05 Qwen3.6-35B-A3B Q4、batch 1,在一台 40 核、128 GB 的 M5 Max 上,按 prompt 长度看 prefill 吞吐、按上下文长度看 decode 吞吐。在从 256 到 128K token 的十个长度上,Lily 在每一个记录点都更快:prefill 吞吐是 MLX-LM 的 1.12–1.42 倍,decode 吞吐是 1.31–1.37 倍。对比用的是 MLX-LM 最快的直接生成路径。两条横轴都用对数刻度;两条纵轴都不从零开始。

为本地平台而造

Apple Silicon 不是一块缩小的数据中心 GPU。它是一个完整的本地推理平台,有自己的硬件和软件特性。统一内存给单个节点一个非常高的上限,决定它能装下多大的模型和状态。M5 的 Neural Accelerator 吸收 prefill 里稠密的矩阵工作。向量 ALU 处理 decode 里那部分受带宽限制、复用率低的剩余工作。

Qwen 还带来更多专门化的机会:把专家路由和循环状态留在 GPU 上、消掉不必要的中间值、让互不相干的工作重叠、复用共享的 KV 数据,并让 kernel 适配负载形状。

有了针对模型和平台的优化,一台 Mac 就能高效地跑一个大型稀疏模型。后续工作会把覆盖面拓宽到更多模型、芯片和服务负载,并把在这一种配置上验证过的机制,变成更通用的运行时策略。

更大的原则是:让引擎同时匹配模型的架构和硬件具体的计算与内存路径。随着前沿开放权重模型和硬件不断演进,高性能的本地推理会越来越依赖那些为两者量身定制的引擎,而不是那些把两者的差异抽象掉的引擎。

判断收口延伸

Indigo 的结论

快了多少只对这一个模型、这一块芯片成立,搬不走;真正的信号是本地 decode 已经撞上内存带宽的墙,软件榨干之后,再提速得靠硬件,而不是更聪明的代码。

需要记住的几件事

  1. 专门优化能多换来约 20% 到 35% 的速度,代价是每个模型、每块芯片都要单独做一遍工程。
  2. 「专用胜过通用」正一层层往下走:模型权重、上下文、推理引擎;通用的那一层会变成大路货。
  3. Apple Silicon 不是缩小版的数据中心 GPU,而是完整的本地推理平台。

放回主线

证实

约束正从算法层迁到物理层 这条判断在端侧的量化证据:软件榨干之后,提速靠内存带宽,不靠算法。

补充

开放权重的安全政治学:门禁还是竞争 35B 开源模型在笔记本上跑得又快又准,说明开源能力的扩散已经收不回来。

证实

Sarah Guo(Conviction) 云端前沿模型加本地隐私处理,正是「能力普及之后会被用得更多」的产品化。

补充

Lin Qiao(Fireworks):post-training 是你保住品味的方式 同一条「专用胜过通用」的规律出现在不同层,这次到了推理引擎。

对 Rewired Index 意味着什么

端侧和混合推理是一条独立的主题,隐私、成本、延迟三样都指向它。Perplexity 未上市;Apple 是平台底座;另一头是内存带宽的供应链。

什么会让我改口

有人只靠软件,把同一块芯片上单用户 decode 的速度再提高一个量级;或者这套专门优化被证明能直接用到别的模型上。

读完了。Indigo 对这篇的判断在这两处:

Mind · In / Out · In · Essay

Optimizing On-Device Inference for Apple Silicon

perplexity.ai · 2026-09-02

On a MacBook, a 35B open model beats Apple's own framework. The real signal is that it hit the memory-bandwidth wall.

Indigo's conclusion

The speedup holds for one model on one chip and doesn't travel. The real signal is that local decode has hit the memory-bandwidth wall: with software squeezed dry, further gains come from hardware, not smarter code.

How to read this A vendor's engineering report, meant partly to show off, sell Hybrid Compute and recruit. But it is unusually honest: controlled comparisons, failed attempts published, outputs checked for consistency, hardware limits measured. High credibility. Skip the kernel details; the three signals it yields are the point.

What to remember

  1. Specializing buys roughly 20% to 35% more speed, at the cost of separate engineering for every model and every chip.
  2. “Specialized beats general” keeps moving down the stack: model weights, then context, then the inference engine. The general layer becomes a commodity.
  3. Apple Silicon is not a shrunken data-center GPU; it is a complete platform for local inference.

Breakdown · 6 steps

  1. 01

    A purpose-built engine beats the general framework on a Mac

    A Rust runtime plus Metal kernels written for Qwen, with PyTorch and MLX both removed. Across ten settings, prefill averages 1.23x faster and decode 1.35x. Read this part →

  2. 02

    Three workload shapes decide the optimizations

    MoE expert groups vary in size, attention reads an ever longer cache, Gated DeltaNet is a fixed-size loop. Prefill can reuse weights; single-user decode cannot. Read this part →

  3. 03

    Three rules: match the phase, move fewer bytes, read the shape

    Pick matrix or vector paths by inference phase, turn expert, loop and attention work into the least data movement, then choose layouts from measured workload shapes. Read this part →

  4. 04

    Prefill: dequantize inside the matrix multiply, keep routing off the CPU

    On a 512-token prompt, moving dequantization into the grouped matrix multiply sped prefill up 77.4%; keeping routing on the GPU added 89%; keeping loop state in registers added 5.6%. Read this part →

  5. 05

    Decode: move as few bytes per token as possible

    Token handoff stays off the CPU, independent kernels overlap, intermediate writes are fused away, GQA packing cuts 8 cache reads to 2, long contexts get a fixed block layout. Read this part →

  6. 06

    The main operations are already at the hardware limit

    Speculative decoding and several other changes didn't speed things up end to end. The conclusion: an engine has to fit both the model's structure and the hardware's compute and memory paths. Read this part →

What it means for Rewired Index

On-device and hybrid inference is its own theme, pulled by privacy, cost and latency. Perplexity is private; Apple is the platform underneath; the other end is the memory-bandwidth supply chain.

What would change my mind

someone lifts single-user decode on the same chip by another order of magnitude with software alone, or this specialization is shown to carry straight over to other models.

How to read this

A vendor's engineering report, meant partly to show off, sell Hybrid Compute and recruit. But it is unusually honest: controlled comparisons, failed attempts published, outputs checked for consistency, hardware limits measured. High credibility. Skip the kernel details; the three signals it yields are the point.

Breakdown · 6 steps
  1. A purpose-built engine beats the general framework on a Mac
  2. Three workload shapes decide the optimizations
  3. Three rules: match the phase, move fewer bytes, read the shape
  4. Prefill: dequantize inside the matrix multiply, keep routing off the CPU
  5. Decode: move as few bytes per token as possible
  6. The main operations are already at the hardware limit
01

A purpose-built engine beats the general framework on a Mac

A Rust runtime plus Metal kernels written for Qwen, with PyTorch and MLX both removed. Across ten settings, prefill averages 1.23x faster and decode 1.35x.

Hybrid Compute on Apple silicon orchestrates a task between frontier intelligence in the cloud and a local model on the Mac. Cloud models handle research and reasoning, while a local model works with private files and apps on the Mac.

For this division of labor to feel seamless, local inference must keep pace with the rest of the task. That requires an engine that can process prompts quickly and sustain a high token-generation rate.

Lily, our lightweight local inference engine, is built specifically for Apple silicon and Qwen3.6-35B-A3B, with separate optimizations for prefill and decode. A standalone demo is publicly available on GitHub.

Introduction

A common way to run LLMs on a Mac is with MLX, Apple’s open-source machine learning framework for Apple silicon. Its companion library, MLX-LM, adds the components needed to load and generate text with a wide range of language models. Together, MLX and MLX-LM provide an off-the-shelf, general-purpose stack for local LLM inference.

Qwen3.6-35B-A3B is a sparse, hybrid model: it uses mixture-of-experts (MoE) routing and combines fixed-size recurrent states with full attention. These architectural choices reduce the amount of computation required, but they also create irregular workloads. Tokens route to different expert weights, and recurrent states are sequential in nature.

MLX-LM already selects optimized kernels for inference phases and common workload shapes, but its reusable operations must support many model architectures. An engine dedicated to Qwen can specialize at the model and runtime level, coordinating kernels, data movement, and scheduling around the model’s fixed structure.

Lily implements this specialization end to end in a single process. A Rust runtime loads the model checkpoint and manages the session state and generation loop, an OpenAI-compatible chat-completions API accepts requests and streams tokens, and custom Metal kernels execute Qwen-specific operations. Neither PyTorch nor MLX is in the execution path.

FIG. 01Where generality and specialization sit in the two inference stacks. MLX-LM describes the model as composable MLX array operations, which MLX schedules through reusable kernels. Lily instead places the model structure, phase-specific execution plans, and kernel selection in a single Rust runtime built around Qwen and Apple silicon.We measure prefill and decode performance separately. Prefill throughput captures how quickly the engine processes the prompt; decode throughput captures how quickly it generates output tokens.

We benchmark Qwen3.6-35B-A3B on a single MacBook Pro powered by an M5 Max with a 40-core GPU and 128 GB of unified memory. Across ten prompt lengths for prefill and ten context lengths for decode, from 256 to 128K tokens (K = 1,024), the engine averages 1.23× MLX-LM’s prefill throughput and 1.35× its decode throughput. At a 4K-token prompt and a 4K-token decode context, the custom engine reaches 5,749.9 prefill tokens per second and 186.6 decode tokens per second, compared with 4,737.5 and 140.9 for MLX-LM. Across a multi-turn session, these time savings accumulate with each additional model call.

FIG. 02Arithmetic mean throughput across ten equally weighted lengths from 256 to 128K tokens. Lily averages 4,156 prefill tokens/s versus 3,388 for MLX-LM (1.23×), and 170.0 decode tokens/s versus 126.4 (1.35×). Prefill varies prompt length; decode varies context length.Next, we explain how Qwen’s architecture creates model-specific optimization opportunities on Apple silicon. We then walk through the resulting prefill and decode changes. We also cover where additional optimization stops paying off before closing with an end-to-end comparison against MLX-LM.

02

Three workload shapes decide the optimizations

MoE expert groups vary in size, attention reads an ever longer cache, Gated DeltaNet is a fixed-size loop. Prefill can reuse weights; single-user decode cannot.

Qwen-specific optimization opportunities on Apple silicon

Qwen creates three distinct workload shapes

Qwen3.6-35B-A3B contains 35 billion parameters but activates only about 3 billion for each token. A router scores 256 expert subnetworks and selects eight, alongside one shared expert that processes every token. This sparse MoE design reduces computation but produces uneven work: experts receive different numbers of tokens, and each token requires weights from a different combination of experts.

Qwen also combines 10 full-attention layers with 30 Gated DeltaNet layers. These two layer types retain earlier information in different ways.

The attention layers use grouped-query attention (GQA). Qwen has 16 query heads and two key–value (KV) heads, with eight query heads sharing each KV head. Sharing makes the KV cache smaller and allows cached data to be reused across query heads. The cache still stores new keys and values for every token, so each decode step reads more data as the context grows.

Gated DeltaNet instead compresses earlier information into a fixed-size recurrent state. A learned gate controls how much of the existing state to retain, while a delta update incorporates information from the current token. The model defines these updates recurrently, so each token depends on the state produced by the preceding token. During prefill, however, an engine can evaluate the same computation in two ways. It can scan through the tokens directly while carrying the state forward, or reorganize the updates into blocks that expose more matrix operations and token-level parallelism. Which approach is faster depends on the model dimensions, workload, and hardware.

Together, these structures create three computational patterns: uneven expert groups, attention over a growing cache, and a fixed-size recurrence that can be evaluated directly or in blocks.

Apple silicon provides different paths for different workloads

Prefill processes many prompt token activation rows at once. The local workload considered here typically decodes one request at a time (batch 1) and processes one new row per step. This difference changes how the same model weights are used. Prefill can reuse each block of weights across hundreds or thousands of rows. Decode largely cannot, since each new token requires another pass through the weights.

Apple silicon places the CPU and GPU behind unified memory, a single physical memory pool accessible to both. This allows the model to remain resident without maintaining a separate GPU copy, but it does not make data movement free. Reading weights and intermediate values still consumes memory bandwidth, while registers and other on-chip storage are faster but much smaller.

The M5 GPU also provides different compute paths. Prefill’s linear layers use general matrix–matrix multiplication (GEMM), applying a weight matrix to many rows at once. Compatible GEMMs can use the Neural Accelerator in each GPU core through Metal 4 tensor operations. Batch-1 decode instead uses general matrix–vector multiplication (GEMV), applying the same weights to one row. With little weight reuse, GEMV is limited mainly by memory bandwidth and is better suited to the GPU’s vector arithmetic logic units (ALUs) than to Neural Accelerators designed for matrix operations with greater data reuse.

These execution paths are not unique to Lily. MLX operates over the same unified memory and selects optimized matrix and vector kernels according to the workload shape. MLX-LM’s Qwen implementation already groups expert work, evaluates Gated DeltaNet with a fused recurrent Metal kernel, and uses GQA-aware attention. These capabilities are the shared starting point for efficient Qwen inference on Apple silicon.

03

Three rules: match the phase, move fewer bytes, read the shape

Pick matrix or vector paths by inference phase, turn expert, loop and attention work into the least data movement, then choose layouts from measured workload shapes.

Optimization strategy

Lily’s narrower scope allows it to coordinate these shared execution paths around Qwen’s exact architecture and dimensions. It uses phase-specific GPU paths, maps Qwen’s expert, recurrent, and attention workloads to minimize data movement, and selects kernels and layouts from the measured workload shape. The strategy has three parts:

Match the GPU path to the inference phase. Use matrix-oriented execution when prefill can reuse weights across many rows, and vector-oriented execution when batch-1 decode processes one row at a time.

Map Qwen’s structure onto the GPU while minimizing data movement. Keep weights compressed until they are used, organize routed expert work without returning to the CPU, retain Gated DeltaNet state on chip through its recurrent scan, and reuse the KV data shared by grouped-query attention.

Adapt kernels to the workload shape. Within each phase, select tile sizes, execution layouts, and attention paths from the available row count, the rows’ distribution across experts, the operation’s dimensions, and the current context length.

The following sections explain these choices. For optimizations evaluated in matched ablations on an M5 Max, we estimate their effects by comparing otherwise identical engine configurations that differ only in the optimization under study. Because these experiments compare versions of our engine against itself, they explain mechanisms rather than decompose the final results against MLX-LM.

04

Prefill: dequantize inside the matrix multiply, keep routing off the CPU

On a 512-token prompt, moving dequantization into the grouped matrix multiply sped prefill up 77.4%; keeping routing on the GPU added 89%; keeping loop state in registers added 5.6%.

Prefill: reuse weights and keep routing on the GPU

Prefill exposes many token rows at once, but Qwen routes those rows unevenly across experts and updates recurrent state through the sequence. Its optimizations fall into three groups: organize sparse expert work around the routed rows, keep the Gated DeltaNet scan on chip, and divide long prompts into bounded chunks.

FIG. 03Execution and data residency for one bounded prefill chunk through a Qwen layer. Attention extends the KV cache, while Gated DeltaNet carries its working recurrent state in registers. Expert-routing metadata remains on the GPU, Q4 weights stay packed until they are dequantized inside the grouped GEMM, and temporary activations are limited to the current chunk.

Optimize sparse expert computation

Dequantize weights during matrix multiplication

The Qwen3.6-35B-A3B checkpoint uses groupwise affine 4-bit quantization. Each weight is stored as a 4-bit integer code, while every group of 64 weights shares a bfloat16 scale and bias used to reconstruct its values. This reduces the 35-billion-parameter model from roughly 70 GB of bfloat16 weights to a 19.4 GB checkpoint, making it practical to keep the model resident on the Mac.

The Metal 4 tensor operation used for matrix multiplication consumes bfloat16 operands rather than the packed 4-bit representation. Before multiplication, the GPU must reconstruct the weights in bfloat16. The optimized grouped GEMM in Lily performs this conversion one small weight tile at a time and holds the result in on-chip threadgroup memory only long enough to multiply it by the routed activation rows. Accumulation uses 32-bit floating point, and the output is written in bfloat16. The complete expanded weight array is never created in unified memory.

In the ablation, dequantization runs as a separate operation: it expands the 4-bit weights into a bfloat16 array in unified memory, after which the matrix kernel reads that array back. At a 512-token prompt, moving dequantization into the grouped GEMM increased end-to-end prefill throughput by 77.4% by eliminating this intermediate write and read.

Keep expert routing on the GPU

The grouped GEMM needs the activation rows assigned to each expert to be stored together. After selecting eight experts per token, a histogram counts how many assignments went to each expert. A prefix scan turns those counts into starting offsets, a scatter step places rows into their expert groups, and a block map lists the fixed-size matrix blocks that the grouped GEMM must process.

The optimized path keeps this entire sequence in one command buffer, an ordered batch of GPU operations, for each prompt chunk. An ablation instead pauses so the CPU can inspect the routing intermediates and submit the next operation. Keeping the histogram and prefix scan on the GPU adds two kernels but removes CPU–GPU synchronization inside each MoE layer.

At a 512-token prompt, enabling GPU-resident routing increased end-to-end prefill by 89%. This also shows why kernel count alone can be misleading: the faster route launches more kernels but never waits for the CPU inside the layer.

Match tile size to expert load

At a 2K-token prompt, routing every token to eight of 256 experts produces 16,384 token–expert assignments, or an average of 64 activation rows per expert. The actual distribution is uneven: some experts receive many rows, while others receive few.

The grouped GEMM divides each expert’s output into tiles, which are small rectangular blocks of a matrix multiplication’s output. Each tile is assigned to one GPU threadgroup. On Apple silicon’s GPUs, a threadgroup contains one or more simdgroups, each consisting of 32 threads that execute instructions in lockstep.

Larger tiles spread setup cost across more rows and expose more parallel work, but part of a large tile remains idle when an expert receives only a few rows. Tile size and simdgroup count are therefore coupled.

An ablation fixes the tile at 16 rows. Against that control, enabling the 32-row tile with four simdgroups improved end-to-end prefill by 13.2% at 2K tokens.

Keep recurrent state on chip

During prefill, each Gated DeltaNet layer scans the prompt in order while carrying its recurrent state forward. With register residency disabled, the ablation uses a blockwise scan. At a 2K-token prompt, that path moves 256 MiB (mebibytes) of state per layer and repeatedly stops cooperating threads at barriers, synchronization points where all participating threads must wait for one another.

The recurrent state is a matrix. The optimized kernel assigns each column to one simdgroup. The simdgroup divides the column among its threads, loads the column into their registers once, and carries the state through the entire scan. The threads exchange intermediate results through simdgroup operations rather than threadgroup memory, on-chip storage shared across a threadgroup. The completed state is written back only after the scan.

The state and its gate use a 32-bit floating-point format because small rounding errors compound across sequential updates. Query and key activations remain in bfloat16.

At a 2K-token prompt, enabling the register-resident scan improved end-to-end prefill by 5.6%. Expert GEMMs accounted for about 90% of prefill time. The sequential scan does not expose enough reusable matrix work to benefit from the Neural Accelerators.

Bound temporary memory with prompt chunking

The runtime processes a long prompt as a sequence of bounded chunks rather than keeping temporary data for every prompt token in memory at once. Model weights remain resident in unified memory, while the recurrent state and KV cache carry context from one chunk to the next. No earlier context is discarded.

Without chunking, temporary activation arrays grow with the full prompt and compete with model weights, recurrent state, and the KV cache for unified memory. Chunking keeps only one segment’s temporary values live at a time, then releases or reuses that storage before processing the next segment. This caps peak working memory and allows the engine to process longer prompts without changing the model’s output.

Chunked prefill is popular in many engines and is crucial for serving long multi-turn trajectories in these memory-constrained environments. Total prefill time for attention layers remains quadratic on prompt length, with some added overhead from repeated KV loads of earlier chunks.

05

Decode: move as few bytes per token as possible

Token handoff stays off the CPU, independent kernels overlap, intermediate writes are fused away, GQA packing cuts 8 cache reads to 2, long contexts get a fixed block layout.

Decode: minimize the bytes moved per token

Batch-1 decode processes one new row at a time. With little weight reuse, its throughput depends mainly on how many bytes the engine moves for each token. The decode changes fall into four groups: optimize the one-row weight path, keep each step on the GPU, reduce intermediate and state traffic, and read the attention cache efficiently.

FIG. 04Data flow for one batch-1 decode step and two mechanisms that reduce idle time and cache traffic. (A) The GPU streams Q4 weights and model state through attention, Gated DeltaNet, routing, and fused expert kernels, then writes the selected token directly into the next step’s input slot while sending a copy to the CPU. (B) Dependency-aware scheduling allows independent kernels to overlap. (C) GQA packing lets four query heads share each KV-row load, reducing eight independent requests to two shared loads.

Optimize the one-row weight path

MLX already dispatches one-row work to specialized matrix–vector kernels. Because Lily does not use MLX, the custom runtime must provide the same basic strategy. Our row-parallel GEMV is designed for one activation row. A simdgroup cooperates on the output while reading different parts of the weight matrix in parallel.

Keep each decode step on the GPU

Keep the token handoff on the GPU

Each decode step ends by selecting the next token; the following step begins with that token as input. Sending the selection to the CPU and then back to the GPU adds a synchronization point to every token. Our runtime instead alternates between two command buffers and two GPU-resident token slots. The GPU selects the highest-scoring token and writes its token ID directly into the input slot for the next decode step, while the CPU prepares subsequent work.

Overlap independent GPU work

In one recorded batch-1 decode step, generating a token launched 795 GPU kernels. Their dependencies formed 555 sequential stages, leaving some kernels free to run concurrently. Metal’s serial execution mode nevertheless ran every kernel in order.

The optimized decode path records the actual data dependencies in a concurrent Metal pass. Independent kernel launches can run at the same time when GPU resources allow. A barrier is inserted only when later work requires an earlier result.

Reduce intermediate and state traffic

Separate kernels often materialize an intermediate: one kernel writes a temporary result to memory, and the next reads the result back. The optimized decode path fuses four chains: the two expert input projections with their gated activation; the expert output projection with its routing score and the shared-expert result; query and key preparation before attention; and the recurrent update with its normalization. Each fused kernel keeps temporary values in registers instead of sending them through memory.

Fusion also shortens the dependency graph: when an intermediate write disappears, so does the barrier that protected its consumer.

Read the attention cache efficiently

Coalesce attention-cache reads

Attention reads keys and values from the KV cache during every decode step. In the ablation, neighboring GPU threads do not always request neighboring bytes, forcing the memory system to serve more separate transactions. Enabling coalesced loads makes adjacent threads request adjacent bytes so the hardware can combine their reads.

On the bfloat16 configuration, coalescing increased key bandwidth from 33.8 to 47.9 GB/s, increased value bandwidth from 42.0 to 61.8 GB/s, and improved end-to-end decode by 2.1% at a 3,840-token context.

Pack query heads to reuse KV rows

Grouped-query attention lets eight query heads share one KV head. In the ablation, each query head runs in a separate simdgroup, so all eight independently request the same cached KV row. The optimized kernel packs four query heads into one threadgroup, which loads each KV row once and reuses it across four attention calculations. A second threadgroup handles the remaining four heads.

This technique, commonly called GQA packing, performs the same arithmetic and produces identical output bytes while reducing eight independent KV requests to two shared loads. Against the unpacked ablation, it improved end-to-end decode throughput by 23.8% at a 32K-token context.

Switch attention layouts at long contexts

Every decode step in a full-attention layer scans the existing KV cache. A fixed-block layout divides that cache into equal pieces that the GPU can process in parallel. Its extra scheduling is not worthwhile when the cache is small, but the fixed-block layout balances the work more evenly as the context grows.

For this model, the runtime keeps the general attention path below 32K tokens and uses the fixed-block path at 32K or longer. The switch applies when each head has 256 values and eight query heads share a KV head; other shapes remain on the general path. An ablation disables this switch and always uses the general path. Enabling the fixed-block route improved end-to-end decode by 7.7% at 32K, 27.4% at 64K, and 40.2% at 128K.

06

The main operations are already at the hardware limit

Speculative decoding and several other changes didn't speed things up end to end. The conclusion: an engine has to fit both the model's structure and the hardware's compute and memory paths.

Limits of further optimization

Some changes improved an isolated operation but did not improve end-to-end inference.

Speculative decoding, which uses a smaller model to propose tokens for the full model to verify, made batch-1 decode 18% slower. Verification processed groups of two to five rows, an inefficient shape for this hardware, and the rows often selected different experts, increasing the amount of expert-weight data read. Reducing the drafter’s output vocabulary improved drafter throughput by 4.7–5.1%, but did not make the complete speculative loop faster. This result is workload-specific: our batched Qwen deployment on Blackwell uses speculative decoding under different conditions.

Other experiments included reducing GPU launches, overlapping entire phases, using larger prefill tiles, applying broader fusion, accelerating the router, and combining the output projection with token selection. None improved the complete inference loop.

Measurements of the hardware limits also showed little remaining headroom in the main prefill and decode operations. The MoE GEMMs and GEMVs reached 97.9% and 90.3% of the fastest sustained weight-read rates for their access patterns. Removing arithmetic from the sparse GEMV changed throughput by only 0.2%, confirming that weight reads rather than computation were the limiting resource. Prefill’s matrix multiplication similarly reached 93% of the theoretical matrix limit in isolation and 80–86% inside the tested models.

End-to-end performance

The end-to-end comparison loads identical 4-bit checkpoint bytes in both engines and runs one request at a time on one 40-core, 128 GB M5 Max. Within each round, the two engines run in alternating order to reduce bias from background load and changes in chip temperature. We compare against MLX-LM’s fastest direct-generation path, not its server, so the measurement focuses on model execution rather than serving overhead.

The sweep covers ten prompt lengths for prefill and ten context lengths for decode, from 256 to 128K tokens. Prefill throughput first rises as the engine spreads fixed setup costs across more tokens. Prefill peaks around a 4K-token prompt, then falls because the ten full-attention layers perform more work as the prompt grows. Decode remains nearly flat at short contexts and declines once reading the growing KV cache becomes significant. The custom engine is faster at every recorded length.

Because specialized execution can change the order of floating-point operations, we also checked numerical consistency against MLX-LM. In a teacher-forced comparison, both engines predicted the next token from the same reference prefix at each of 192 positions, preventing earlier differences from affecting later inputs. Lily’s perplexity was only 0.04% higher, and it selected the same top-ranked token at 96.35% of the tested positions.

FIG. 05Prefill throughput by prompt length and decode throughput by context length for Qwen3.6-35B-A3B Q4, batch 1, on one 40-core, 128 GB M5 Max. Across ten lengths from 256 to 128K tokens, Lily is faster at every recorded point: 1.12–1.42× MLX-LM’s prefill throughput and 1.31–1.37× its decode throughput. The comparison uses MLX-LM’s fastest direct-generation path. Both horizontal axes use logarithmic scales; neither vertical axis starts at zero.

Built for the local platform

Apple silicon is not a smaller datacenter GPU. It is a complete local inference platform with its own hardware and software characteristics. Unified memory gives a single node a very high ceiling on how much model and state it can hold. The M5 Neural Accelerators absorb the dense matrix work in prefill. The vector ALUs handle the bandwidth-bound, low-reuse remainder in decode.

Qwen adds further opportunities for specialization: keep expert routing and recurrent state on the GPU, eliminate unnecessary intermediates, overlap independent work, reuse shared KV data, and adapt kernels to the workload shape.

With model- and platform-specific optimization, one Mac can run a large sparse model efficiently. Future work will widen coverage across models, chips, and serving workloads, and turn the mechanisms validated here on one configuration into more general runtime policy.

The broader principle is to match the engine to both the model’s architecture and the hardware’s specific compute and memory paths. As frontier open-weight models and hardware evolve, high-performance local inference will increasingly depend on engines tailored to both rather than ones that abstract away their differences.

Where Indigo landsFurther

Indigo's conclusion

The speedup holds for one model on one chip and doesn't travel. The real signal is that local decode has hit the memory-bandwidth wall: with software squeezed dry, further gains come from hardware, not smarter code.

What to remember

  1. Specializing buys roughly 20% to 35% more speed, at the cost of separate engineering for every model and every chip.
  2. “Specialized beats general” keeps moving down the stack: model weights, then context, then the inference engine. The general layer becomes a commodity.
  3. Apple Silicon is not a shrunken data-center GPU; it is a complete platform for local inference.

Back on the long-running theses

confirms

Constraints are moving from algorithms to physics Measured evidence for this view on the device: once software is squeezed dry, speed comes from memory bandwidth, not algorithms.

adds to

The safety politics of open weights: gatekeeping or competition A 35B open model runs fast and accurately on a laptop: open capability has spread too far to pull back.

confirms

Sarah Guo (Conviction) Frontier models in the cloud plus private processing on the device is “more access means more use” turned into a product.

adds to

Lin Qiao (Fireworks): post-training is how you keep your taste The same “specialized beats general” pattern at another layer, this time the inference engine.

What it means for Rewired Index

On-device and hybrid inference is its own theme, pulled by privacy, cost and latency. Perplexity is private; Apple is the platform underneath; the other end is the memory-bandwidth supply chain.

What would change my mind

someone lifts single-user decode on the same chip by another order of magnitude with software alone, or this specialization is shown to carry straight over to other models.

Finished. Indigo's take on this piece is in two places: