Heterogeneous Inference(异构推理)

使用不同类型的加速器分别处理推理流水线的不同阶段,以同时优化吞吐和延迟。

动机

推理不是单一负载——prefill 和 decode 对硬件要求完全不同:

  • Prefill:compute 密集,适合大 batch GPU 吞吐优化
  • Decode:memory-bandwidth 密集,小 batch,延迟敏感

即使同为 decode 阶段,不同操作的特性也不同:

  • Attention:stateful(动态 KV cache),memory-bound,GPU 利用率不随 batch 提升
  • FFN / MoE Expert:stateless,compute-bound(dense)或 sparse(MoE),利用率随 batch 提升

单一架构无法同时最优化所有操作。

Vera Rubin + LPX 异构方案

两种使用模式

1. Attention FFN Disaggregation(AFD)

  • GPU 负责:Attention(decode 阶段,stateful,需要大量 HBM 存 KV cache)
  • LPU 负责:FFN / MoE expert(stateless,确定性架构适配静态工作负载)
  • Ping-pong pipeline 掩盖 GPU↔LPU 通信延迟
  • 源自 Megascale Infer 2504.02263 和 Step-3

2. Speculative Decoding

  • LPU 运行:Draft model 或 MTP layer(利用低延迟)
  • GPU 验证:Main model warm prefill k 个 draft tokens
  • 通常 1.5-2× output tokens per decode step
  • Draft model 需要 KV cache → 使用 FPGA 附加 DDR5(256 GB/FPGA)
  • GPU 侧负载感知调度见 DSpark Speculative Decoding(confidence + batch 容量动态截断 verify)

AFD 为什么 work

MoE 稀疏性 → 每个 expert effective batch 小 → 解耦后 GPU HBM 全给 KV cache → 更多 token → expert batch 增大 → FFN 回到 compute-intensive

Agentic AI 的推动

  • Agent 推理循环中延迟跨步骤累积
  • 稳定的 per-token 性能和强 tail-latency 表现至关重要
  • 需要 ~1000+ tokens/sec/user

实证基础

Understanding Inference Scaling For Llms 系统量化了 prefill/decode 的正交资源需求,为异构推理提供了实证基础:

  • Prefill: compute-bound, 适合高 TFLOP GPU
  • Decode: bandwidth-bound, 适合 memory-centric 架构(如 LPU SRAM)
  • 见 Prefill Decode Divergence 的详细分析

相关页面

书 Ch.9:阶段配硬件,复用定放置(2026-09)

Ch.9 给了两组可算边界,补 GPU+LPU 叙事:

  1. P/D 异构卡:H20 算力约 A100 一半、HBM 约两倍 → 做 decode;A100 做 prefill。4+4 分离 4.55 req/s vs 共置 3.02(1.51×)。对调角色腰斩。同构上分离不提高吞吐。
  2. 专家 CPU vs GPU:低复用(每专家 1 行)CPU 就地读 DRAM 快于搬 36 MiB 权重;AVX 交点 ~72 行、AMX ~689 行。跨 NUMA 125 GB/s 时 AMX 路径会改读 bound。Engram 大表在 H100 上预取窗口 69 μs,放主机内存即可(Ch.6)。
  • PDD — 跨 DC H100/H200 角色映射,BCR 最高 +37.5%

FPGA–GPU LRM 投机推理(2026-09-28)

除 GPU+LPU 的 prefill/decode 分工外,HeteroReason 展示 draft@FPGA + PRM/target@GPU 的另一条异构轴:相对同构 GPU 延迟最高约 1.42×、能效最高约 1.57×(U280/V80 × RTX 3090)。

消费卡 EP 与异构训练规划(2026-10-02)

ThunderEP 证明在无 NVLink/P2P 的 RTX 40/50 PCIe 系统上,专用 EP 通信仍可相对 NCCL/vLLM 拉开(dispatch 2.00×,prefill 最高 1.66×)。HAPMoE 把异构轴推到 MoE 训练自动并行(e2e 最高 3.2×)。

比特级 CPU–GPU MoE 卸载(2026-10-05)

RapidMoE(EuroSys’27)把专家拆为低比特 W^Q(GPU)+ 残差 W^R(CPU),只让少数关键专家在 CPU 补精度:decode 相对 SOTA 卸载系统最高 3.5×、prefill 最高 2.1×;DeepSeek-V3 峰值 DRAM 240 GB vs KTransformers 385 GB。与 ThunderEP 互补(权重 vs EP 通信)。

端侧 UMA 多智能体(2026-10-06)

EdgeAgent(ASPLOS’27):Apple M4 UMA(120 GB/s)上 decode 期 CPU+GPU 共跑因抢同一总线仅 0.99×(大矩阵),故改为零拷贝列切分 TP + SME2 kernel(vs Batch-SD 1.29×)、per-agent 动态草稿预算、工具停顿 suspend-and-yield;[1,100] s 工具延迟下 makespan 1.77×。与数据中心侧 RapidMoE 的 CPU–GPU 分工相对照:端侧瓶颈是共享带宽而非 PCIe。

AMX CPU 专家执行引擎(2026-10-07)

HiNa-MoE(PACT’26)补的是 CPU–GPU MoE 卸载里 CPU 那一半的算子:不改权重布局(与 GPU 库兼容),在页交错下做 NUMA-aware 切分,decode MV→MM 用上 AMX。FFN kernel vs IPEX/KTransformers 平均 1.73×/1.68×;双路 6430+A6000 batch-1 decode vs KTransformers 平均 1.22×。与 RapidMoE 的比特级切分可叠加。

Citations

[PDD] arXiv:2609.13161

[1] nvidia-groq3-lpx-blog-2026-04.md [2] [raw/articles/GTC 2026 – The Inference Kingdom Expands.md](raw/articles/GTC 2026 – The Inference Kingdom Expands.md) [3] CHIPSMORE_CIM_Chiplets_LLM_Inference_2026.pdf [4] LEAP_IMC_NoC_LLM_Inference_2026.pdf [5] DynaNDE_Near_Data_Expert_Scheduling_2026.pdf [6] Ch.9 — 李博杰《AI Infra》 [7] AI-Infra-Book.pdf [n] arXiv:2609.28717 — HeteroReason [8] arXiv:2609.40093 — ThunderEP [9] arXiv:2609.39350 — HAPMoE [10] arXiv:2610.01265 — RapidMoE [11] arXiv:2610.03394 — EdgeAgent [12] arXiv:2610.05123 — HiNa-MoE