Heterogeneous Inference(异构推理)
使用不同类型的加速器分别处理推理流水线的不同阶段,以同时优化吞吐和延迟。
动机
推理不是单一负载——prefill 和 decode 对硬件要求完全不同:
- Prefill:compute 密集,适合大 batch GPU 吞吐优化
- Decode:memory-bandwidth 密集,小 batch,延迟敏感
即使同为 decode 阶段,不同操作的特性也不同:
- Attention:stateful(动态 KV cache),memory-bound,GPU 利用率不随 batch 提升
- FFN / MoE Expert:stateless,compute-bound(dense)或 sparse(MoE),利用率随 batch 提升
单一架构无法同时最优化所有操作。
Vera Rubin + LPX 异构方案
两种使用模式
1. Attention FFN Disaggregation(AFD)
- GPU 负责:Attention(decode 阶段,stateful,需要大量 HBM 存 KV cache)
- LPU 负责:FFN / MoE expert(stateless,确定性架构适配静态工作负载)
- Ping-pong pipeline 掩盖 GPU↔LPU 通信延迟
- 源自 Megascale Infer 2504.02263 和 Step-3
2. Speculative Decoding
- LPU 运行:Draft model 或 MTP layer(利用低延迟)
- GPU 验证:Main model warm prefill k 个 draft tokens
- 通常 1.5-2× output tokens per decode step
- Draft model 需要 KV cache → 使用 FPGA 附加 DDR5(256 GB/FPGA)
- GPU 侧负载感知调度见 DSpark Speculative Decoding(confidence + batch 容量动态截断 verify)
AFD 为什么 work
MoE 稀疏性 → 每个 expert effective batch 小 → 解耦后 GPU HBM 全给 KV cache → 更多 token → expert batch 增大 → FFN 回到 compute-intensive
Agentic AI 的推动
- Agent 推理循环中延迟跨步骤累积
- 稳定的 per-token 性能和强 tail-latency 表现至关重要
- 需要 ~1000+ tokens/sec/user
实证基础
Understanding Inference Scaling For Llms 系统量化了 prefill/decode 的正交资源需求,为异构推理提供了实证基础:
- Prefill: compute-bound, 适合高 TFLOP GPU
- Decode: bandwidth-bound, 适合 memory-centric 架构(如 LPU SRAM)
- 见 Prefill Decode Divergence 的详细分析
相关页面
- Nvidia Groq 3 Lpx — LPX 实体
- Nvidia Vera Rubin Nvl72 — GPU 侧
- Disaggregated Inference — 解耦推理概念
- Lpu Architecture — LPU 架构
- Prefill Decode Divergence — Prefill vs Decode 资源分歧
- FlashDecoding++ — 单 GPU decode kernel 优化(与异构分 tier 互补)
- Understanding Inference Scaling For Llms — 推理 scaling 系统分析
- Heterogeneous Computing for Agents — Agent 推理的 OI/CF 异构框架
- Cache-Resident LLC Inference — GB 级 LLC 上 cache-resident CPU 推理
- CXL Tiered Memory — CXL 分层内存与页迁移
- CHIPSMORE — 异构在 RRAM-ACIM(静态)vs SRAM-DCIM(LoRA/动态),不是 GPU+LPU
- LEAP — 异构在 IMC / NMC / INC;LEAP-D 再按 PD 重配宏
- DynaNDE — NPU vs NDP 专家调度异构;vs MoNDE 2.6×/2.2×
- HYDRA — 封装内 PD×算子 chiplet 异构 DSE
- AI Infra Book Ch.9 — A100/H20 与 CPU/GPU 专家交点
书 Ch.9:阶段配硬件,复用定放置(2026-09)
Ch.9 给了两组可算边界,补 GPU+LPU 叙事:
- P/D 异构卡:H20 算力约 A100 一半、HBM 约两倍 → 做 decode;A100 做 prefill。4+4 分离 4.55 req/s vs 共置 3.02(1.51×)。对调角色腰斩。同构上分离不提高吞吐。
- 专家 CPU vs GPU:低复用(每专家 1 行)CPU 就地读 DRAM 快于搬 36 MiB 权重;AVX 交点 ~72 行、AMX ~689 行。跨 NUMA 125 GB/s 时 AMX 路径会改读 bound。Engram 大表在 H100 上预取窗口 69 μs,放主机内存即可(Ch.6)。
- PDD — 跨 DC H100/H200 角色映射,BCR 最高 +37.5%
FPGA–GPU LRM 投机推理(2026-09-28)
除 GPU+LPU 的 prefill/decode 分工外,HeteroReason 展示 draft@FPGA + PRM/target@GPU 的另一条异构轴:相对同构 GPU 延迟最高约 1.42×、能效最高约 1.57×(U280/V80 × RTX 3090)。
消费卡 EP 与异构训练规划(2026-10-02)
ThunderEP 证明在无 NVLink/P2P 的 RTX 40/50 PCIe 系统上,专用 EP 通信仍可相对 NCCL/vLLM 拉开(dispatch 2.00×,prefill 最高 1.66×)。HAPMoE 把异构轴推到 MoE 训练自动并行(e2e 最高 3.2×)。
比特级 CPU–GPU MoE 卸载(2026-10-05)
RapidMoE(EuroSys’27)把专家拆为低比特 W^Q(GPU)+ 残差 W^R(CPU),只让少数关键专家在 CPU 补精度:decode 相对 SOTA 卸载系统最高 3.5×、prefill 最高 2.1×;DeepSeek-V3 峰值 DRAM 240 GB vs KTransformers 385 GB。与 ThunderEP 互补(权重 vs EP 通信)。
端侧 UMA 多智能体(2026-10-06)
EdgeAgent(ASPLOS’27):Apple M4 UMA(120 GB/s)上 decode 期 CPU+GPU 共跑因抢同一总线仅 0.99×(大矩阵),故改为零拷贝列切分 TP + SME2 kernel(vs Batch-SD 1.29×)、per-agent 动态草稿预算、工具停顿 suspend-and-yield;[1,100] s 工具延迟下 makespan 1.77×。与数据中心侧 RapidMoE 的 CPU–GPU 分工相对照:端侧瓶颈是共享带宽而非 PCIe。
AMX CPU 专家执行引擎(2026-10-07)
HiNa-MoE(PACT’26)补的是 CPU–GPU MoE 卸载里 CPU 那一半的算子:不改权重布局(与 GPU 库兼容),在页交错下做 NUMA-aware 切分,decode MV→MM 用上 AMX。FFN kernel vs IPEX/KTransformers 平均 1.73×/1.68×;双路 6430+A6000 batch-1 decode vs KTransformers 平均 1.22×。与 RapidMoE 的比特级切分可叠加。
Citations
[PDD] arXiv:2609.13161
[1] nvidia-groq3-lpx-blog-2026-04.md [2] [raw/articles/GTC 2026 – The Inference Kingdom Expands.md](raw/articles/GTC 2026 – The Inference Kingdom Expands.md) [3] CHIPSMORE_CIM_Chiplets_LLM_Inference_2026.pdf [4] LEAP_IMC_NoC_LLM_Inference_2026.pdf [5] DynaNDE_Near_Data_Expert_Scheduling_2026.pdf [6] Ch.9 — 李博杰《AI Infra》 [7] AI-Infra-Book.pdf [n] arXiv:2609.28717 — HeteroReason [8] arXiv:2609.40093 — ThunderEP [9] arXiv:2609.39350 — HAPMoE [10] arXiv:2610.01265 — RapidMoE [11] arXiv:2610.03394 — EdgeAgent [12] arXiv:2610.05123 — HiNa-MoE