Paper
-
HBF - SNU/KAIST/UIUC — agent harness 定 KV 寿命,短命进 HBM、长命进 HBF;HBF 寿命 vs HBM-first 1.19–3.13×(3.3–12.2 device-years)
-
DynaCore: Shape-Adaptive Array + Disaggregated Quantization - Duke — 脉动 MEU 非对称重塑 + Split-K,prefill W8A8 / decode W4A16;TTFT vs FIGLUT/Planaria 3.50×/2.97×、TPOT 36.55×/8.02×
-
T-CCL: TMA-Based Collective Communication - Chalmers SC’26 WS — 搬运+规约交给 TMA;vs NCCL 最高 2.4×(受限预算 3.42×),SM 占用更少;vLLM 最高 1.31×
-
NCCL M2N: Layout-Aware Tensor Resharding - NVIDIA — RL 权重 M→N 重分片集体;单层 FFN-MoE 最高 7.9×;DeepSeek-V3 256 GB200 权重同步 5.78→2.77 s
-
VLA Workload Characterization for Embodied AI - KAIST ASPLOS’27 — batch-1 闭环 VLA 刻画;动作张量维度定访存/计算受限;Orin 降频省能 24–32%;重叠执行 −56 pp / 3.9× / 5.1× 无 Pareto 最优
-
HiNa-MoE: CPU AMX MoE Inference - 国防科大 PACT’26 — AMX 非侵入 MoE 算子;FFN 最高 3.37×、端到端 decode 最高 2.09×
-
LLM Inference Parallelism: Compute–Comm Trade-offs - Dell — TP/PP/HB 解析模型;prefill 选 PP(TP8 约 40% TTFT 在 NCCL)、decode 选 TP;decode 有效 NVLink ~150 GB/s
-
Terracotta: Flexible DRAM Interface & Controller - ETH SAFARI MICRO 2026 — 自定义命令 + 可编程内存控制器;>96% 定制收益,面积 0.03%、功耗 0.56%
-
Divide and Conquer: MCM GPU Disaggregation - Cantabria — MCM GPU 扩展策略×拓扑 DSE;16-chiplet Torus/256 SM 性能 2.40×、能耗 −4.45×;Ring 在 64 chiplet 跌破单片 10%
-
RailWave: EP Rail & Incast Scheduling - 中山大学等 — DeepEP 下 RailBalance + 循环置换波次;GLM-4.5-Air 回放 P50 通信 H800 2.02–5.84×、H20 1.74–4.36×
-
AFORE: AFD Expert Reconfiguration - HKUST 等 — AFD 微批级专家重配 + NVLink 迁移重叠;吞吐 +10.1–17.6%、P95 ITL −7.1–9.5% vs 最强基线
-
EdgeAgent: UMA Multi-Agent Edge Inference - 中山大学+中国移动 ASPLOS’27 — M4 UMA 零拷贝 CPU–GPU TP + 动态草稿 + 工具停顿让出;1.29×/1.77×
-
MegaFlux: Skew-Resilient MoE Megakernels - Princeton+NVIDIA — megakernel 内运行时热专家复制;8×B200 前向/反向几何均值 1.45×/1.28×(峰值 2.14×/2.64×)
-
GPU-Initiated Communication Dissected - Koç+fal — IBGDA vs CPU proxy;发起 0.7 µs/完成 4.0 µs;库额外最高 4.6 µs;~3000 连接 all-to-all 丢 59% 消息率
-
RapidMoE: Residual Offloading MoE Inference - 清华 EuroSys’27 — 比特级 CPU–GPU 卸载;decode 最高 3.5×、prefill 最高 2.1×
-
Characterizing HBF for LLM Serving - Berkeley/Furiosa — HBM–HBF–host + buffered 调度;完成时间 −36.1–87.0%;能耗最高 −55.8%;寿命 4.77→14.82 年
-
ThunderEP: PCIe Consumer-GPU MoE EP - 首尔大学 — dispatch/combine vs NCCL 2.00×/1.53×;prefill 最高 1.66×、decode 最高 1.26×
-
HAPMoE: Heterogeneity-Aware MoE Parallelism - 异构集群 MoE 自动并行;e2e 吞吐最高 3.2×;非均匀 PP 再 +78%;搜索 <1 分钟
-
Mixture-of-Kittens: MoE Megakernel for NVL72s - Stanford/Cursor — NVL72;vs 最强公开基线最高 2.37×;512 GPU 生产 e2e 1.41×
-
Janus: Agentic Serving with SSD-Centric Sparse KV - 上交大等 — TTFT 最高 1.57–3.69×(均值 1.22–1.85×);关键路径 SSD I/O <6.5%
-
Purlin: Separating Orchestration from Collectives Datapath - Stanford/NVIDIA — 延迟最高 5.14×、带宽 4.50×;SGLang 离线均值 1.13×、在线最高 2.85×
-
SPLASH: Switching Parallel Layouts of Attention - 中科院计算所 — 热切换 TP/DP/CP/DOP;吞吐 1.3–1.73×;中位切换 <0.51% step(异文于 HBF-SPLASH)
-
SpecStream: Resource-Efficient Speculative Decoding with Streamed KV - 西交大 — A800;vs 卸荷基线吞吐 Qwen3 1.41× / InternLM2.5 1.32×;同 GPU 每 GPU 吞吐均值 +55.4%
-
SPIMOE: Hybrid Sparse Reasoning MoE on Heterogeneous PIM - 北航 — ICCAD’26;vs A100 最高 8.35×;MoE FFN vs PIMoE 最高 3.33×
-
RR-Evict: Round-Robin Prefix Cache Eviction for Agentic Serving - UCSD — vs LRU P99 TTFT 最高 −75.4%;P99 uncached token 最高 −65.7%
-
EAServe: Encode-Aware Disaggregated MLLM Serving - UGA 等 — PACT’26;goodput vs Dynamo 最高 4.3×、vs vLLM 1.7×;A100 4.91 req/s
-
DynBranch: Speculative Subgraph Reuse for Agentic Serving - NUS — 32B@4×H200;vs 最强基线延迟最高 −32%;vs 无复用底 −46–66%
-
The KV Cache Is the New Memory Wall (SoK) - Singh — 五域 taxonomy;70B@128k +42 GB;H100 crossover b=32→13.4k
-
HBF-Sim: Extensible HBF Simulator for GPU Memory - 华东师大/上海创智 — 媒体吞吐最高 15.94×;page-service 放大 −41.9%;条带化 kernel 最高 3.94×
-
HeteroReason: FPGA–GPU Speculative Reasoning - Imperial/清华/Bristol — MICRO’26;延迟 1.01–1.42×、能效 1.25–1.57×;回退最高 +4.2%
-
When Fancy Eviction Fails: Prefix-Cache Replacement - Harvard — 14 算法;partial-node compute-aware vs LRU TTFT −19.9%、prefill +18.8%
-
Flux: OCS Optimal Scheduling for LLM Training - imec/Antwerp — vs RotorNet/BvN;iteration 最高 10×、峰值 NIC buffer >1000×
-
Co-Fabric: Unified xPU Interconnection - IEIT — 64-xPU 3D-Mesh vs RoCE;延迟 >50%↓、带宽 2–5×、R1 +30–80%、成本 −80%
-
HBF for Agentic LLM - 热集 HBM + 冷池 HBF;14 ms TBT、resume +≈0.1 ms、会话 24×;−7.6 kW/8-GPU
-
D Elasticity for Agentic Serving - Meta — decode 租约弹性 prefill;吞吐 geomean +16.2–17.4%,高负载最高 +43.4%
-
EMA: Elastic Memory Across GPUs - 伯克利 — 机内互借 HBM;吞吐最高 +52%,达 2× 容量静态的 96%
-
Tessera: Dynamic Block-Sparse Attention Runtime - NUS — 逻辑 mask↔物理 tile;BSA 最高 6.79×;720p 50-step 1.22–2.08×
-
SPECTRA: Speculative Decoding on Reconfigurable Tiles - Columbia — 20-tile FPGA;tile 最高 2.09×、系统级再最高 1.25×
-
SPLASH: Sparse Attention × High-Bandwidth Flash - NUS — HBM+HBF KV;100 ms TPOT 下 3.5–11.4× decode/GPU
-
Die Scaling Breaks GPU Fine-grained Scheduling - 上交/NUS/NVIDIA — 远端 HBM +67%;kernel 1.22×、LLM decode +14.3%
-
AHRR: Agents + HLS Abstraction for Chip Design - UCLA — AHRR vs Direct RTL 2.6× geomean;ICCAD’26
-
CARDAN: Scratchpad MoE Multi-Engine Dataflow - UC Merced/Yotta — Trainium3;HBM read −26–50%;batch-1 1.15–1.31×、batch-16 1.70×
-
Weave: Dynamic SM Scheduling for MoE Overlap - 上交/NUS/阿里 — 4×H100 EP4;MoE layer 2.89×、端到端 1.33×(摘要汇总)
-
COMET: FPGA Packet Tracking for EC-RDMA WAN - 东北大学/Microsoft — Agilex 7 400 Gb/s;资源 <1%;连接数 vs SDR 6×
-
DeepSeek DSec: Agentic Sandbox Infrastructure - DeepSeek — 300 万 sandbox/日;>380K 并发;>5,000 创建/s
-
MeshKV: NoC KV Cache Fabric for Tiled Decode - UCLA/Columbia — 片上 KV NoC;流量最高 −58%、KV 利用率 2.1×、多流吞吐 1.9×
-
HBFlex: Full-HBF Memory for Fine-Grained LLM KV - 北大/阿里 — 全 HBF;吞吐 vs FlashAccel 最高 1.58×、vs H3 最高 3.30×
-
Ask the Tool: Progress-Aware KV for Agentic Serving - 清华/阿里云 — 工具 progress;p90 TTFT after tool vs LRU −20.7%/−20.8%
-
Fathom: Per-Query Bit-Plane Scan for Offloaded KV - 独立 — host 卸荷;Qwen3-8B@1M vs 136-bit 扫 1.67×;同 time −18% 字节
-
PipeSwift: Pipeline Parallel Agentic Serving - 清华等 — JCT-aware PP+MTP;vs SGLang EP 1.21–1.45×、vs vLLM PP2 1.60–2.33×、vs PD-disagg 1.14–1.54×(64 H800)
-
Nested BSP - Huawei 廖恒 — Nested BSP + UB 端到端 peer;与 τ Scaling law 配对
-
Budgeted Express-Mesh - 清华 — traffic-aware express;高负载 Tornado Greedy vs Random +50.7%
-
World Model Hardware Accelerator (WMHA) - 独立研究 — DiT VLIW;MSE×23;sky130 68.4 mm²;调度 1.484×
-
Trillion-Parameter MoE in a Box: HBF Memory Provisioning - Huawei — 权重驻 HBF 后状态层 1.4–4.0 s⁻¹(vs HBM3e 33.3);HBF×6 暴露 2.30 TB/s
-
BOOST: Concurrent Host+HBM Access for LLM Inference - GT/NVIDIA/Stanford — Grace Hopper iso-batch TPOT +4.3%、高吞吐 +31%(vs prefetch +15%)
-
PDD: Cross-Datacenter Prefill-Decode Disaggregation - Infinigence/清华等 — Prefill+RLD+MD 跨 DC;H100×H200 BCR 最高 +37.5%
-
UNISON: Near-Memory Session KV Scheduler for LLM Agents - 复旦 — SPEAR+TIDE;hit +0.3–23.1%、AMAT −22–51%、TTFT −58–89%;28nm 0.169 mm²
-
Vortex: Extreme Compression for Efficient LLM Inference - Duke — 脉动 bi-flow VQ+稀疏;相对 SOTA 8.03×–23.7× 加速、5.68×–12.5× 能耗(仿真)
-
Dissecting GPU Utilization for LLM Inference on Hopper - KTH — H100 NVL vLLM/FA3;decode GMMA fill 1.6–12.5%,SOL 冷 92% vs decode 7.9%
-
RoofLang: AI-Driven Architecting of LLM Inference Systems - 行云智理/MSR — DSL+roofline 仿真;V4 系峰值 decode 3.5–39.5×;B300 agent +6.23–50.1%
-
Composable CXL Memory as K8s Shared Memory for LLM Serving - Seagate — DRA+DAX 组合 CXL;Qwen2.5-7B 跨节点 prefix TTFT 5.5–36.6×,sharing gap 1–4%
-
Entwine: Tiled Computation and Fine-Grained GPU Communication - 中科院 — tile×SM 通信预算;GEMM–RS vs NCCL geomean 1.232×(最高 1.433×)
-
Fengshui: Chiplet Ecosystem and Bespoke Accelerator Codesign - 密歇根 — 8 chiplet 池联合 BASIC;能量/EDP×$ 相对同构 −48.5–97.8%
-
SAGE: Semantic-Aware Geographic Error Recovery for AI Data Movement - 城大香港 — 语义分级×地理检查点;Garnet vs 34-hop 延迟 −28% / Ψ_del −30.1%
-
WaferTrans: IOMMU-free VA Translation for Wafer-scale GPUs - 清华 — 片上 PPD 去掉 CPU-IOMMU;vs Trans-FW 平均 2.5×(Seq/Adj 3.1×)
-
HDA-MoE: Hybrid Parallelism for MoE on 3D NMP - 北大/阿里 DAMO — 离线 hybrid 放置 + 在线调度;vs TP 1.1–3.4×、vs HD-MoE 1.1–1.3×
-
Sharing a Fabric with Collective Communication - NTU/LLNL — 存储×集体同 fabric;Lustre 同 TC all-reduce 最高 145×;DYAD vs Lustre 7.4×
-
Hot Chips 2026 Handy HBM Tutorial - Objective Analysis — HBM 吃 3× DDR 晶圆面积;DRAM 产能十年未涨;PIM/base-die 被推理推上台
-
Hot Chips 2026 Samsung HBM Base Die - Samsung — HBM4/4E B-die 改 4 nm logic;cHBM→aHBM→zHBM(WoW+HCB 取消 2.5D interposer)
-
Hot Chips 2026 SK hynix HBM Packaging - SK hynix — HBM4 12Hi 量产/16Hi Qual;HyB 才能 ≥20Hi、pitch <18 μm;i-HBM 热阻 >30% ↓
-
Hot Chips 2026 d-Matrix Raptor 3D-DRAM - d-Matrix — 1-Hi logic-on-top;自称 ≈20× BW/mm²、13.5× 更好 mW/GB/s vs HBM4;ISCA 2026 指针
-
Hot Chips 2026 OXMIQ HBF - OXMIQ — HBF 是低 α/低 β 容量点;72-GPU 机柜 ~14× 容量 / ~0.6× 带宽;HBM for the rack
-
Hot Chips 2026 NVIDIA NVLink Fusion - NVIDIA — NVL72 全铜 72 GPU;3.6 TB/s per GPU、900 GB/s C2C、28.8 TB/s/switch tray;CHI Fusion
-
Hot Chips 2026 Pistil 20-Chiplet SLM - Harvard/Google/Lockheed — 16 nm 2.5D flower;512 MB / 51.2 GB/s;vs Jetson Nano 最高 7.6× 吞吐
-
Hot Chips 2026 NVIDIA Rubin GPU - NVIDIA — NVLink 6;72 GPU / 3.6 TB/s all-to-all;100 MW factory 2 ZFLOPS / 11 PB / 800 PB/s
-
Hot Chips 2026 AMD Instinct MI455X - AMD — 12× HBM4 432 GB / 23.3 TB/s;UALoE 3.6 TB/s bi-dir
-
Hot Chips 2026 AMD Helios UALoE - AMD — 72-GPU rack;UALoE load-store / ESUN;1.8 TB/s/dir
-
Hot Chips 2026 Intel Crescent Island - Intel — Xe3p;160 GB LPDDR5x;Memory Fabric,不是 packet NoC
-
Hot Chips 2026 NVIDIA Vera CPU - NVIDIA — C2C 1.8 TB/s;SCF 3.4 TB/s;1.5 TB SOCAMM
-
Hot Chips 2026 Cerebras CS-4 - Cerebras — 三片 WSE-3T;Direct Wafer Links + RoCE;CS-6 指向 3D DRAM
-
Hot Chips 2026 Meta MTIA 400 - Meta — 2D mesh + leaky-bucket;1.2 TB/s SU
-
Hot Chips 2026 Microsoft Maia 200 - Microsoft — GNOC / Ethernet ATL;FCQ 4 卡;口号 chip→6k
-
Maia 200 SDLA 归档全文 - Microsoft arXiv:2608.24664 — ATLv2 接收端驱动;Hamming Mesh 20+8 / 4 plane → 6144;8 芯 Allgather 78%/94% SoL
-
Hot Chips 2026 Google TPU 8 - Google — 8i Boardfly / 8t OCS torus
-
Hot Chips 2026 SambaNova SN50 - SambaNova — dataflow RDU;2 TB/s Ethernet SU
-
Hot Chips 2026 NVIDIA BlueField-4 - NVIDIA — Astra 7.2 Tb/s;KV G1–G4
-
Hot Chips 2026 NVIDIA Groq 3 LPX - NVIDIA — 256 LPU / 128 GB SRAM / 350 ns C2C
-
Hot Chips 2026 NVIDIA Spectrum-X Multiplane - NVIDIA — 8 plane × 4 rail;512k @ 1.6T
-
Hot Chips 2026 OpenAI Jalapeño - OpenAI — 128 @ 600 GB/s / 2048 @ 200 GB/s Ethernet SU
-
Hot Chips 2026 Broadcom Thor Ultra - Broadcom — 800G eRoCE / MRC++ / RCCC
-
Synchronization Tax in GPU Scale-Up Domains - Cornell — 8-GPU 域集体通信 >50% 是 barrier 等待;最优带宽随域规模下降(512 vs 8 GPU 为 2.06×)
-
Thermal Tuning Overhead in Wafer-Scale Optical Interconnects for LLM MoE - Georgia Tech — 晶圆级 DWDM MRR 热光 stall;铁电相对热光 Mixtral/Qwen-MoE/LLaMA-MoE 2.7×/3.8×/3.3×(四层 proxy)
-
HYDRA: Heterogeneous Chiplet DSE for Hybrid LLM Serving - UW–Madison/Ulsan — 2.5D 异构 chiplet 上 Hybrid LLM serving 的宏架构+运行时联合 DSE;平均 1.55× 吞吐、TTFT −43.7%,最高 2.3×
-
HCCL: Collective Communication for Meta MTIA 300 - SC 2026 自称 — 包内 NIC chiplet + ME/NMC 卸载集体;机柜内最高 940 GB/s,重叠 GEMM 降幅 <0.5%;推理 PUT 集体 <6 μs
-
ReXpert: ReRAM Near-Memory FFN Pool for Disaggregated MoE - HKUST/阿里云 — 驻留 expert + core 内组播;occupancy 0.328→0.519;iso-compute vs H20 FFN 9.5×、权重搬运能 20×
-
DASH: Dual-Path HBF for MoE LLM Inference - KAIST — GPU–HBF 直连 + HBM 基座中继;五模型几何均值吞吐 1.90× vs RelayOnly;代表负载 1.94× 吞吐 / 1.90× E2E
-
FLINT: Workload-Driven HBF Substrate for LLM Inference - Huawei/ETH — burst-buffer + phantom-plane + 只读 FTL;decode 吞吐 vs SSD/HBM-only/H3 为 1205×/2.2×/6.2×(仿真)
-
CHIPSMORE: CIM Chiplets for Multi-Mode Multi-Request LLM Inference - NUS — RRAM-ACIM+SRAM-DCIM + IPCN DMAC;vs H100 Mistral-7B INT8 最高 2.38× 吞吐、27× 能效(仿真)
-
LEAP: IMC-NoC LLM Inference with Balanced Dataflow - NUS — IMC+NMC+INC;LEAP-D vs H100 1.52× 吞吐 / 24.91× 能效(仿真)
-
DynaNDE: Near-Data Expert Scheduling for Batched MoE - IIT — NPU–NDP 分析模型调度;vs MoNDE prefill/decode 2.6×/2.2×(仿真)
-
Scaling Inference Prefill with High-Radix Photonic Interconnects - 3D 光子 scale-up;高 batch 2.1–3.2×、跨 pod 生产平台 2.2–4.5×(分析)
-
AInfer-PD: In-Place Prefill–Decode Multiplexing for MoE - Ant — turnstile+DeepEP 相位隔离;vs Normal −7.1–22.5%、vs SGLang −24.8–32.9%
-
BASP: Batch-Aware Sequence Parallelism - Clemson — Ulysses A2A 按 micro-batch 子组;Llama/Qwen 相对 Ulysses 1.17–1.32×(8×A100)
-
CREDIT: DSMEM Inter-CTA Tiling - UW–Madison — DSMEM reduction-reuse + 成本模型 91.7%;5090/H100 几何均值 1.466×/1.318×
-
Einsummable: Automatic Multi-GPU Parallelism - Rice — join-agg 自动分解;LLaMA block 几何均值 8.97 ms vs PyTorch 13.65 / vLLM 14.87
-
CIERA: Cross-Iteration Exponent Reuse Allgather - UVA/Anyscale — MoE 无损指数复用 Allgather;OLMoE@16GPU vs ZeRO-3 3.70×、vs ZeRO++ 3.68×
-
REACT: Tuning Collective Patterns in Shared AI Clusters - UIUC/Meta/IBM — NCCL shim 拥塞改写集体 pattern;算法带宽 +13–38%,ns-3 最高约 +75%
-
FlexPosit: Tunable Fractional Precision LLM Accelerator - UVA/SJTU — Posit+bit-serial;vs BitMoD 最高 1.8× 吞吐/1.2× 能,vs OliVe 1.5×/2.0×(16 nm)
-
LogicFolding - Huawei — W2W HB LogicFolding;Kirin2026 密度 +55%,iso-perf NPU/GPU/CPU −66%/−58%/−41%
-
DICE: Detailed Inter-Chiplet End-to-End PHY Modeling - Uppsala — gem5 运行时 QC-LDPC/PAM4 chiplet PHY;相对 HeteroGarnet IPC 平均偏移 6.8%、最高 27.6%;9454P 跨 die 最大 C2C RMSE 89.5 vs 141.2 cycle
-
C2C-Explorer: Chip-to-Chip Interconnect DSE for LLM Systems - DAC 2026 — LLM 轨迹驱动的 scale-up C2C 仿真+贝叶斯 DSE;FPGA 时序误差 2.46–8.23%;DeepSeek combine goodput +44.1%、buffer −98.4%
-
Fovea: Physical-Implication-Aware Wafer-Scale DSE - 清华 — 物理可行域 + Decision Domain;70 对 LLM 训练全部找回参考最优,相对穷尽参考平均 4.13×、最高 7.80×
-
ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM MoE - WSU ESWEEK-26 — FeFET-NAND PNM 存 expert、DRAM-PNM 做 attention、分层树 NoC;相对 H3D-T 最高 15.7× TBT、9.8× 能效
-
3DLS: A 3D Logic-Stacked Architecture for Disaggregated LLM Serving - KAIST IEEE CAL 2026 — logic-on-logic 把 KVT 与 decode AllReduce 物理隔离;相对共享 D2D 最高 1.49× 吞吐、60.2% 更低 E2E
-
Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet Architectures - UNC/UMN 3.5D 晶圆级 NoP-Tree + 专家共激活布局;Qwen3/OLMoE/DeepSeek post-training 1.92× / 2.37× / 2.17×
-
Network Design for Wafer-Scale Systems with Wafer-on-Wafer Hybrid Bonding - ETH Iff et al. — WoW 放置即拓扑;相对 mesh-like baseline 吞吐最高 +250%、延迟 -36%、每字节能量 -38%
-
22580: From GPT-2 to Kimi K3, Explained - Ali (@waterloo_intern, Baseten) 2026-07-27 X 长文;从 GPT-2 attention 一路演化到 Kimi K3;核心论点”过去七年 LLM 真正的变化不是规模 22,580×,而是 attention 状态空间从 O(N) 到 O(1) 的选择/衰减/reset 范式”
-
MegaScale-Infer - MegaScale-Infer:MoE disaggregated attention/FFN serving,ping-pong pipeline + M2N 通信库,1.90× 吞吐提升
-
Resilient AI Supercomputer Networking using MRC and SRv6 - MRC+SRv6+multi-plane Clos:三管齐下的 100K+ GPU AI 训练网络容错方案,OpenAI/Microsoft 生产验证
Summary
- A Cloud-Scale Characterization of Remote Procedure Calls - SOSP 2023 Google — 700 天 fleet 级 RPC 剖析:>10K 方法、毫秒级延迟、RPC/CPU 比年增 30%;尾延迟由 RPC tax 主导
- A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators - FlooNoC 扩展:multicast/归约/barrier + DCA 借 Snitch FPU;router +16.9% 面积,4×4 mesh 上 multicast 5.3×、reduction 2.8×,SUMMA GEMM 最高 3.8×
- A Preliminary Architecture for a Basic Data-Flow Processor - Dennis & Misunas (ISCA 1975) 基本数据流处理器:decider/T-gate/merge 条件迭代、Decision Units、Instruction Cell 两级存储作活跃指令 cache
- AI Accelerators for Large Language Model Inference: Architecture Analysis and Scaling Strategies - 首篇 LLM 推理加速器跨架构定量横评:五类(GPU/Systolic/SRAM-centric/WSE/Deterministic pipeline)六大操作域评估;expert parallelism 8.4× 参数-计算比但 2.1× 延迟方差
- Alibaba HPN: A Data Center Network for Large Language Model Training - SIGCOMM 2024 阿里云 — LLM 训练专用 2-tier dual-plane DCN,15K GPU/Pod;+14.9% 训练吞吐,non-stacked dual-ToR 防单点
- Aurelia: CXL Fabric with Tentacle - WORDS 2023 — 将寻址/路由/传输层 networking 化扩展 CXL fabric;解决 PBR 单路径与 PCIe 拥塞(RDMA 延迟可 spike 3×)
- Batude Monolithic 3D Review 2011 - Batude et al. ICCAD 2011 Low-Temperature 3D Sequential Integration;Monolithic vs TSV-based 路线对比 + port 假设与商业现实剖析
- Balfour Tiled CMP NoC Tradeoffs - Balfour & Dally MICRO 2006 — CMP NoC area/energy/delay Pareto;wormhole、2-stage、mesh sweet spot
- Cache-Resident LLM Inference in GB-Scale LLCs - KAUST cache-resident CPU inference — weight/attention domain split + sub-operator sync; 2.04–11.51× TPOT vs llama.cpp on Llama-3.2-3B/2-7B
- CODE PLAN: Scaling Code-Form Planning for LLM Reasoning - Wen et al. — code-form pseudocode plans auto-mined at scale; 2M-example training; 25.1% relative gain on 13 multi-step reasoning benchmarks
- Constable: Safely Eliminating Load Instruction Execution - ISCA 2024 Best Paper:likely-stable load 识别 + RMT/AMT 监控;12.4 KB/core;+5.1% perf、-3.4% 动态功耗、SMT +8.8%;与 EVES LVP 正交至 8.5%
- CosMoS: Architectural Support for Cost-Effective Data Movement in a Disaggregated Memory Systems - ACM JETCAS 2025 — 解耦内存系统硬件热页预测/调度迁移,+20% vs SOTA、+86% vs 基线;保护关键路径 cache miss
- Dally Virtual-Channel Flow Control - Dally IEEE TPDS 1992 — VC 原典;物理通道时分复用破死锁环;吞吐 vs VC 数
- DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation - DeepSeek 半自回归 speculative decoding + 负载感知 confidence verify;离线 τ +16–31%,V4 生产 per-user +57–85% vs MTP-1,开源 DeepSpec
- DynaX: Dynamic X:M Sparse Attention Acceleration - ASPLOS ‘25 DynaX — dynamic X:M structured attention pruning + block scheduling; 89–92% sparsity at <1% accuracy loss; 1.99× speedup vs Sanger on BERT
- Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks - MIT 65nm 168-PE CNN 加速器:Row Stationary dataflow、四级存储、GIN 组播 NoC、RLC+data gating;AlexNet 35 frames/s @278mW、0.0029 DRAM access/MAC
- FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow Switching - NEST+BIRRD 可重构加速器,RIR 在归约中做 arbitrary layout reorder;Layoutloop dataflow-layout 联合搜索,ResNet-50 1.27–2.89× 延迟、FPGA 2.65–3.91× 吞吐
- Feero & Stan 3D Mesh NoC - Feero et al. Microelectronics J. 2008 — 3-D Mesh NoC 拓扑基线:直径短 1/3、port 7、面积 +40%、TSV pitch vs KOZ trade-off
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning - IO-aware 精确 attention 第二代:减 non-matmul、seq 维并行、warp split-Q;相对 FlashAttention ~2×,A100 73% 峰值 TFLOPs/s、GPT 训练 225 TFLOPs/s
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision - Hopper H100:TMA/WGMMA producer-consumer、2-stage GEMM-softmax 流水线、FP8 block quant+incoherent processing;FP16 740 TFLOPs/s、相对 FA2 1.5–2×
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness - IO-aware tiling + online softmax + 反向重算;O(N) 内存、IO-optimal HBM 访问;GPT-2 attention 7.6×、BERT 15% 快于 MLPerf 纪录、Path-X 16K 61.4%
- FlashDecoding++: Faster Large Language Model Inference on GPUs - 异步 unified-max softmax + M=8 flat GEMM 双缓冲 + FastGEMV/CUTLASS 启发式 dataflow;decode 相对 HF 最高 4.86×、FlashDecoding 平均 1.37×,NVIDIA/AMD 双平台
- FlashMoE: Fast Distributed MoE in a Single Kernel - NeurIPS 2025:单 persistent GPU kernel 融合 MoE 计算与 NVSHMEM RDMA;actor 调度 + payload-efficient dispatch;8×H100 最高 6× 延迟、5.7× 吞吐、93% SM(FP32 vs FP16 基线)
- FlexInfer: Flexible On-Device LLM Offloading - FlexInfer — async prefetch + balanced memory locking + flexible tensor retention for budget-adaptive edge LLM inference; 10.6–12.5× vs prior offload on Llama2-70B
- HCache: Fast State Restoration in LLM Serving - EuroSys ‘25 HCache — restore conversational state from hidden activations; 6× less compute than recompute, 2× less I/O than KV offload; up to 5.73× TTFT gain
- Heterogeneous Computing for AI Agent Inference - Zhao & Liu — OI/CF framework beyond roofline for agent inference; snowballing contexts (300K–1M tokens) expose memory capacity wall and system heterogeneity need
- Hoskote 5GHz Mesh Polaris - Intel Polaris 80-core 5GHz Mesh NoC(ISSCC/JSSC 2007–08)— 工业频率、message class、fault-tolerant XOR 路由
- Hybrid Bonding 3D Integration Recent - 综述整合:Cu-Cu 直接键合 ~1 μm pitch,TSMC SoIC/Samsung X-Cube/Intel Foveros/SK hynix HBM4 已量产;3D NoC 设计假设被根本改写
- HyperMR: Efficient Hypergraph-enhanced Matrix Storage on Compute-in-Memory Architecture - SIGMOD 2025 — CIM 矩阵存储超图建模 + 两阶段划分,优化通信/累加成本;100% 矩阵有效优化,合成查询 +29.65%
- Katti TSV Technology Roadmap 2010 - Katti et al. IEEE Comm. Mag. 2010 — TSV 综述原典:via-first/middle/last 工艺、KOZ、寄生 R/C、热密度、良率模型;3D NoC 物理层标准参考
- Kim Adaptive Routing High-Radix Clos - Kim/Dally/Abts SC 2006 — high-radix Clos + DisPERoute;负载均衡自适应 vs mesh+DOR
- Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective - LLM 推理加速硬件最完整综述:HBM-assisted vs SRAM-based 双路线、Quantization/Sparse/Speculative/Paged 关键专题、典型 FPGA/SoC 数据点
- LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference - FPGA hybrid spatial-temporal dataflow:Macro Dataflow Kernels + state-machine 调度 + multi-FPGA ring;解决”spatial dataflow 在 decode 串行依赖下利用率不足”;双节点 1.67× A100、四节点 2.52× A100
- M5: Mastering Page Migration and Memory Management for CXL-based Tiered Memory Systems - ASPLOS 2025 — CXL 控制器内 HPT/HWT 硬件热页追踪平台,比 CPU 驱动 ANB/DAMON 识别更准(+47% 热度),内存密集型应用 +14% 性能
- MAERI: Enabling Flexible Dataflow Mapping over DNN Accelerators via Reconfigurable Interconnects - ASPLOS 2018 首个 flexible interconnect DNN 加速器:ART + Distribution Tree + tiny switches;任何 layout/任意 dataflow 都能映射;8-459% 利用率提升
- Mixed Precision Training - Micikevicius et al. (ICLR 2018) — FP16 training with FP32 master weights, loss scaling, and FP32 accumulation; ~2× memory savings, no hyperparameter change
- MOCAP: Wafer-Scale Chunked Pipelining for Prefill-Only LLM Inference - Tsinghua MOCAP — MBKR + LBCP chunked pipeline on wafer-scale chips for prefill-only workloads; 76.4% lower latency and 3.24× throughput vs GPipe
- Multi-Branch Self-Drafting for LLM Inference Acceleration - AAAI-25 Self-Draft — multi-branch in-model drafting without extra draft model; 2.0–3.2 tokens/step and ~2× throughput vs AR decode
- Near-Optimal Wafer-Scale Reduce - WSE Reduce/AllReduce 首次系统研究:性能模型(<4% 误差)、5 种算法(Auto-Gen ≤1.4× 下界)、3.27× 快于 vendor
- NVIDIA NVLink Hopper Blackwell - Hopper/Blackwell NVLink + NVSwitch — 固定 fat-tree、NVL72、每 GPU 带宽代际翻倍
- Optimization of Collective Reduction Operations - Rabenseifner ICCS 2004:MPI Reduce/AllReduce 五算法(tree、doubling、RHD、binary blocks、ring)与 (p,n) 选择;占 MPI 时间 >40%;长向量相对 vendor 最高 100×
- Optimizing the Parallelism of Communication and Computation in Distributed Training Platform - ICA3PP 2023 — Torus-Ring 分层训练平台上重叠通信与计算,ResNet50 +23.8–25.6%,Transformer +11.7–12.8%
- PANDA: Adaptive Prefetching and Decentralized Scheduling for Dataflow Architectures - ACM TACO 2025 — 应用自适应 prefetch + PE 去中心化调度;相对 Plasticine 1.90×、REVEL 2.53× geomean 性能
- pHost: Distributed Near-Optimal Datacenter Transport Over Commodity Network Fabric - UC Berkeley CoNEXT 2015 — 主机端 RTS/token 分布式调度,商品交换机上接近 pFabric FCT(±4%),比 Fastpass 快 3.8×
- Plasticine: A Reconfigurable Architecture For Parallel Patterns - Stanford CGRA 直接支持 Map/FlatMap/Fold/HashReduce;64 PCU+64 PMU @28nm 112.8mm²、12.3 TFLOPS;相对 Stratix V 最高 76.9× Perf/W、DHDL 数分钟编译
- PRESERVE: Prefetch Weights and KV-Cache in Distributed LLM Serving - Huawei PRESERVE — overlap HBM→L2 weight/KV prefetch with collective comm; up to 1.6× E2E speedup; optimal L2 104 MB yields 1.25× perf/$
- RDMA over Ethernet for Distributed AI Training at Meta Scale - SIGCOMM 2024 Meta — 专用 backend RoCE 网络设计/运维:ECMP→流量工程,DCQCN→collective 库接收端准入;千级–32K GPU 集群
- Route Packets, Not Wires - Dally & Towles DAC 2001 — NoC 奠基;packet-switched on-chip vs dedicated wires;五决策话语体系
- SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts - SN40L RDU(TSMC 5nm, 1040 PCU+PMU, 638 BF16 TFLOPS, 520 MiB SRAM+64 GiB HBM+1.5 TiB DDR)+ Samba-CoE(150 个 8B expert, 1T 总参);streaming dataflow 编译期融合数百 op;vs DGX H100 3.7× speedup
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN Training - HPCA 2020 MAERI 团队的 sparse + training 延伸:Flex-DPE + FAN(Forwarding Adder Network)+ global NoC;任意 GEMM 形状 + 任意稀疏度;5.7× vs systolic、3× vs 稀疏加速器
- SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on Mobile - ASPLOS 2024 反向思路:编译期把 layout transformation 消除掉(不靠灵活 NoC);4 类算子分类 + 2.5D 内存;2.8× vs DNNFusion、6.9× vs TVM、7.9× vs MNN
- SpaDA: A Spatial Dataflow Architecture Programming Language - 空间数据流语言 place/dataflow/compute + GT4Py→CSL 优化编译;WSE-2 14× 减码、collective 1.04× 手写 CSL、260 TFlop/s stencil、82× GEMV vs A100
- SuperInfer: SLO-Aware Rotary Scheduling on Superchips - UIUC SuperInfer — RotaSched + DuplexKV on GH200 NVLink-C2C; up to 74.7% higher TTFT SLO attainment under KV pressure vs PCIe offload stacks
- TileLoom: Automatic Dataflow Planning for Tile-Based Languages - MLIR 编译框架:Triton/Helion → spatiotemporal dataflow planning + df 硬件模型;Tenstorrent Wormhole/Blackhole 上 FlashAttention ~2× TTNN、Mamba Scan 最高 27× unfused
- TPU v4 Optically Reconfigurable - Jouppi et al. ISCA 2023 — TPU v4 pod OCS 可重构拓扑;4096-chip scale-up
- Understanding Inference Scaling for LLMs - Reasoning-centric LLM 推理系统瓶颈分析:Capacity Trap, Reasoning Cliff, DP→TP Transition, Prefill-Decode Divergence(8B-671B H200 实测)
- Understanding Silent Data Corruptions in a Large Production CPU Population - SOSP 2023 — 阿里云 >100 万 CPU、32 个月 SDC 实测:故障率 3.61‱;提出 Farron 优先测试 + 温控缓解
- Venus: A Versatile Deep Neural Network Accelerator Architecture Design for Multiple Applications - DAC 2023 NoC fission/fusion 多 DNN 并行 serving:分布式 buffer + flexible NoC 按 workload 动态 morph;首个 runtime multi-tenancy 适配 layout 工作
- Voxel: 3D-Stacked AI Chip Efficiency for LLM Inference - Voxel 编译器感知 3D AI 芯片仿真框架:LLM prefill/decode 软硬件协同探索,Graphcore IPU 验证误差 ≤6.8%
- WaferLLM: Large Language Model Inference at Wafer Scale - 首个晶圆级 LLM 推理系统:PLMR 设备模型 + MeshGEMM/MeshGEMV + KV shift;WSE-2 上 E2E 10–20× SGLang/A100 集群、MeshGEMV 606× 单 A100
- Æthereal Network on Chip - Philips Æthereal NoC(IEEE MDT 2005)— contention-free TDM 电路交换提供 GS;GS+BES 组合;分布式/集中编程;四种路由器面积对比