Bundle Update Log
2026-10-08
Watch (morning)
- Watch: 2026-10-08 Asia/Shanghai AI infra 论文巡检。arXiv cs.AR/cs.DC new = Wed 10/7(cs.AR 18 新 + 交叉/替换;cs.DC 18 新 + 交叉/替换)。口径含 WSE/NoC/NoW/SoC/3D IC、LLM architecture/interconnect/accelerator/chiplet 与 agentic AI architecture / chip design。去重:HBF 已有 characterizing-hbf / hotcold / hbflex / dash 页,Lachesis 放置依据(寿命)不同;Terracotta 2610.06475 昨日已入库;其余新候选 grep 未命中。
- Ingest: Lachesis PDF + stub(arXiv:2610.08378);DynaCore(2610.07443);T-CCL(2610.07098,SC’26 Workshops);NCCL M2N(2610.07516)。量化数字均复核自本地原文 PDF。
- Creation (papers): Lachesis(HBF 寿命 vs HBM-first 1.19–3.13×,3.3–12.2 device-years);DynaCore(TTFT vs FIGLUT/Planaria 3.50×/2.97×,TPOT 36.55×/8.02×);T-CCL(vs NCCL 最高 2.4×,受限预算 3.42×;vLLM 最高 1.31×);NCCL M2N(单层最高 7.9×;DeepSeek-V3 权重同步 5.78→2.77 s)。
- Update: End-to-End Memory Data Path、LLM Distributed Training Collectives、M2N Communication、DNN Accelerator Systolic Dataflow、Prefill-Decode Divergence。
- Indexes: 手动同步
papers/index.md、raw/papers/index.md;未跑generate_indexes.py。 - Considered not ingested: Ofan (2610.07230,胖树目的地感知交换机负载均衡,CCT 膨胀降 16–39×,下轮候选);fbnic (2610.07644,Meta 自研多主机网卡运维,NSDI’27,偏运维);NeMo-DCR (2610.08430,agentic RL 增量 bit-exact refit,1T 3% 变化 150 s vs 87.5 min);TRANSIT (2610.07593,CPU DRAM 透明扩展训练显存);ECO (2610.08373,AFD 能耗配置 BO,−40.5%);Cascadia (2610.07219,975B MoE 跑在 11 台 AI PC);TraceDSE (2610.07191,MLCAD’26,agent 做异构 SoC DSE);ActTune (2610.08444,VLA 精度/DVFS);SVRF (2610.07078)、Trail (2610.08483)、gem5 计算型 DRAM (2610.08186)、DySCo (2610.08268)、多 agent 持久记忆评测 (2610.07782) 等非本轮主线。
2026-10-07
Watch (morning)
- Watch: 2026-10-07 Asia/Shanghai AI infra 论文巡检。arXiv cs.AR/cs.DC new = Tue 10/6(cs.AR 23 条含交叉/替换;cs.DC 68 条)。口径含 WSE/NoC/NoW/SoC/3D IC、LLM architecture/interconnect/accelerator/chiplet 与 agentic AI architecture / chip design。去重:2608.24637(晶圆级光互连热调谐)、2609.39131(HBF)已有页,2609.14643(BigMoMo)log 已提;新候选 grep 均未命中。
- Ingest: VLA workload characterization PDF + stub(arXiv:2610.05062,ASPLOS’27);HiNa-MoE(2610.05123,PACT’26);Parallelism compute–comm trade-offs(2610.05305);Terracotta(2610.06475,MICRO 2026 扩展版)。量化数字均复核自本地原文 PDF。
- Creation (papers): VLA Workload Characterization(Orin 降频省能 24–32%、Thor 1–4%;重叠执行精度 −56 pp/速度 3.9×/能耗 5.1×);HiNa-MoE(FFN 最高 3.37×;端到端 decode 最高 2.09×);Parallelism Trade-offs(prefill TP8 约 40% TTFT 在 NCCL;decode 有效 NVLink ~150 GB/s);Terracotta(保留 >96% 定制收益;面积 0.03%、功耗 0.56%)。
- Update: Heterogeneous Inference、Prefill-Decode Divergence、Parallelism Transition Point、DRAM and Memory System。
- Indexes: 手动同步
papers/index.md、raw/papers/index.md;未跑generate_indexes.py。 - Considered not ingested: PhaseGate (2610.04537,M4 UMA 上按 LLM 阶段准入 CPU 检索;检索吞吐 2.0×,NeurIPS workshop,与 EdgeAgent 同向);Nexus (2610.05709,云-边 agent 执行平台,软件);SparseCraft (2610.05037,LLM agent 闭环改 Gemmini RTL;周期少 2.1×,A3@MICRO workshop,用 agent 做设计而非为 agent 设计硬件);MOLT (2610.05748,serving 与微调细粒度共享显存);LearnSched (2610.06212,agent 自进化调度);RetainZ (2610.05554,ZNS SSD 检查点放置);RL-PDN (2610.06148,供电网络 RL 优化);Alkaid (2610.04808)、AID (2610.04801)、CommuteProp (2610.05105)、SyclKittens (2610.04277)、HLS 修复/VHDL 基准等非本轮主线。
2026-10-06
Watch (morning)
- Watch: 2026-10-06 Asia/Shanghai AI infra 论文巡检。arXiv cs.AR/cs.DC new = Mon 10/5(cs.AR 13 新 + 4 交叉 + 3 替换;cs.DC 16 新 + 交叉/替换;Tue 10/6 列表上海 9 点仍未放出)。口径含 WSE/NoC/NoW/SoC/3D IC、LLM architecture/interconnect/accelerator/chiplet 与 agentic AI architecture / chip design。去重 MegaFlux/GPU-initiated comm/RapidMoE/HBF/ThunderEP/HAPMoE 等,新候选 grep 均未命中。
- Ingest: Divide and Conquer MCM GPU PDF + stub(arXiv:2610.03061);RailWave(2610.03415);AFORE(2610.03203);EdgeAgent(2610.03394,ASPLOS’27)。量化数字均复核自本地原文 PDF。
- Creation (papers): MCM GPU Divide and Conquer(16-chiplet Torus/256 SM vs 同算力 SOTA MCM 性能 2.40×、能耗 −4.45×;Ring 16/64 chiplet 跌破单片 50%/10%);RailWave(P50 通信 H800 2.02–5.84×、H20 1.74–4.36×;同需求 incast 延迟 +67.1%);AFORE(吞吐 +10.1–17.6%、P95 ITL −7.1–9.5% vs 最强基线;迁移 0 暴露);EdgeAgent(UMA 执行 1.29×、极端工具停顿 makespan 1.77×)。
- Update: Flattened Butterfly、Interconnection Network Design Space、LLM Distributed Training Collectives、Disaggregated Inference、Heterogeneous Inference。
- Indexes: 手动同步
papers/index.md、raw/papers/index.md;未跑generate_indexes.py。 - Considered not ingested: Hardware-Native Joint Sparse-Quantization for Trillion-Scale MoE (2610.02241,SpTC 稀疏量化 grouped GEMM;B200 kernel 最高 1.65×,偏压缩算法);FlashAttention on Blackwell Fixed-Shift Softmax (2610.02229,TLX 几何均值 +7.3%,kernel 增量);Backside Clock Meshes for 2 nm BSPDN (2610.02401,背面时钟网格 skew −45%,物理设计);RAPID Row-Parallel PIM in DRAM (2610.02502);Coda coding-agent serving (2610.03088,准入调度软件);VenusRL agentic RL (2610.03286,最高 4.24× 训练,软件框架);ServeTwin (2610.02732,分布式 serving 模拟器);WakeKV (2610.02713)、BCR (2610.02233)、CORE (2610.02235) KV 压缩/复用算法;CUDA→Tenstorrent Blackhole MLIR 迁移 (2610.02658);100 MW AI 集群功耗管理 (2605.24461 替换版)。
2026-10-05
Watch (morning)
- Watch: 2026-10-05 Asia/Shanghai AI infra 论文巡检。arXiv cs.AR/cs.DC new = Fri 10/2(Mon 10/5 美东列表上海早晨尚未放出;相对 10/2 早报当时仅到 Thu 10/1)。口径含 WSE/NoC/NoW/SoC/3D IC、LLM architecture/interconnect/accelerator/chiplet 与 agentic AI architecture / chip design。去重 HBF-char/ThunderEP/HAPMoE/MoK/Purlin 等;cs.AR 交叉列表 2608.24637(晶圆级光互连热调谐)已有页。
- Ingest: MegaFlux PDF + stub(arXiv:2610.00671);GPU-Initiated Communication Dissected(2610.01380);RapidMoE(2610.01265,EuroSys’27)。量化数字均复核自本地原文 PDF。
- Creation (papers): MegaFlux(8×B200 前向/反向几何均值 1.45×/1.28×,峰值 2.14×/2.64×;vLLM DeepSeek-V4-Pro prefill 中位 1.13–1.26×);GPU-Initiated Communication Dissected(最小 GPU 路径发起 0.7 µs/完成 4.0 µs;库额外最高 4.6 µs;~3000 连接 all-to-all 丢 59% NIC 消息率);RapidMoE(decode 最高 3.5×、prefill 最高 2.1×;峰值 DRAM 240 GB vs KTransformers 385 GB)。
- Update: MegaMoE Kernel、LLM Distributed Training Collectives、Heterogeneous Inference、Network Interface and System Design。
- Indexes: 手动同步
papers/index.md、raw/papers/index.md;未跑generate_indexes.py。 - Considered not ingested: Leto (2610.00687,训练 in-place 故障恢复;恢复快 3.6–6.5×,偏可靠性软件);MoE-CORE (2610.01950,NPU 内存受限专家驻留/预取,与 RapidMoE 同向但对比口径不齐);Serving a Revisable World (2610.01160,可中断 agent 的版本化执行,vLLM 控制面软件;修订后 TTFT 中位 −17.1%);ePACT (2610.01784,能耗承诺跟踪);ShatterQuant (2610.00207,16nm 块级混合精度脉动 Transformer 加速器,评测在 ViT/DiT 非 LLM);EdgeDAE (2610.00311,VLA 扩散动作 FPGA-GPU);Redundancy Meets Synergy (2610.00558,MoE 专家选择算法);CONFERM CGRA / Catscan / ZTA-Q 等非本轮主线。
2026-10-02
Watch (morning)
- Watch: 2026-10-02 Asia/Shanghai AI infra 论文巡检。arXiv cs.AR/cs.DC new = Thu 10/1(相对 10/1 早报当时仅到 Wed 9/30)。口径含 WSE/NoC/NoW/SoC/3D IC、LLM architecture/interconnect/accelerator/chiplet 与 agentic AI architecture / chip design。去重 MoK/Janus/Purlin/SPLASH-layouts/SpecStream/SPIMOE/RR-Evict 等。
- Ingest: Characterizing HBF PDF + stub(arXiv:2609.39131);ThunderEP(2609.40093);HAPMoE(2609.39350)。量化数字均复核自本地原文 PDF。
- Creation (papers): Characterizing HBF(完成时间相对 HBM-only −36.1–87.0%;能耗最高 −55.8%;寿命 4.77→14.82 年);ThunderEP(dispatch/combine vs NCCL 2.00×/1.53×;prefill 最高 1.66×、decode 最高 1.26×);HAPMoE(e2e 训练吞吐最高 3.2×;非均匀 PP 再最高 +78%;搜索 <1 分钟)。
- Update: End-to-End Memory Data Path、Memory Hierarchy and Cache、LLM Distributed Training Collectives、Heterogeneous Inference。
- Indexes: 手动同步
papers/index.md、raw/papers/index.md;未跑generate_indexes.py。 - Considered not ingested: SpecScale (2609.39334,TTS 投机服务;MATH 上吞吐 vs naive/FastTTS 2.18×/1.87×,相对 SpecStream/SPECTRA/DynBranch 增量偏搜索树服务软件);Provenance-blind KV (2609.38706,共享 KV 正确性/合约,非体系结构主线);Cascadia (2609.38697,AIPC 无控制面 serving);HPC-for-Agents (2609.38723,测量/展望);Vosti (2609.38981,确定性推理形式化);NDS (2609.38454,通用近数据 strand);Cobalt/CadenceRL/DScale/vSkipper/MEDEM/Joint MoE sim 等昨日已列理由仍适用;cs.AR STELLA CGRA / HLS pragma LLM 等非本轮主线。
2026-10-01
Watch (morning)
- Watch: 2026-10-01 Asia/Shanghai AI infra 论文巡检。arXiv cs.AR/cs.DC new = Wed 9/30(相对 9/30 早报当时仅到 Tue 9/29)。口径含 WSE/NoC/NoW/SoC/3D IC、LLM architecture/interconnect/accelerator/chiplet 与 agentic AI architecture / chip design。去重 SpecStream/SPIMOE/RR-Evict/EAServe/DynBranch/KV-SoK/HBF-Sim 等。
- Ingest: Mixture-of-Kittens PDF + stub(arXiv:2609.36070);Janus(2609.36938);Purlin(2609.36954);SPLASH-layouts(2609.37626,异文于既有 HBF-SPLASH 2609.23816)。量化数字均复核自本地原文 PDF。
- Creation (papers): Mixture-of-Kittens(vs 最强公开基线最高 2.37×;512 GPU 生产 e2e 1.41×);Janus(TTFT 最高 1.57–3.69×,均值 1.22–1.85×;关键路径 SSD I/O <6.5%);Purlin(集体延迟最高 5.14×、带宽 4.50×;SGLang 离线均值 1.13×/最高 1.37×,在线交互最高 2.85×);SPLASH-layouts(吞吐 1.3–1.73×;中位切换 <0.51% step;DOP KV 容量 +27–60%)。
- Update: NVSwitch、LLM Distributed Training Collectives、Memory Hierarchy and Cache、End-to-End Memory Data Path、Disaggregated Inference、FlashMoE Kernel。
- Indexes: 手动同步
papers/index.md、raw/papers/index.md;未跑generate_indexes.py。 - Considered not ingested: Cobalt (2609.36959,专家共激活布局;相对 MoK 同日 MoE 训练增量偏 EP 放置);CadenceRL (2609.36899,异构 RL rollout 调度);DScale (2609.37532,block-diffusion 投机,相对 SpecStream/SPECTRA/DSpark 增量偏软件);vSkipper (2609.37062,层跳过 serving 插件);MEDEM (2609.37399,多引擎加速器 DSE);MemExplorer (2604.16007 replaced,agentic NPU 异构内存综合);Scepsy (2604.15186 replaced,agentic 工作流 GPU 分配);ParaAnya (2609.36522,扩散并行采样缓存);Joint MoE topology sim (2609.37828,仿真剖析);FP64/INT8/FP4 Ozaki (2609.37693,数值仿真);昨日已列 EfficientAgent/PackServe/CascadeEP/TopoEP/VarioPath 等仍适用。
2026-09-30
Watch (morning)
- Watch: 2026-09-30 Asia/Shanghai AI infra 论文巡检。arXiv cs.AR/cs.DC new = Tue 9/29(Wed 9/30 美东列表上海早晨尚未放出)。相对 9/29 早报已扫 Fri 9/25+Mon 9/28。口径含 WSE/NoC/NoW/SoC/3D IC、LLM architecture/interconnect/accelerator/chiplet 与 agentic AI architecture / chip design。去重 EAServe/DynBranch/KV-SoK/HBF-Sim/HeteroReason/Fancy Eviction 等。
- Ingest: SpecStream PDF + stub(arXiv:2609.33184);SPIMOE(2609.34612,ICCAD’26);RR-Evict(2609.32278)。量化数字均复核自本地原文 PDF。SCHEMA 标签新增
pim。 - Creation (papers): SpecStream(vs 卸荷基线 Qwen3 1.41× / InternLM2.5 1.32×;同 GPU 每 GPU 吞吐均值 +55.4%);SPIMOE(vs A100 最高 8.35×;MoE FFN vs PIMoE 最高 3.33×);RR-Evict(vs LRU P99 TTFT 最高 −75.4%、P99 uncached 最高 −65.7%)。
- Update: DSpark Speculative Decoding、Memory Hierarchy and Cache、End-to-End Memory Data Path、Disaggregated Inference、Fancy Eviction。
- Indexes: 手动同步
papers/index.md、raw/papers/index.md;未跑generate_indexes.py。 - Considered not ingested: EfficientAgent (2609.33762,KV 卸荷何时划算的特征/策略研究,相对 SpecStream/RR-Evict 增量偏软件);PackServe (2609.33224,agentic SLO 调度);CascadeEP (2609.33252,MoE prefill 异步 EP);Tessera RAG-KV (2609.32999,与既有 Tessera BSA 同名异文,需求驱动 RAG KV);Hybrid Attention on NPUs (2609.32114);MpFA Blackwell QK4V8 (2609.33135);TopoEP (2609.35481);VarioPath PCIe AlltoAllv (2609.34340);Torch-PIM / PolyCIM / MorphAtt;AgentReplay;OLED-MoE;TempoKV;WavePP;Spexis;SmartNIC 建模;昨日已列 EAAC/PipeDRAM/CXL-SSD 等仍适用。
2026-09-29
Watch (morning)
- Watch: 2026-09-29 Asia/Shanghai AI infra 论文巡检。补扫 arXiv cs.AR/cs.DC 美东列表 Fri 9/25 + Mon 9/28(9/28 早报当时只到 Wed 9/24;Sat–Sun 无新增 AR 日栏)。口径含 WSE/NoC/NoW/SoC/3D IC、LLM architecture/interconnect/accelerator/chiplet 与 agentic AI architecture / chip design。去重 HBF-Sim/HeteroReason/Fancy Eviction/Flux/Co-Fabric/HotCold/Crossflow/EMA/Tessera/SPECTRA/SPLASH 等。
- Ingest: EAServe PDF + stub(arXiv:2609.31551,PACT’26);DynBranch(2609.31047);KV Cache Memory Wall SoK(2609.30854)。量化数字均复核自本地原文 PDF。
- Creation (papers): EAServe(goodput vs Dynamo 最高 4.3×、vs vLLM 1.7×;A100 4.91 req/s);DynBranch(vs 最强基线延迟最高 −32%;vs 无复用底 −46–66%);KV Cache Memory Wall(70B@128k +42 GB;crossover b=32→13.4k)。
- Update: Disaggregated Inference、Prefill-Decode Divergence、Memory Hierarchy and Cache、End-to-End Memory Data Path。
- Indexes: 手动同步
papers/index.md、raw/papers/index.md;未跑generate_indexes.py。 - Considered not ingested: PipeDRAM (2609.30998,MICRO’26 PUD 通用,非 LLM 主线);AI Datacenter HW Survey (2609.26829,昨日已跳);EAAC (2609.30099);CXL-SSD KV (2609.26828);PatchKV (2609.26219,后缀编辑 KV 恢复偏软件);Dynamo Fast Recovery (2609.25451);Power-Aware PD (2609.24639);Cross-Model Autoscaling (2609.29160);KernelOPT/KREX;Blackwell agent 能耗剖析 (2609.29707,测量研究);Xtrace;GRADE-RTL/AgenticSizing/Agentic-IC3;BitNet/Mamba CGLA 等既有跳过理由仍适用。
2026-09-28
Watch (morning)
- Watch: 2026-09-28 Asia/Shanghai AI infra 论文巡检。arXiv cs.AR/cs.DC 覆盖到 Wed 9/24(Fri–Sun 美东列表上海早晨尚未放出);相对 9/25 早报已扫 Flux/Co-Fabric 与 9/24 HotCold/Crossflow/EMA/Tessera。口径含 WSE/NoC/NoW/SoC/3D IC、LLM architecture/interconnect/accelerator/chiplet 与 agentic AI architecture / chip design。
- Ingest: HBF-Sim PDF + stub(arXiv:2609.29246);HeteroReason(2609.28717,MICRO’26);Fancy Eviction(2609.28870)。量化数字均复核自本地原文 PDF。
- Creation (papers): HBF-Sim(媒体 15.94×;page-service −41.9%;条带化 3.94×);HeteroReason(延迟 1.01–1.42×;能效 1.25–1.57×;回退 +4.2%);Fancy Eviction(TTFT −19.9%;prefill +18.8%)。
- Update: End-to-End Memory Data Path、Memory Hierarchy and Cache、Disaggregated Inference、DSpark Speculative Decoding、Heterogeneous Inference。
- Indexes: 手动同步
papers/index.md、raw/papers/index.md;未跑generate_indexes.py。 - Considered not ingested: AI Datacenter Hardware Survey (2609.26829,综述有用但今日优先仿真器/异构/生产轨迹三篇);EAAC compiler–HW co-design (2609.30099,通用加速器栈);Cross-Model Autoscaling (2609.29160,MaaS 调度软件);Analytical Power-Aware PD Provisioning (2609.24639,运维配比);Dynamo Fast Recovery (2609.25451,运行时恢复);KernelOPT/KREX agent 核优化;GRADE-RTL/AgenticSizing/Agentic-IC3(相对 AHRR 增量弱);BitNet/Mamba CGLA、Tetris RNB、CXL-SSD KV、SARA、MicroQonv、MVP CV SoC 等 9/25 已跳理由仍适用。
2026-09-25
Watch (morning)
- Watch: 2026-09-25 Asia/Shanghai AI infra 论文巡检。arXiv cs.AR/cs.DC/cs.NI 覆盖 Tue 9/22 及窗口内增量(相对 9/24 早报已扫 HotCold/Crossflow/EMA/Tessera);去重 SPECTRA/SPLASH/Die Scaling/AHRR/CARDAN/Weave/COMET/DSec 等。口径含 WSE/NoC/NoW/SoC/3D IC、LLM architecture/interconnect/accelerator/chiplet 与 agentic AI architecture / chip design。
- Ingest: Flux PDF + stub(arXiv:2609.25949);Co-Fabric(2609.25560)。量化数字均复核自本地原文 PDF。
- Creation (papers): Flux(iteration 最高 10×;峰值 NIC buffer >三个数量级);Co-Fabric(延迟 >50%↓;带宽 2–5×;R1 +30–80%;互连成本 −80%、功耗 −5%)。
- Update: TPU v4 OCS、LLM Distributed Training Collectives、Interconnection Network Design Space、Interconnection Network Protocol Stack、NVSwitch、AI Infra Supernode。
- Indexes: 手动同步
papers/index.md、raw/papers/index.md;未跑generate_indexes.py。 - Considered not ingested: BitNet CGLA (2609.27453,host-bound 2.52 tok/s);Mamba CGLA (2609.27437);Tetris RNB photonic scheduling (2609.25434,与 Flux 互补但今日增量弱);CXL-SSD KV (2609.26828);KV Working Set (2609.27746);SARA (2609.26763);MicroQonv;MVP CV SoC;quantum/DPD/fuzzing/tapeout education;9/24 已 ingest 的 HotCold/Crossflow/EMA/Tessera 不重复。
2026-09-24
Watch (morning)
- Watch: 2026-09-24 Asia/Shanghai AI infra 论文巡检。arXiv cs.AR/cs.DC submittedDate 覆盖 Tue 9/22–Wed 9/23(相对 9/23 早报已扫到 Mon 9/21);去重 SPECTRA/SPLASH/Die Scaling/AHRR/CARDAN/Weave/COMET/DSec/MeshKV/HBFlex 等。口径含 WSE/NoC/NoW/SoC/3D IC、LLM architecture/interconnect/accelerator/chiplet 与 agentic AI architecture / chip design。
- Ingest: Hot–Cold HBM/HBF PDF + stub(arXiv:2609.25782);Crossflow(2609.27085);EMA(2609.27040);Tessera(2609.25869)。量化数字均复核自本地原文 PDF。
- Creation (papers): HBF(14 ms TBT;resume +≈0.1 ms;会话 24×;−7.6 kW/8-GPU);Crossflow(吞吐 geomean +16.2–17.4%;高负载最高 +43.4%);EMA(吞吐最高 +52%;达 2× 容量静态的 96%);Tessera(BSA 最高 6.79×;50-step 环 1.22–2.08×)。
- Update: End-to-End Memory Data Path(修 SPLASH sources 错位 + HotCold/EMA)、Disaggregated Inference、Prefill-Decode Divergence、GPU SIMT。
- Indexes: 手动同步
papers/index.md、raw/papers/index.md;未跑generate_indexes.py。 - Considered not ingested: SARA SLO 资源分配 (2609.26763,排队论/部署优化偏软件);KV Cache Working Set 容量规划 (2609.27746,在线规划方法);GRADE-RTL / AgenticSizing / Agentic-IC3(EDA/验证 agent,相对昨 AHRR 增量弱);Toki HBM FPGA 剖析 (2609.26551,工具框架);DSAC near-threshold TPU (2609.26644,NTC 时钟,非 LLM 主线);Multi-kW 3D HI Power (2609.24904,昨已跳);nondeterminism GEMM (2609.25624,数值可复现内核);PyTorch→NPU agent 转换 (2609.27249)。
2026-09-23
Watch (morning)
- Watch: 2026-09-23 Asia/Shanghai AI infra 论文巡检。arXiv cs.AR/cs.DC submittedDate 扫到 Mon 9/21;去重 CARDAN/Weave/COMET/DSec/MeshKV/HBFlex/Ask the Tool/Fathom 等。口径含 WSE/NoC/NoW/SoC/3D IC、LLM architecture/interconnect/accelerator/chiplet 与 agentic AI architecture / chip design。
- Ingest: SPECTRA PDF + stub(arXiv:2609.24847);SPLASH(2609.23816);Die Scaling GPU Scheduling(2609.24270);AHRR(2609.21157)。量化数字均复核自本地原文 PDF。
- Creation (papers): SPECTRA(tile 最高 2.09×;系统级再最高 1.25×);SPLASH(100 ms TPOT 下 3.5–11.4×);Die Scaling(远端 HBM +67%;decode +14.3%);AHRR(vs Direct RTL 2.6× geomean)。
- Update: GEMM vs GEMV、DSpark Speculative Decoding、DNN Systolic、End-to-End Memory Data Path、GPU SIMT、DSA Tradeoffs。
- Indexes: 手动同步
papers/index.md、raw/papers/index.md、concepts/index.md;未跑generate_indexes.py。 - Considered not ingested: Multi-kW 3D HI Power Delivery (2609.24904,方法综述缺本轮量化增量);WaveletECO (2609.23444,ECO agent/本地 9B,偏 EDA 工具链);MoSim network contention sim (2609.23278,DT 仿真器);QEffect CUDA Graph FP8 PP (2609.23536,runtime 契约偏软件);Presage prefetch agent (2609.22636,预取搜索);MiX VLM (2609.19683,9/21 已跳过);NSP Nested Sequence Parallelism (2609.22755,训练 SP 软件);NPU vehicle FINN (2609.24757);AWE FP4 GEMM (2609.24519);ScaleMPA RRT* (2609.24497);CIM SAR skipping (2609.24288);Measured Joules Learned Routes (2609.23085,能耗路由软件);MCP-GRANITE (2609.24161,agent 接口评测);XDNA FlashAttention case (2609.21264,编译案例);Verification Reward Model (2609.22347);Quality over Quantity Verilog data (2609.22765)。
2026-09-22
Watch (morning)
- Watch: 2026-09-22 Asia/Shanghai AI infra 论文巡检。arXiv cs.AR/cs.DC submittedDate 扫到 Sat 9/19;去重 MeshKV/HBFlex/Ask the Tool/Fathom/PipeSwift/Nested Parallel/Express-Mesh/WMHA 等近期条目。口径含 WSE/NoC/NoW/SoC/3D IC、LLM architecture/interconnect/accelerator/chiplet 与 agentic AI architecture / chip design。
- Ingest: CARDAN PDF + stub(arXiv:2609.21137);Weave(2609.21483);COMET(2609.21774);DeepSeek DSec(2609.22978)。量化数字均复核自本地原文 PDF。
- Creation (papers): CARDAN(HBM read −26–50%;batch-1 1.15–1.31×);Weave(MoE layer 2.89×;端到端 1.33×,摘要汇总);COMET(400 Gb/s;资源 <1%;连接数 vs SDR 6×);DSec(300 万 sandbox/日;>380K 并发;>5,000/s)。
- Update: DNN Accelerator Systolic Dataflow、LLM Distributed Training Collectives、Interconnection Network Protocol Stack、DSec Sandbox Platform。
- Indexes: 手动同步
papers/index.md、raw/papers/index.md、concepts/index.md;未跑generate_indexes.py。 - Considered not ingested: ExoFlow (2609.21427,通用 Ray dataflow fault tolerance,未给 LLM/AI accelerator 增量);Towards Efficient Serverless LLM Serving (2609.22358,K8s/Knative 启发式调度,缺体系结构增量);From Deployment Hell to Stability (2609.22064,MLOps framework 经验);Efficient Wide-Area Communication for Federated Foundation Model Training (2609.22673,联邦优化/仿真,非互连硬件);The Cost-Efficiency of AI Inferencing (2609.22597,经济模型);EdgeServerlessBench (2609.22330,通用 edge benchmark);Agent Model Measurement (2609.21943,代码测量 agent 应用);DaYu (2609.22090,DPU stream analytics,非本轮 AI/LLM workload)。9/21 已列教程、VLM 格式、纯 agent/评测与软件 runtime 等跳过理由仍适用。
2026-09-21
Watch (morning)
- Watch: 2026-09-21 Asia/Shanghai AI infra 论文巡检。cs.AR/cs.DC API 到 Thu 9/17(Fri–Mon 美东列表上海早晨尚未放出)。已 ingest 的 HBFlex/Ask the Tool/Fathom/PipeSwift/Nested Parallel/Express-Mesh/WMHA 等不重复。口径含 agentic AI architecture / chip design。补扫发现 MeshKV 2609.19207(9/16 投稿)在 9/18 早报漏列。
- Ingest: MeshKV PDF →
raw/papers/MeshKV_NoC_KV_Cache_Fabric_2026.pdf+ stubraw/papers/meshkv-noc-kv-cache-fabric.md(arXiv:2609.19207, 2026-09-16, cs.AR)。 - Creation (papers): MeshKV(流量 −58%;KV 利用率 2.1×;多流 1.9×)。
- Update: Torus、End-to-End Memory Data Path、Interconnection Network Design Space。
- Indexes: 手动同步
papers/index.md(+1)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: The Life of a Token 教程 (2609.19924);SiliconBench 桌面统一内存评测 (2609.19169);MiX VLM 微缩放格式 (2609.19683);PixelFlow DiT serving (2609.20723);Do AI Agents Understand Computer Architecture? / AutoTuring 评测 (2609.19387);Rosetta 多代理建分析模型 (2609.19376);Shared KV / Hybrid-State LMCache 正确性笔记 (2609.15021/15030);DeepSeek-V4-Flash AMD gfx90a 工程笔记 (2609.15627);Token Latency Fairness (2609.18112);Xronos 边缘 CPU TP (2609.19909);ASRB K8s 路由 (2609.20497);Agentic Autoscaling 文本分类 (2609.14898);9/18 已跳过 Rect3D/Locus/HCL/GPU ISA/VeriBug/GeoMesh/SSD-LLaMA 等仍适用。BusyBarn 仍无公开全文 PDF。
2026-09-18
Watch (morning)
- Watch: 2026-09-18 Asia/Shanghai AI infra 论文巡检。cs.AR/cs.DC pastweek 到 Thu 9/17(Fri 9/18 美东列表上海早晨尚未放出)。已 ingest 的 PipeSwift/Nested Parallel/Express-Mesh/WMHA/Trillion HBF/BOOST/PDD/UNISON 等不重复。口径含 agentic AI architecture / chip design。
- Ingest: HBFlex PDF →
raw/papers/HBFlex_Flexible_Memory_HBF_LLM_2026.pdf+ stubraw/papers/hbflex-flexible-memory-hbf-llm.md(arXiv:2609.18675, 2026-09-16, cs.AR)。 - Ingest: Ask the Tool PDF →
raw/papers/Ask_Tool_Progress_Agent_KV_Serving_2026.pdf+ stubraw/papers/ask-tool-progress-agent-kv-serving.md(arXiv:2609.18849, cs.DC)。 - Ingest: Fathom PDF →
raw/papers/Fathom_Sparse_Decoding_Offloaded_KV_2026.pdf+ stubraw/papers/fathom-sparse-decoding-offloaded-kv.md(arXiv:2609.17652, cs.LG;DC 列表)。 - Creation (papers): HBFlex(vs FA 1.58× / vs H3 3.30×);Ask the Tool(p90 TTFT −20.7%/−20.8%);Fathom(1M 1.67×;−18% bytes)。
- Update: End-to-End Memory Data Path(HBFlex/Fathom)、Disaggregated Inference(Ask the Tool)、CXL Tiered Memory(Fathom host 卸荷对照)。
- Indexes: 手动同步
papers/index.md(+3)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: Rect3D 3D-IC floorplan CAD (2609.18946);Locus ZKP (2609.18846);HCL MXFP4 (2609.18792);GPU ISA encoding (2609.18662);VeriBugBench/RTL Trojans/FAME/Lyapunov/WARD/REQAP/Analog pin/PLC ladder/analog EML/FairCompressAgent/BLADE;OAK/Ermes/Fluid Notarization/Vigil/AUPE/COMPASS-ABS/DiverseFT/UAV Edge/GeoMesh geo-training (2609.18388, WAN 训练非 WSE/NoC)/Zero-I/O FT/Token Latency Fairness/SSD-LLaMA consumer MoE (2609.18110, 消费级 SSD 栈)/Thunderbolt RDMA/ASPIRE speculative/vidax video JAX mesh/GroupKV dLLM/State P2P;Wed 已跳过 OptiPrime/FSNIC/SpecLens/CGRA/ScaleLUT/FINNAS/Carry-Through/DT-RAID/BusyBarn 等仍适用。
2026-09-17
Watch (morning)
- Watch: 2026-09-17 Asia/Shanghai AI infra 论文巡检。cs.AR recent 到 Wed 9/16(15 篇;Thu 9/17 美东列表上海早晨尚未放出)。cs.DC Wed 9/16 同步覆盖。已 ingest 的 Trillion MoE HBF/BOOST/PDD/UNISON/Vortex/Hopper Util/RoofLang 等不重复。口径含 agentic AI architecture / chip design。
- Ingest: PipeSwift PDF →
raw/papers/PipeSwift_Pipeline_Parallel_Agentic_Serving_2026.pdf+ stubraw/papers/pipeswift-pipeline-parallel-agentic-serving.md(arXiv:2609.16491, 2026-09-15, cs.DC)。 - Ingest: Nested Parallel PDF →
raw/papers/Nested_Parallel_von_Neumann_Nested_BSP_2026.pdf+ stubraw/papers/nested-parallel-von-neumann-nested-bsp.md(arXiv:2609.16787, Huawei 廖恒)。 - Ingest: Budgeted Express-Mesh PDF →
raw/papers/Budgeted_Express_Mesh_NoC_2026.pdf+ stubraw/papers/budgeted-express-mesh-noc.md(arXiv:2609.17057, cs.AR/cs.NI)。 - Ingest: WMHA PDF →
raw/papers/World_Model_Hardware_Accelerator_WMHA_2026.pdf+ stubraw/papers/world-model-hardware-accelerator-wmha.md(arXiv:2609.16244, cs.AR)。 - Creation (papers): PipeSwift(JCT 1.21–1.45× / 1.60–2.33× / 1.14–1.54×);Nested Parallel(Nested BSP+UB peer);Budgeted Express-Mesh(Tornado +50.7%);WMHA(MSE×23;1.484×;68.4 mm²)。
- Update: Disaggregated Inference(PipeSwift JCT/PP)、Topology Optimization Variants + Torus(Express-Mesh)、UnifiedBus + AI Infra Supernode(Nested Parallel)、DNN Systolic(WMHA)。
- Indexes: 手动同步
papers/index.md(+4)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: OptiPrime HE-MPC 私有推理 (2609.16898, 出核心 WSE/NoC/LLM serving 口径);Cognitive Admission Control agentic (列表错位,API 为 PipeSwift);XMPIaaS / BOA ANNS / INT8 portable / SpecLens Verilog / CGRA / ScaleLUT / FINNAS / FSNIC / Carry-Through Checksum / DT-RAID / Nested BSP 同窗联邦/FL/农业 DS2;Tue 9/15 已跳过 GVA/VAMP/InplaceKV/AgentKV/FlashGPU-sim/BigMoMo/BrainScaleS/Cnuas 等仍适用。BusyBarn 仍无公开全文 PDF。
2026-09-16
Watch (morning)
- Watch: 2026-09-16 Asia/Shanghai AI infra 论文巡检。cs.AR recent 到 Tue 9/15(24 篇;Wed 9/16 美东列表上海早晨尚未放出)。已 ingest 的 Vortex/Hopper Util/RoofLang/Fengshui/Entwine/SAGE/Composable CXL 等不重复。口径含 agentic AI architecture / chip design。
- Ingest: Trillion MoE HBF PDF →
raw/papers/Trillion_Param_MoE_HBF_Memory_Provisioning_2026.pdf+ stubraw/papers/trillion-param-moe-hbf-memory-provisioning.md(arXiv:2609.15636, 2026-09-14, cs.AR)。 - Ingest: BOOST PDF →
raw/papers/BOOST_Concurrent_Host_HBM_LLM_Inference_2026.pdf+ stubraw/papers/boost-concurrent-host-hbm-llm-inference.md(arXiv:2609.13592, 2026-09-11, cs.DC/cs.AR)。 - Ingest: PDD PDF →
raw/papers/PDD_Cross_Datacenter_Prefill_Decode_Disaggregation_2026.pdf+ stubraw/papers/pdd-cross-datacenter-prefill-decode-disaggregation.md(arXiv:2609.13161, cs.AR/cs.DC;Tue 9/15 列表)。 - Ingest: UNISON PDF →
raw/papers/UNISON_Near_Memory_Scheduler_LLM_Agents_2026.pdf+ stubraw/papers/unison-near-memory-scheduler-llm-agents.md(arXiv:2609.09643, 2026-09-09, cs.AR;先前作 NMP KV 跳过,今日按 agentic 硬件口径入库)。 - Creation (papers): Trillion MoE HBF(状态层 1.4–4.0 s⁻¹;HBF×6 → 2.30 TB/s);BOOST(TPOT +4.3% / 吞吐 +31%);PDD(跨 DC BCR +37.5%);UNISON(hit +0.3–23.1%、AMAT −22–51%、TTFT −58–89%;0.169 mm²)。
- Update: Disaggregated Inference(PDD+对照)、End-to-End Memory Data Path(BOOST/HBF/UNISON)、CXL Tiered Memory(BOOST/UNISON 对照)、Heterogeneous Inference(PDD)。
- Indexes: 手动同步
papers/index.md(+4)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: Grouped Value Attention (2609.13285, 算法侧 KV 表示压缩);Dynamic HBM Repartitioning / VAMP (2609.13537, serving 运行时边界);InplaceKVCache (2609.14507, KV 抽象格式);AgentKV (2609.14872, cs.LG 软件淘汰);FlashGPU-sim (2609.15311, 仿真器);AMD Matrix Cores (2609.14845, 数值模型);BigMoMo (2609.14643, 手机 MoE);DVFS SLM (2609.13153);NPU Eval v1.0 (2609.13166);BrainScaleS chiplet NoC (2609.13563, 神经形态非 LLM);mKernel (2609.13585);ETCInfer 热调度 (2609.15230);OpWeave (2609.14237);Cnuas (2609.15889);HBF Sucks 等先验跳过仍适用。BusyBarn 仍无公开全文 PDF。AccelForge/PATTON/py-kvcache/TinyML/HeatCache 等 9/14–15 跳过项仍适用。
2026-09-15
Watch (morning)
- Watch: 2026-09-15 Asia/Shanghai AI infra 论文巡检。cs.AR/cs.DC recent 到 Mon 9/14(Tue 9/15 美东列表上海早晨尚未放出)。已 ingest 的 Fengshui/Entwine/SAGE/Composable CXL 与 WaferTrans/HDA-MoE 等不重复。
- Ingest: Vortex PDF →
raw/papers/Vortex_Extreme_Compression_LLM_Inference_2026.pdf+ stubraw/papers/vortex-extreme-compression-llm-inference.md(arXiv:2609.12208, 2026-09-10, cs.AR)。 - Ingest: Hopper Utilization PDF →
raw/papers/Dissecting_GPU_Utilization_LLM_Inference_Hopper_2026.pdf+ stubraw/papers/dissecting-gpu-utilization-llm-inference-hopper.md(arXiv:2609.12923, 2026-09-11, cs.PF/cs.AR)。 - Ingest: RoofLang PDF →
raw/papers/RoofLang_AI_Driven_LLM_Inference_Architecting_2026.pdf+ stubraw/papers/rooflang-ai-driven-llm-inference-architecting.md(arXiv:2609.12551, 2026-09-11, cs.DC)。 - Creation (papers): Vortex(脉动 bi-flow VQ+稀疏;8.03×–23.7× / 5.68×–12.5×);Hopper Utilization(GMMA fill 1.6–12.5%;SOL 92%→7.9%);RoofLang(V4 峰值 decode 3.5–39.5×;B300 agent +6.23–50.1%)。
- Update: DNN Systolic(Vortex)、GPU SIMT(Hopper util)、FlashAttention-3(FA3 serving 利用率)、GEMM vs GEMV(fragment fill + bi-flow)、Disaggregated Inference(RoofLang+M 杠杆)、End-to-End Memory Data Path(KV 压缩/流量)。
- Indexes: 手动同步
papers/index.md(+3)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: py-kvcache (2609.11744, 外部 KV/NVMe 表征;有数字但相对既有 CXL/disagg 页增量偏 serving 连接器);AccelForge (2609.11906, DSE 框架);PATTON (2609.11392, PIM runtime);TinyML multi-exit (2609.11939);PyTorch operator profiling (2609.11938);VLA robot factories (2609.12075);HeatCache (2609.12449, 热调度);ForgeMegakernel (2609.12379, codegen);Argus (2609.12299, 测量编排);CHERI/病理 BEACON/SNN/Ising/HLS agents/电网/三角格点/AMEND/UNISON/REACH/HBFSim 等先验跳过仍适用。BusyBarn 仍无公开全文 PDF。
2026-09-14
Ingest: bojieli《深入理解 AI Infra》
- Ingest: 李博杰开源书(Apache-2.0)
bojieli/ai-infra-book。只拉manuscripts/*.md(gh api/ raw.githubusercontent.com),未 clone Git LFS。 - Creation (raw): 源 stub — 许可、仓库/PDF/站点、章节清单;不收录正文。
- Creation (entities): 深入理解 AI Infra — 从约束推导设计、五问、阅读地图。
- Creation (analyses): AI Infra Book — 前言 + 12 章摘要(优先深度 Ch.6/7/9/10)。
- Creation (concepts, 2/5 cap): Constraint-Driven AI Infra Design、AI Infra Supernode。Engram / 分层集合通信 / 多 rail 溢出写在章节页与既有概念补丁中。
- Update: Disaggregated Inference(Ch.9 PD/AF)、NVLink fabric(Ch.6–7 域/超售)、LLM Collectives(环/树、分层 AR、1024 卡)、Interconnection Design Space(环面 vs 交换)、Heterogeneous Inference(A100/H20 + CPU/GPU 专家)、CXL Tiered Memory(远端读类比)、Clos(QM9700 超售/rail)、Network-on-Wafer(OpenTallas 对照)、UB(OpenURMA μs)。
- SCHEMA: 增 tag
book、methodology、supernode、datacenter。 - Indexes: 手工同步
analyses/index.md、entities/index.md、concepts/index.md、raw/articles/index.md。未跑generate_indexes.py。未部署 Pages。 - Deferred:
experiments/、calculations/代码树、archive/reviews、逐字章节转写;Ch.11 工具环境/API 成本、Ch.12 鹊桥/WAN 只做薄摘要。OpenTallas ROM 晶圆数字为书中分析、非实测硅。
Watch (morning)
- Watch: 2026-09-14 Asia/Shanghai AI infra 论文巡检。cs.AR/cs.DC recent 到 Fri 9/11(周末无新表;Mon 9/14 美东列表上海早晨尚未放出)。补扫 9/10–9/11(9/11 早报失败窗口)。已 ingest 的 WaferTrans/HDA-MoE/Sharing-a-Fabric/CIERA/REACT/FlexPosit/Huawei τ 等不重复。
- Ingest: Fengshui PDF →
raw/papers/Fengshui_Chiplet_Ecosystem_BASIC_Codesign_2026.pdf+ stubraw/papers/fengshui-chiplet-ecosystem-basic-codesign.md(arXiv:2609.10970, 2026-09-10, cs.AR)。 - Ingest: Entwine PDF →
raw/papers/Entwine_Tiled_Computation_Fine_Grained_GPU_Comm_2026.pdf+ stubraw/papers/entwine-tiled-computation-fine-grained-gpu-comm.md(arXiv:2609.11562, 2026-09-10, cs.DC)。 - Ingest: SAGE PDF →
raw/papers/SAGE_Semantic_Aware_Geographic_Error_Recovery_AI_2026.pdf+ stubraw/papers/sage-semantic-aware-geographic-error-recovery.md(arXiv:2609.10126, 2026-09-09, cs.AR)。 - Ingest: Composable CXL PDF →
raw/papers/Composable_CXL_Memory_K8s_LLM_Serving_2026.pdf+ stubraw/papers/composable-cxl-memory-k8s-llm-serving.md(arXiv:2609.10790, 2026-09-09, cs.DC)。 - Creation (papers): Fengshui(8 chiplet 池;能量/EDP×$ −48.5–97.8%);Entwine(tile×SM;vs NCCL 1.232×);SAGE(语义重放;vs 34-hop −28%/Ψ_del −30.1%);Composable CXL(K8s DRA+DAX;TTFT 5.5–36.6×)。
- Update: Interconnection Design Space(Fengshui+SAGE)、Protocol Stack(SAGE)、LLM Collectives(Entwine)、NVLink fabric(Entwine)、CXL Tiered Memory(Composable CXL)、Disaggregated Inference(Fengshui+CXL)。
- Indexes: 手动同步
papers/index.md(+4)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: AccelForge (2609.11906, 加速器建模框架/DSE);PATTON (2609.11392, 商品 PIM runtime,无新 fabric/NoC);AMEND (2609.09823, GPU-PIM attention);UNISON (2609.09643, NMP KV 调度);REACH (2609.10861, HBM ECC);HBFSim (2609.09800, HBF 仿真工具);三角格点 NoC 路由 (2609.09746, 纯理论无 LLM);电源/数据中心电网 (2609.11649);CHERI/病理 BEACON/FlexSpIM SNN/Shift-Accumulate Attention;MoE Comp-Comm Overlap (2609.07536, 已于 9/10 跳过);Tools-CC-Bench;Epoch diffusion MoE serving;EStream mobile NPU。BusyBarn 仍无公开全文 PDF。
2026-09-10
Public PDF links
- Fix: Quartz/
generate_site.py发布papers/但不发布raw/二进制,相对raw/papers/*.pdf在 https://lukebest.github.io/ 会 404。有已知 arXiv id 的论文页,把正文 PDF: 与# Citations里的本地 PDF 链接改为https://arxiv.org/pdf/<id>(frontmattersources:仍保留本地 ingest 路径)。无 arXiv 的会议论文 / Hot Chips 幻灯改为纯文本路径(非公开本地路径),不编造 arXiv id。未跑generate_indexes.py。
Watch (morning)
- Watch: 2026-09-10 Asia/Shanghai AI infra 论文巡检。cs.AR/cs.DC recent 到 Wed 9/9(Thu 9/10 美东列表上海早晨尚未放出)。昨日无增量日已扫过的 Mon 9/7 项不重复。
- Ingest: WaferTrans PDF →
raw/papers/WaferTrans_IOMMU_free_Wafer_Scale_GPU_2026.pdf+ stubraw/papers/wafertrans-iommu-free-wafer-scale-gpu.md(arXiv:2609.06125, 2026-09-05, cs.AR)。 - Ingest: HDA-MoE PDF →
raw/papers/HDA_MoE_3D_NMP_Hybrid_Parallel_2026.pdf+ stubraw/papers/hda-moe-3d-nmp-hybrid-parallel.md(arXiv:2609.08682, 2026-09-09, cs.AR)。 - Ingest: Sharing a Fabric PDF →
raw/papers/Sharing_Fabric_Collective_Storage_Penalties_2026.pdf+ stubraw/papers/sharing-fabric-collective-storage-penalties.md(arXiv:2609.06506, 2026-09-09, cs.DC)。 - Creation (papers): WaferTrans(IOMMU-free 片上翻译;vs Trans-FW 2.5×);HDA-MoE(3D NMP hybrid 放置+调度;vs TP 1.1–3.4×);Sharing a Fabric(存储×集体争用;DYAD 7.4× vs Lustre)。
- Update: Network-on-Wafer(UM 翻译面)、3D Stacking(3D NMP MoE)、LLM Collectives(fabric×存储)、Interconnection Design Space(一行)。
- Indexes: 手动同步
papers/index.md(+3)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: MoE Comp-Comm Overlap 资源管理 (2609.07536, FLUX/COMET SM residency,无新 fabric/NoC PHY);Poseidon (2609.06086, 异构集群并行搜索);EStream (2609.06551, 手机 NPU MoE prefill);Interface-Aware KV NVM (2609.05764, 片上 NVM 量化接口);Photonic chiplet HT (2609.06796, 安全威胁综述);Flash KV for recsys (2609.07175);Multichip Ising Pegasus (2609.07907);Gutenberg NDP (2609.06691);DejaVu unified-memory SoC (2609.05635);Tools-CC-Bench (2609.08739, 基准);MonoMoE/Budgeting Bytes 等 9/9 已跳过项仍适用。BusyBarn 仍无公开全文 PDF。
2026-09-09
Watch (morning)
- Watch: 2026-09-09 Asia/Shanghai AI infra 论文巡检。无增量。cs.AR/cs.DC recent/new 仍停在 Mon 9/7(Tue 9/8 与 Wed 9/9 美东列表上海早晨尚未放出)。昨日已 ingest 的 CIERA/REACT/FlexPosit/Huawei τ 与 BASP/CREDIT/Einsummable/Photonic Prefill/AInfer-PD/LEAP/DynaNDE/CHIPSMORE/Sync Tax 等不重复。
- Indexes: 无 papers 变更。未跑
generate_indexes.py。 - Considered not ingested: MonoMoE (2609.04244, H200 量化 MoE fused megakernel,无 NoC/fabric 增量);Budgeting Bytes (2609.04238, 边缘存储字节/token roofline,非互连);TreeFI/EOSQR/Proton irradiation/APEX-RBD/HCST/edge DVFS/GreenPipe/Serverless CVM/Atlas;LevelSyn EDA、systolic beamforming、Barnacle 区块链、JuPyLive、Iapetus 卫星 ViT。BusyBarn 仍无公开全文 PDF(仅 artifact/Zenodo)。先验跳过 Beacon/AXI4/survey 等仍适用。
2026-09-08
Watch (morning)
- Watch: 2026-09-08 Asia/Shanghai AI infra 论文巡检。cs.AR/cs.DC recent 列表到 Mon 9/7(Tue 9/8 美东列表上海早晨尚未放出)。已 ingest 的 BASP/CREDIT/Einsummable/Photonic Prefill/AInfer-PD/LEAP/DynaNDE/CHIPSMORE/Sync Tax 等不重复。
- Ingest: CIERA PDF →
raw/papers/CIERA_Cross_Iteration_Exponent_Reuse_Allgather_2026.pdf+ stubraw/papers/ciera-cross-iteration-exponent-reuse-allgather.md(arXiv:2609.04609, 2026-09-04, cs.DC)。 - Ingest: REACT PDF →
raw/papers/REACT_Tuning_Collective_Patterns_Shared_AI_Clusters_2026.pdf+ stubraw/papers/react-tuning-collective-patterns-shared-clusters.md(arXiv:2609.04417, 2026-09-03, cs.NI/cs.DC)。 - Ingest: FlexPosit PDF →
raw/papers/FlexPosit_Tunable_Fractional_Precision_LLM_2026.pdf+ stubraw/papers/flexposit-tunable-fractional-precision-llm.md(arXiv:2609.04724, 2026-09-04, cs.AR)。 - Ingest: Huawei τ PDF →
raw/papers/Huawei_Tau_Chip_LogicFolding_Thermal_2026.pdf+ stubraw/papers/huawei-tau-chip-logicfolding-thermal.md(arXiv:2609.04287, 2026-09-02, cs.AR)。 - Creation (papers): CIERA(无损指数复用 Allgather;OLMoE@16 3.70×/3.68×);REACT(拥塞改写集体;+13–38% / ns-3 ~75%);FlexPosit(Posit bit-serial;vs BitMoD 1.8×/1.2×);Huawei τ(LogicFolding HB;NPU −66% iso-perf)。
- Update: LLM Collectives(CIERA+REACT),3D Stacking(τ/LogicFolding),DNN Systolic(FlexPosit 一行)。SCHEMA 补
collective/allreduce/distributed。 - Indexes: 手动同步
papers/index.md(+4)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: CIM Analog Softmax (2609.04266, 电路级 softmax、无 NoC/fabric 增量);Atlas compound AI deploy (2609.04513, 编排);KV low-rank adapt (2609.04263);Prefix-cache×量化 divergence (2609.04748, serving 可复现);GreenPipe/TreeFI/Proton/APEX-RBD/EOSQR/edge DVFS/BF16 AMX/progressive compression/MemGuard/量子/FL/电网柔性;Sep 7 已跳过 Para-Pipe/FlowTT/Analog Photonic Interposer/Latency-Aware Multi-Agent。BusyBarn 仍无公开全文 PDF。先验跳过 Beacon/AXI4/survey 等仍适用。
2026-09-07
Watch (morning)
- Watch: 2026-09-07 Asia/Shanghai AI infra 论文巡检。cs.AR/cs.DC recent 列表到 Fri 9/4(周一美东新稿上海早晨尚未放出;Labor Day 窗口)。已 ingest 的 Photonic Prefill/AInfer-PD/LEAP/DynaNDE/CHIPSMORE/Sync Tax 等不重复。
- Ingest: BASP PDF →
raw/papers/BASP_Batch_Aware_Sequence_Parallelism_2026.pdf+ stubraw/papers/basp-batch-aware-sequence-parallelism.md(arXiv:2609.03151, 2026-09-04, cs.DC)。 - Ingest: CREDIT PDF →
raw/papers/CREDIT_DSMEM_Inter_CTA_Tiling_2026.pdf+ stubraw/papers/credit-dsmem-inter-cta-tiling.md(arXiv:2609.01864, 2026-09-02, cs.DC)。 - Ingest: Einsummable PDF →
raw/papers/Einsummable_Multi_GPU_Parallelism_2026.pdf+ stubraw/papers/einsummable-multi-gpu-parallelism.md(arXiv:2609.03905, 2026-09-04, cs.DC)。 - Creation (papers): BASP(Ulysses 子组 A2A;1.17–1.32×);CREDIT(DSMEM reduction-reuse;5090/H100 1.466×/1.318×);Einsummable(自动 join-agg;LLaMA block 8.97 vs 13.65/14.87 ms)。
- Update: LLM Collectives(SP/Ulysses 子组 + 自动分解),NVLink fabric(训练侧关 NVLink 域),GPU SIMT(DSMEM 一行)。
- Indexes: 手动同步
papers/index.md(+3)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: Para-Pipe (2609.04168, 边缘异构 SoC 算子流水,非 LLM fabric);FlowTT (2609.03459, DLRM TT embedding GPU kernel);Analog Photonic Interposer (2609.03125, 视觉 sensor↔模拟加速器);Latency-Aware Multi-Agent LLM on Heterogeneous GPUs (2609.03335, serving 编排);Einsummable 同窗 Barnacle/JuPyLive/区块链/FL/5G 等出范围;Sep 4 已跳过 AceSpec/Atlas/NOVA/CREDIT 当时仅扫过标题——今日全文入库;Characterizing multi-tenancy (2609.00817) 仍表征文。BusyBarn 仍无公开全文 PDF。先验跳过 Beacon/AXI4/survey 2608.28048 等仍适用。
2026-09-04
Watch (morning)
- Watch: 2026-09-04 Asia/Shanghai AI infra 论文巡检。cs.AR recent/new = Thu 9/3 提交(含 cross);已 ingest 的 LEAP/DynaNDE/CHIPSMORE/Sync Tax/FLINT/Maia/光互连/HYDRA/ReXpert/HCCL/DASH/DICE/Fovea/C2C/ThAME/3DLS/Mozart 与全部 hc2026-* 不重复。
- Ingest: Photonic Prefill PDF →
raw/papers/Scaling_Inference_Prefill_High_Radix_Photonic_2026.pdf+ stubraw/papers/scaling-inference-prefill-photonic.md(arXiv:2609.01821, 2026-09-01, cs.DC/cs.AR)。 - Ingest: AInfer-PD PDF →
raw/papers/AInfer_PD_InPlace_Prefill_Decode_MoE_2026.pdf+ stubraw/papers/ainfer-pd-inplace-prefill-decode-moe.md(arXiv:2609.00993, 2026-09-01, cs.DC)。 - Creation (papers): Photonic Prefill(3D 光子 4× SU BW / 1152 pod;高 batch 2.1–3.2×、跨 pod 2.2–4.5×);AInfer-PD(Ant;turnstile+DeepEP 相位隔离;vs Normal −7.1–22.5%、vs SGLang −24.8–32.9%)。
- Update: Disaggregated Inference(同池 P/D 复用一行 + 光学 DES),NVLink NVSwitch Scale-Up Fabric(光学 1152 pod 对照),NVIDIA CPO Roadmap(推理 prefill 量化)。
- Indexes: 手动同步
papers/index.md(+2)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: AceSpec (2609.02514, edge-cloud WAN speculative,出 WSE/NoC/NoW);Atlas 3DGS VR (2609.02352);NOVA eNVM on-chip training (2609.01948);RunSoC automotive (2609.01614);Batch Before You Time EDA (2609.02470);H3DNAS / FORGE MCU / GadIR / HDL repair;CREDIT DSMEM;MeanField GPU scheduling;Characterizing multi-tenancy AI training (2609.00817, 表征);Just Talk Once split FL (2609.01457)。先验跳过仍适用:Block-Diffusion/VARA/HBQ、Beacon、AXI4、LLM-H、survey 2608.28048、Gen-TAS、SNN、MeshReduce-U、Redwood、Ankhdjet、TerraceMoE、FPGA Transformer survey。BusyBarn 仍无公开全文 PDF。
2026-09-03
Watch (morning)
- Watch: 2026-09-03 Asia/Shanghai AI infra 论文巡检。cs.AR new/recent = Wed 9/2 提交(12+cross);Thu 9/3 美东列表上海早晨可能尚未放出。已 ingest 的 CHIPSMORE/Sync Tax/FLINT/Maia/光互连/HYDRA/ReXpert/HCCL/DASH/DICE/Fovea/C2C/ThAME/3DLS/Mozart 与全部 hc2026-* 不重复。
- Ingest: LEAP PDF →
raw/papers/LEAP_IMC_NoC_LLM_Inference_2026.pdf+ stubraw/papers/leap-imc-noc-llm-inference.md(arXiv:2609.00857, 2026-09-01, cs.AR;ICCAD’25 扩展)。 - Ingest: DynaNDE PDF →
raw/papers/DynaNDE_Near_Data_Expert_Scheduling_2026.pdf+ stubraw/papers/dynande-near-data-expert-scheduling.md(arXiv:2609.00407, 2026-09-01, cs.AR)。 - Creation (papers): LEAP(NUS;IMC+NMC+INC;LEAP-D 片上 PD;vs A100 ≥2.55×/≥71.94×,vs H100 1.52×/24.91×);DynaNDE(IIT;NPU–NDP 分析模型调度;vs MoNDE prefill/decode 2.6×/2.2×)。
- Update: Disaggregated Inference(LEAP-D 片上行),Heterogeneous Inference(IMC/NMC/INC;NPU–NDP),Interconnection Network Design Space(IRCU INC),Collective-Capable NoC(对照 LEAP IRCU),CXL Tiered Memory(MoE CXL-NDP)。
- Indexes: 手动同步
papers/index.md(+2)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: Block-Diffusion edge (2609.01084, LPDDR/systolic 压缩,无 NoC/NoW/chiplet 互连增量);VARA ReRAM (2609.00421, 通用激活稀疏 IMC);HBQ (2609.00450, MICRO 量化);SPEC CPU 2026 EPYC (2609.01527);FPGA Transformer survey (2609.01212);Analog-DB / FALCON / JENGA / energy-law / version-space / SILK replace。先验跳过仍适用:Beacon 2608.30932、AXI4 monitor、LLM-H、survey 2608.28048、Gen-TAS、SNN、MeshReduce-U、Redwood、Ankhdjet、TerraceMoE。BusyBarn(ISCA 2026)仍仅 artifact/Zenodo,无公开全文 PDF。3DLS 已入库不重复。
2026-09-02
Watch (morning)
- Watch: 2026-09-02 Asia/Shanghai AI infra 论文巡检。cs.AR new 页 = Tue 9/1 提交;Wed 9/2 美东列表上海早晨可能尚未放出。已 ingest 的 Sync Tax/FLINT/Maia/光互连/HYDRA/ReXpert/HCCL/DASH/DICE/Fovea/C2C/ThAME/3DLS/Mozart 与全部 hc2026-* 不重复。
- Ingest: CHIPSMORE PDF →
raw/papers/CHIPSMORE_CIM_Chiplets_LLM_Inference_2026.pdf+ stubraw/papers/chipsmore-cim-chiplets-llm-inference.md(arXiv:2608.30509, 2026-08-31, cs.AR)。 - Creation (papers): CHIPSMORE(NUS;RRAM-ACIM+SRAM-DCIM + IPCN in-network DMAC;分层 KV;非复制多请求层流水;vs H100 Mistral-7B INT8 最高 2.38× 吞吐 / 27× 能效)。
- Update: Disaggregated Inference(层流水共享权重一行),Heterogeneous Inference(CIM PE 异构),Interconnection Network Design Space(IPCN compute-in-interconnect)。
- Indexes: 手动同步
papers/index.md(+1)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: Beacon (2608.30932, LLM multi-agent chiplet HW-DSE,方法文、无新 fabric PHY/拓扑数,与 HYDRA 域重叠);AXI4 transaction monitoring (2608.30435, Benini mixed-crit SoC);LLM-based HW development hierarchical IRs (2608.30659, EDA);Clock-gating MSP430 (2608.30954);FABO routing (2608.30268);adiabatic systolic (2608.30058);SNN memory (2608.30444);genomic storage thesis (2608.31004);KORD (2608.30379)。先验跳过仍适用:survey 2608.28048、Gen-TAS、TerraceMoE gate-fail、VPP、CE-MoE、Blackwell CC、MeshReduce-U、edge Hydra 2608.25053、Intelligent Network WAN 2608.26453(出 WSE/NoC/NoW 范围)。BusyBarn(ISCA 2026)仍无公开全文 PDF。
2026-09-01
Watch (morning)
- Watch: 2026-09-01 Asia/Shanghai AI infra 论文巡检。无增量。cs.AR recent/new 到 Mon 8/31(5 条;Tue 9/1 美东列表上海早晨尚未放出)。已 ingest 的 Sync Tax/FLINT/Maia/光互连/HYDRA/ReXpert/HCCL/DASH/DICE/Fovea/C2C/ThAME/3DLS/Mozart 与全部 hc2026-* 不重复。
- Indexes: 无 papers 变更。未跑
generate_indexes.py。 - Considered not ingested: AI Hardware Accelerators survey (2608.28048, 综述无新一作互连数);Gen-TAS (2608.28160, FPGA-GPP 任务分配);Neuromorphic numerical solvers (2608.28387);DeepSeq3 (2608.28188, EDA GNN);RVV 1.0 HPC (2608.28097);TerraceMoE (2608.27874, MoE 分层 A2A 代价模型,step-level gate 失败、实测未进 hierarchical 区);VPP (2608.26523, Ascend chunked-prefill 软件流水);CE-MoE layer reconfig (2608.28511, 减 A2A 的模型层型,非 fabric);Blackwell CC TEE (2608.26575);LLM energy token/request (2608.28044);cache LAH/S4-FIFO (2608.27975)。Fri 8/28 已跳过项(MeshReduce-U/SNN/Redwood/Ankhdjet/HOLMES/LLM-EDA 等)仍跳过。BusyBarn(ISCA 2026)仍仅 artifact/IEEE,无公开全文 PDF。
2026-08-31
Watch (morning)
- Watch: 2026-08-31 Asia/Shanghai AI infra 论文巡检。cs.AR recent 到 Fri 8/28;Mon 8/31 美东列表上海早晨可能尚未放出。已 ingest 的 FLINT/Maia/光互连/HYDRA/ReXpert/HCCL/DASH/DICE/Fovea/C2C/ThAME/3DLS/Mozart/Iff 与全部 hc2026-* 不重复。
- Ingest: Synchronization Tax PDF →
raw/papers/Synchronization_Tax_GPU_Scale_Up_Domains_2026.pdf+ stubraw/papers/synchronization-tax-gpu-scale-up.md(arXiv:2608.22503, 2026-08-23, cs.DC)。 - Creation (papers): Synchronization Tax(Cornell;8-GPU 域集体 >50% 是 barrier 等待;增广 Hockney T=pα+qS/B+τ,B* 随域规模下降)。
- Update: NVLink fabric(带宽缩放 vs 域规模张力),LLM Collectives(墙钟含与 B 无关的 τ)。
- Indexes: 手动同步
papers/index.md(+1)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: MeshReduce-U (2608.26220, GNN/不规则 mesh NoC 编译器,非 LLM fabric);SNN multicast (2608.26223);Redwood (2608.26418, AI 设计加速器 EDA);Ankhdjet (2608.26206, 三值 CiROM);HOLMES yield (2608.26758);LLM EDA orchestration (2608.27184);DNA storage;SILK TOCTOU;vision generative edge。3D-IC Benchmark (2608.25155) 与 edge Hydra (2608.25053) 已于 8/28 跳过。BusyBarn 仍无公开 arXiv PDF。2603.22774 CPU slowdowns 替换版在窗口外。昨日跳过的 Simthesizer/NOVA/FlashAccel/TMR/SPICE 仍跳过。
2026-08-28
Watch (morning)
- Watch: 2026-08-28 Asia/Shanghai AI infra 论文巡检。cs.AR new/recent 到 Thu 8/27(7 篇新稿 + 替换);Fri 8/28 列表尚未放出。已 ingest 的 Maia/光互连/HYDRA/ReXpert/HCCL/DASH/DICE/Fovea/C2C/ThAME/3DLS/Mozart/Iff 与全部 hc2026-* 不重复。2608.24637 v2 数字未变(2.7×/3.8×/3.3×),不重入库。
- Ingest: FLINT PDF →
raw/papers/FLINT_HBF_LLM_Inference_2026.pdf+ stubraw/papers/flint-hbf-llm-inference.md(arXiv:2608.25062, 2026-08-25)。 - Creation (papers): FLINT(HBF 基座 burst-buffer / phantom-plane / 只读 FTL;级联不是 DASH 双路径)。
- Update: DASH,OXMIQ HBF,TSV,DRAM,3D Stacking。
- Indexes: 手动同步
papers/index.md(+1)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: 3D-IC Benchmark Suite (2608.25155, UCLA Gupta,CATCH→3Dblox 物理设计测试集,无 LLM/NoC/NoW);edge Hydra 表征 (2608.25053, AGX SoC + llama.cpp,与已入库 HYDRA 2608.19395 不是同一篇);APT DiT 剪枝 (2608.25380);BOOSTEDSOSA (2608.25346);Syn2Logic (2608.25536);Ising anneal (2608.26100)。cs.NI 近窗 5G/UAV/IoT/BGP/SlimTCP,无 LLM 互连。昨日跳过的 Simthesizer/NOVA/FlashAccel/TMR/SPICE/FIBER/HyperCut 与 2608.24637 v2 仍跳过。BusyBarn 仍无公开 PDF。
2026-08-27
Watch (morning)
- Watch: 2026-08-27 Asia/Shanghai AI infra 论文巡检。cs.AR new/recent 到 Wed 8/26(Tue 8/25 提交;API 最新 2608.24664 15:05 UTC)。Thu 8/27 列表尚无更新提交。cs.NI 近窗多为 5G/量子/IoT;MemChannel CXL pooling 与 WiCi 无线 GPU 不入库。已 ingest 的 HYDRA/ReXpert/HCCL/DASH/DICE/Fovea/C2C/ThAME/3DLS/Mozart/Iff 与全部 hc2026-* 不重复。
- Ingest: Maia 200 全文 PDF →
raw/papers/Maia_200_Software_Defined_Dataflow_2026.pdf+ stubraw/papers/maia-200-sdla.md(arXiv:2608.24664, 2026-08-25)。 - Ingest: 晶圆级光互连热调谐 PDF →
raw/papers/Thermal_Tuning_Wafer_Scale_Optical_Interconnect_LLM_MoE_2026.pdf+ stubraw/papers/wafer-scale-optical-interconnect-moe-thermal.md(arXiv:2608.24637, 2026-08-25)。 - Creation (papers): Maia 200 SDLA(归档全文,链到 HC 幻灯页,不重复峰值 TOPS);晶圆级光互连热 stall。
- Update: HC Maia 200(加归档指针;幻灯 <1 µs 与全文 ~4 µs 不混用),Network-on-Wafer(光 interposer 近亲),LLM Collectives,Protocol Stack(ATLv2),3D Stacking。
- Indexes: 手动同步
papers/index.md(+2)、raw/papers/index.md。未跑generate_indexes.py。 - Considered not ingested: Simthesizer (2608.24650, serving 仿真器+agent,无互连架构);Pipeline-Native Transformers (2608.23841, CPU decode 带宽);Elastic KV (2608.23658);MemChannel CXL (2608.21731, cs.NI pooling transport);WiCi (2608.24204)。昨日跳过的 NOVA/FlashAccel/TMR/SPICE/FIBER/HyperCut/VIPER/M3D SRAM/MCM GPU/SYNTLOG/Optalysys 仍跳过。BusyBarn 仍无公开 PDF。NVHBM 是新闻博客,不是论文。
2026-08-26
Hot Chips 2026 ingest
- Ingest: Day1/Day2 KEEP 幻灯 PDF →
raw/papers/HC2026_*.pdf+ Raw Source stub。数字只取 extract notes。未拷贝 Micron Confidential。 - Creation (papers): Rubin, MI455X, Helios UALoE, Crescent Island, Vera CPU, CS-4, MTIA 400, Maia 200, TPU 8, SN50, BlueField-4, Groq 3 LPX, Spectrum-X Multiplane, Jalapeño, Thor Ultra。教程/海报 7 页见同日上一小节。
- Update: Vera Rubin NVL72, Groq 3 LPX, Cerebras WSE, NVLink fabric, HCCL(MTIA 400 2D mesh / 1.2 TB/s SU), CXL Tiered Memory(Diamond Rapids CXL 3.0 1LM/Flat2LM 三行), TPU v4 OCS(链 TPU 8)。
- Schema: 公司加
microsoft / openai / broadcom / intel / sambanova;网络加ualink / ualoe。 - Indexes: 手动同步
papers/index.md(+15)、entities/index.md、raw/papers/index.md。未跑generate_indexes.py。 - Skip as full papers: Micron Confidential;Opticore;LUTs and Bolts;RISC-V / Canonical / Infineon;Day1 Arm AGI / IBM Z / Diamond Rapids(仅 CXL 三行)/ welcome / awards / Waymo / BosSemi / Fujitsu / Wildcat;Day2 Samsung LPDDR5X-PIM / XCENA MX1 / closing / Versal RF / Versal Premium Gen2。
Hot Chips 2026 tutorials / posters ingest
- Ingest: Handy / Samsung / SK hynix / d-Matrix / OXMIQ / NVIDIA Fusion 教程 PDF + Pistil 海报 PDF →
raw/papers/HC2026_*.pdf+ Raw Source stub。未拷贝、未 ingest Micron Confidential 教程。 - Creation (papers): Handy HBM 开场, zHBM, SK hynix packaging, d-Matrix Raptor, OXMIQ HBF, NVIDIA NVLink Fusion, Pistil。
- Update: 3D Stacking(SK HyB vs HBM4E;Samsung zHBM WoW+HCB), Hybrid Bonding, 3D-Stacked AI Chip(Raptor 1-Hi logic-on-top), DRAM, TSV, Network-on-Wafer(zHBM 不是晶圆级 NoW), DASH(OXMIQ 对照), NVLink fabric, Vera Rubin NVL72。
- Schema: 公司标签加
samsung / sk-hynix / d-matrix / oxmiq;技术加hbf。未加 intel / microsoft / openai。 - Indexes: 手动同步
papers/index.md(+7)、raw/papers/index.md。未跑generate_indexes.py。 - Skip: Micron Confidential;Opticore(photonic compute);LUTs and Bolts(edu ring NoC);RISC-V / Canonical / Infineon。
Watch (morning, no increment)
- Watch: 2026-08-26 Asia/Shanghai AI infra 论文巡检。cs.AR new/recent 聚焦 Tue 8/25(Wed 8/26 列表尚未出);扩展核对 FlashAccel 替换版(2607.10186v2, replaced 2026-08-22)。已 ingest 的 HYDRA/ReXpert/HCCL/DASH 等不重复。
- No increment: 本轮无 WSE/NoC/NoW/3D/chiplet/LLM 互连合格全文。不硬凑入库。
- Considered not ingested: NOVA (2608.22613, 4F² VCT + peri-over-cell 两层 NMP,HB/TSV 仅作 DRAM 堆叠带宽,无 NoC/chiplet/NoW 拓扑增量);FlashAccel (2607.10186v2, GPU 内 HBF 权重+KV 布局/SRAM 预取/管理软件,相对已 ingest 的 DASH 双路径 UCIe 无互连架构增量);SYNTLOG FSM FPGA (2608.23288);systolic PE approx (2608.22378);NoTB RTL (2608.21962);FPGA compression survey (2608.21657);Optalysys photonic compute-in-transit (2608.21536);M3D 6T SRAM 2nm (2608.22741);VIPER PIM DSE (2608.23404);MCM GPU cycle simulator (2608.22602);TherMapNet (2608.21887)。昨日跳过的 TMR/SPICE/FIBER/HyperCut/FPGA NoC CAD/DTX/SLA/H100 load/HBM reliability/MAGMA/TokenPowerSandbox/HCRMap 仍跳过。BusyBarn 仍无公开 PDF。WATOS (2512.12279) 在窗口外。
2026-08-25
- Watch: 2026-08-25 Asia/Shanghai AI infra 论文巡检。cs.AR pastweek 最新到 Mon 8/24(6 篇);Tue 8/25 列表尚未出。昨日已 ingest 的 HYDRA(2608.19395)不重复。
- No increment: 本轮无 WSE/NoC/NoW/3D/LLM 互连合格全文。不硬凑入库。
- Considered not ingested: TMR wide-link NoC router (2608.21288, Benini 组,512-bit/2-cycle/7nm TMR,偏 SEE 可靠性而非 LLM 互连拓扑);SPICE MoE prefetch (2608.21240, PCIe expert offload 投机预取,软件编排);FIBER (2608.19628);HyperCut (2608.19296, tiled NoC 层间调度,唯一偏 LLM 的是 GPT-2 decode);FPGA NoC CAD / DTX / SLA / H100 load / HBM reliability / MAGMA / TokenPowerSandbox / HCRMap 仍跳过。BusyBarn 仍无公开 PDF。
2026-08-24
- Watch: 2026-08-24 Asia/Shanghai AI infra 论文巡检。近 7 天 cs.AR(8/17–8/21)仍多 FPGA/RTL/GPU SIMT;WSE/NoC/NoW/3D/LLM 互连增量里只选出 1 篇有架构实质的全文。已 ingest 的 ReXpert/HCCL/DASH(08-21)、DICE(08-20)、Fovea/C2C-Explorer/ThAME(08-19)、Iff/3DLS/Mozart(08-18)不重复。
- Ingest: HYDRA PDF →
raw/papers/HYDRA_Heterogeneous_Chiplet_DSE_Hybrid_LLM_2026.pdf+ stubraw/papers/hydra-heterogeneous-chiplet-dse-hybrid-llm.md(arXiv:2608.19395, 2026-08-19)。 - Creation (papers): HYDRA。
- Update: C2C-Explorer, Disaggregated Inference, Interconnection Network Design Space, Interconnection Network Protocol Stack, Network-on-Wafer。
- Indexes: 手动同步
papers/index.md(+1)、raw/papers/index.md。 - Considered not ingested: FPGA NoC CAD (2608.17266);DTX (2608.16953);SLA scheduling (2608.16336);H100 global load (2608.15764);FIBER (2608.19628, GPU SIMT/寄存器);HBM reliability (2608.19471);MAGMA (2608.18366, FPGA GMM 视觉);TokenPowerSandbox (2608.18149, 能耗 serving);HCRMap (2607.11586, 3.5D MoE 映射,先前跳过、与 Mozart 重叠)。BusyBarn 仍无公开全文 PDF。本周合格增量只有 HYDRA,不硬凑 2–4 篇。
2026-08-21
- Watch: 2026-08-21 Asia/Shanghai AI infra 论文巡检。近 7–14 天 cs.AR 仍多 GNN/RTL/FPGA/SNN;WSE/NoC/NoW/3D/LLM 互连增量里选出 3 篇有架构实质的全文。已 ingest 的 Iff/3DLS/Mozart(08-18)、Fovea/C2C-Explorer/ThAME(08-19)、DICE(08-20)不重复。
- Ingest: ReXpert PDF →
raw/papers/ReXpert_MoE_ReRAM_Near_Memory_Disaggregated_Serving_2026.pdf+ stubraw/papers/rexpert-reram-nmc-disaggregated-moe.md(arXiv:2608.13962, 2026-08-14)。 - Ingest: HCCL PDF →
raw/papers/HCCL_Collective_Communication_Meta_MTIA_300_2026.pdf+ stubraw/papers/hccl-meta-mtia-300-collective-communication.md(arXiv:2608.00358;abs 自称 SC ‘26,未核程序册)。 - Ingest: DASH PDF →
raw/papers/DASH_Dual_Path_HBF_MoE_LLM_Inference_2026.pdf+ stubraw/papers/dash-dual-path-hbf-moe-inference.md(arXiv:2608.14333, 2026-08-14)。 - Creation (papers): ReXpert, HCCL, DASH。
- Update: Disaggregated Inference, LLM Distributed Training Collectives, DRAM and Memory System, TSV Physical Layer, 3D Stacking Technologies;交叉 ThAME、3DLS、C2C-Explorer、Meta RDMA。
- Schema:
SCHEMA.md公司标签加meta。 - Indexes: 手动同步
papers/index.md(+3)、raw/papers/index.md。 - Considered not ingested: FPGA NoC CAD (2608.17266, FPL 2024 投稿窗口、FPGA 布局布线,无 LLM 架构增量);DTX (2608.16953, 训练脉动/VLIW,无互连/3D/chiplet);Dryas (2608.12934, Enzian ECI 跟踪引擎);HBF Sucks / Potential Applications of HBF / Beyond Capacity 的姊妹短文 (2608.11668/13127/13868, 表征或短 CAL,架构增量已被 DASH 覆盖);Hardware Design and Security (2608.05063, 安全);SLA scheduling (2608.16336);H100 global load (2608.15764);BusyBarn 仍无公开全文 PDF(仅 artifact/IEEE);先前跳过的 SHIFT/HCRMap/SiFAR/CLIP-3D/Chiplet-Contiguous/DeepStack/3D-Flow/DyPNet-MSC/Trivance/200mm InOx/2608.15118 仍无新全文增量。
2026-08-20
- Watch: 2026-08-20 Asia/Shanghai AI infra 论文巡检。近 7–14 天 cs.AR 多为 GNN/RTL/FPGA CAD,WSE/NoC/NoW/3D/LLM 互连增量很少。已 ingest 的 Iff/3DLS/Mozart(08-18)与 Fovea/C2C-Explorer/ThAME(08-19)不重复。
- Ingest: DICE PDF →
raw/papers/DICE_Detailed_Inter_Chiplet_End_to_End_PHY_Modeling_2026.pdf+ stubraw/papers/dice-detailed-inter-chiplet-end-to-end-phy-modeling.md(arXiv:2607.24221;PDF 页眉仍为 ISCA 2026 submission draft)。 - Creation (papers): DICE。
- Update: Interconnection Network Protocol Stack(chiplet PHY/FEC/PAM4)、C2C-Explorer、Network-on-Wafer、UB 物理层。
- Indexes: 手动同步
papers/index.md(+1)、raw/papers/index.md。 - Considered not ingested: Collective Communication for Distributed LLM Systems (2608.15118, 集群级 AR/RS/AG/A2A 教程,与已有 collectives 概念重叠、非 WSE/NoC/NoW/3D);200 mm M3D InOx (2608.09508, 器件/工艺);BusyBarn (ISCA 2026 晶圆级 LLM 映射+BALD,IEEE 付费、无公开 PDF,数字无法核);Ouroboros / FlatAttention / ELMoE-3D / ATLAS 等 3–5 个月前工作超出 1–2 月窗口。先前跳过的 SHIFT/HCRMap/SiFAR/CLIP-3D/Chiplet-Contiguous/DeepStack/3D-Flow/DyPNet-MSC/Trivance 仍无新全文增量。
2026-08-19
- Watch: 2026-08-19 Asia/Shanghai AI infra 论文巡检。检索近 7–14 天并扩展到约 2 个月:WSE / NoC / NoW / 3D IC / LLM 加速器互连。昨日已 ingest 的 Iff/3DLS/Mozart 未重复。
- Ingest: Fovea PDF →
raw/papers/Fovea_Physical_Implication_Aware_Wafer_Scale_DSE_2026.pdf+ stubraw/papers/fovea-physical-implication-aware-wafer-scale-dse.md(arXiv:2608.03285, 2026-08-04)。 - Ingest: C2C-Explorer PDF →
raw/papers/C2C_Explorer_Chip_to_Chip_Interconnect_LLM_2026.pdf+ stubraw/papers/c2c-explorer-chip-to-chip-interconnect-llm.md(arXiv:2608.08611, DAC 2026 自称)。 - Ingest: ThAME PDF(v2)→
raw/papers/ThAME_3D_Memory_Enabled_Heterogeneous_MoE_2026.pdf+ stubraw/papers/thame-3d-memory-enabled-heterogeneous-moe.md(arXiv:2607.17074;昨日仅摘要、今日全文)。 - Creation (papers): Fovea, C2C-Explorer, ThAME。
- Update: Network-on-Wafer(Fovea 可行域/DSE), 3D Stacking Technologies(CBA+HB), 3D-Stacked AI Chip, LLM Distributed Training Collectives, Cerebras WSE, Mozart / WoW / 3DLS 论文页交叉引用。
- Indexes: 手动同步
papers/index.md(+3)。 - Considered not ingested: SHIFT (2606.28754, 计算搬迁 vs 数据搬迁,对比「晶圆级 LLM 服务」但主贡献偏 runtime);HCRMap (2607.11586, 3.5D MoE 映射,与 Mozart 重叠且偏调度);SiFAR (2607.08973, 软件同步-free AllReduce);HyNoC (FPGA VLIW NoC, LLaMA 只作负载);CLIP-3D (2607.12788, 通用 3D-IC 热/布局, 非 LLM);200 mm M3D InOx (2608.09508, 器件/工艺为主);Chiplet-Contiguous Layout / locality simulator (2606.11718/11716, 多 chiplet GPU GEMM);DeepStack / 3D-Flow / DyPNet-MSC / Trivance 仍缺相对 wiki 的新增量或全文无新实质。无增量条目不单列「无增量」,因本轮有三篇 ingest。
2026-08-18
- Watch (first-run): 2026-08-18 Asia/Shanghai 首轮 AI infra 论文巡检。检索 2025–2026(偏近 2–8 周)WSE / NoC / NoW / 3D IC / LLM 加速器。已有页未重复 ingest:FlooNoC collectives、Cerebras/WSE、Voxel、WaferLLM、MOCAP、hybrid bonding 综述。
- Considered not ingested: ThAME (arXiv:2607.17074, ESWEEK-26, 15.7×/9.8× 仅摘要级);DeepStack (2604.04750, 与 Voxel DSE 重叠);3D-Flow FlashAttention hybrid-bond (2602.11016);DyPNet-MSC photonic NoW (ISPASS 2026);Trivance AllReduce (2602.17254);RPU (2602.18568);CHIME (2601.19908)。
- Ingest: Iff et al. WoW NoW PDF →
raw/papers/Network_Design_Wafer_Scale_WoW_Hybrid_Bonding_2026.pdf+ stubraw/papers/network-design-wafer-scale-wow-hybrid-bonding.md(arXiv:2603.05266)。 - Ingest: 3DLS PDF →
raw/papers/3DLS_3D_Logic_Stacked_Disaggregated_LLM_Serving_2026.pdf+ stubraw/papers/3dls-3d-logic-stacked-disaggregated-llm-serving.md(arXiv:2607.01617, IEEE CAL 2026)。 - Ingest: Mozart PDF →
raw/papers/Mozart_35D_Wafer_Scale_MoE_Training_2026.pdf+ stubraw/papers/mozart-35d-wafer-scale-moe-training.md(arXiv:2603.07006)。 - Creation (papers): WoW Network Design, 3DLS, Mozart。
- Creation (concepts): Network-on-Wafer — 三条 WSI 物理路线 + 放置即拓扑。
- Update: 3D Stacking Technologies, 3D-Stacked AI Chip, Disaggregated Inference, Cerebras WSE, Post-Moore Architecture Frontiers, LLM Distributed Training Collectives,
SCHEMA.md(加now / network-on-wafer / wafer-on-wafer)。 - Indexes: 手动同步
concepts/index.md(+1)、papers/index.md(+3)。
2026-08-13
- Ingest: Corner To Corner Route 交互可视化 →
raw/articles/corner-to-corner-route.html+ 摘录raw/articles/corner-to-corner-route.md(6×8 AIC 折叠多环、RBRG 转弯服务、相位约束最短路)。 - Creation: AIC Folded Multi-Ring NoC — floorplan 常量、微边周期表、H→V→H 合法性、对角 194 cyc / 53.7 mm 复现。
- Update: Linear and Ring Topology, Mesh and Torus Topology, Deterministic Routing and DOR, Topology Optimization Variants — 折叠多环/几何 DOR 交叉引用。
- Indexes: 手动同步
concepts/index.md(+1 条)。
2026-07-31
- Layer 1 Ingest (New Study Series): 3D NoC 研究 Phase 1(Layer 1 物理层)开篇,4 raw 学记 →
raw/articles/3d-noc-study-{01-tsv-process-tech,02-monolithic-vs-tsv,03-hybrid-bonding,04-3d-mesh-baseline}.md(侧重 TSV 物理 + 三路线对比 + 商业现实 + Feero baseline)。 - Creation (papers): Katti TSV 2010 — TSV 综述原典入口;Batude Monolithic 2011 — Monolithic 综述入口;Hybrid Bonding Recent — Cu-Cu 直接键合综述;Feero 3D Mesh Stan 2008 — 3-D Mesh NoC 拓扑 baseline。
- Creation (concepts): TSV Physical Layer — TSV 工艺 + KOZ + 寄生 + 热 + 良率 五约束概念页;3D Stacking Technologies — TSV / Monolithic / Hybrid Bonding 三路线对比 + 对 3D NoC 设计含义。
- Update: 3D-Stacked AI Chip + Post-Moore Architecture Frontiers 反向链接在 #5 步补。
SCHEMA.mdtag taxonomy 加3d / tsv / monolithic / hybrid-bonding / through-silicon-via / microbump / cu-cu / packaging / integration / sequential-integration(既有chiplet / physical-layer等保留,无重复)。 - Indexes: 手动同步
concepts/index.md(+2 条)、papers/index.md(+4 条)。
2026-07-30
- Ingest: Ali (@waterloo_intern) ‘22580: From GPT2 to Kimi3, Explained’ (2026-07-27 X 长文) →
raw/articles/22580 From GPT2 to Kimi3, Explained.md(551 行,已存)。 - Creation (entities): Moonshot AI Kimi K3 — K3 模型实体 + 架构骨架 + 与 GPT-2 尺度对照。
- Creation (concepts): Linear Attention Evolution — GPT-2 → Linear Attn → DeltaNet → Gated DeltaNet → KDA 七年演化主线;Attention Residuals — AttnRes 深度方向选择性残差检索;Stable Latent MoE — K3 MoE 框架(latent-space + Quantile Balancing + 898 expert/16+2 active)。
- Creation (papers): Ali 22580 From GPT2 to Kimi3 — 论文摘要页(含全部章节结构 + 关键代码 + 数学公式)。
- Update: WaferLLM System — 新增 §“与 Kimi K3 的同构关系” + 相关概念交叉引用,frontmatter sources/updated 同步;Prefill-Decode Resource Divergence — 新增 Linear Attention Evolution 与 K3 交叉引用(KDA 改变 decode bandwidth 格局),frontmatter 同步。
- Update: SCHEMA.md tag taxonomy 实际未改动(新页 tags 全部命中既有 taxonomy:attention/moe/llm/architecture/moonshot/optimization/kernel/inference 等)。
- Indexes: 手动同步
concepts/index.md(+3 条)、entities/index.md(+1 条)、papers/index.md(+1 条)。
2026-07-24
- Ingest: 金观涛、华国凡《控制论与科学方法论》→
raw/books/控制论与科学方法论-金观涛-华国凡.mobi+ 结构化摘录raw/books/cybernetics-and-scientific-methodology.md(源:/home/luke/下载/控制论与科学方法论.mobi)。 - Creation: Cybernetics and Scientific Methodology, Black-Box Epistemology。
- Update: Architecture Paper Reading Methodology, Architecture Benchmark Methodology, Network Interface and System-Level Design, Quantitative Architecture Fundamentals — 黑箱/负反馈方法论交叉引用。
2026-07-22
- Ingest: 论文精读专项 paper-deepdive Day 1–8 + OVERVIEW →
raw/articles/paper-deepdive-day-01.md…day-08.md、paper-deepdive-overview.md(源:openclawdata/.../paper-deepdive/)。 - Creation: CMP NoC Pareto Design Tradeoffs, High-Radix Clos Adaptive Routing, TPU v4 OCS Reconfigurable Fabric, NVLink NVSwitch Scale-Up Fabric, Paper Deep-Dive Map;papers:
route-packets-not-wires,hoskote-5ghz-mesh-polaris,balfour-tiled-cmp-noc-tradeoffs,dally-virtual-channel-flow-control,kim-adaptive-routing-high-radix-clos,tpu-v4-optically-reconfigurable,nvidia-nvlink-hopper-blackwell。 - Update: WSE Reduce Algorithms, NoC Research Methodology and Case Studies, Virtual Channel Flow Control, Clos and Fat-Tree Topology, Adaptive Routing for NoC, Architecture Paper Reading Methodology, Post-Moore Architecture Frontiers, Near-Optimal Wafer-Scale Reduce, Interconn-Study 21d Knowledge Map, Cerebras WSE, Nvidia Vera Rubin NVL72 — 精读链与 OCS/NVLink 对照交叉引用。
2026-07-21
- Ingest: Dally & Towles 互连网络 Day 19–21 →
raw/articles/interconn-study-21d-day-19.md…day-21.md(源:openclawdata/.../interconn-study-21d/day-19..21.md;21 天计划完结)。 - Creation: Network Interface and System-Level Design, NoC Research Methodology and Case Studies, Interconn-Study 21d Knowledge Map。
- Update: Interconnection Network Design Space, Interconnection Network Protocol Stack, Flow Control Fundamentals, Virtual Channel Flow Control, Mesh and Torus Topology, Architecture Paper Reading Methodology, NoC Router Pipeline Optimizations, Cerebras WSE, Arch-Study 30d Knowledge Map — NI/拥塞、论文案例、21 天地图交叉引用。
2026-07-17
- Ingest: Zotero 新下载 22 篇 PDF →
raw/papers/*.pdf(上次批量 Find Available PDF / arXiv 直下)。 - Creation (papers + raw stubs): MOCAP, SuperInfer, Heterogeneous Computing Agents, pHost, Multi-Branch Self-Drafting, M5 CXL, DynaX, HyperMR, Comp Parallelism, Silent Data Corruptions, Cache-Resident LLC, PRESERVE, FlexInfer, Code-Form Planning, Mixed Precision Training, HCache, Cloud-Scale RPC, CosMoS, PANDA, Aurelia, Alibaba HPN, Meta RDMA.
- Creation (concepts): CXL Tiered Memory, Mixed Precision Training.
- Update: WaferLLM System, Prefill-Decode Divergence, Disaggregated Inference, DSpark Speculative Decoding, LLM Distributed Training Collectives, Heterogeneous Inference,
SCHEMA.md(补 speculative-decoding / dataflow / cxl / rdma)。 - Indexes:
generate_indexes.py重生成。
2026-07-14
- Ingest: Goossens et al. Æthereal NoC PDF →
raw/papers/Aethereal_Network_on_Chip_Concepts_Architectures_Implementations_2005.pdf(Zotero: 4PFJG7KE, IEEE MDT 2005, DOI 10.1109/MDT.2005.99)。 - Creation: Æthereal NoC, aethereal-network-on-chip.md,
raw/papers/aethereal-network-on-chip.md。 - Update: Switching Principles, Flow Control Fundamentals, Virtual Channel Flow Control, NoC Router 微架构, Deterministic Execution, Cerebras Color Mechanism, Interconnection Network Design Space — TDM GS / 确定性通信对照交叉引用。
2026-07-13
- Ingest: Dally & Towles 互连网络 Day 15–18 →
raw/articles/interconn-study-21d-day-15.md…day-18.md(源:openclawdata/.../interconn-study-21d/day-15..18.md)。 - Creation: Flow Control Fundamentals, Virtual Channel Flow Control, NoC Router Pipeline and Allocators, NoC Router Pipeline Optimizations。
- Update: Switching Principles, NoC Router 微架构, Deadlock-Free Routing CDG and Dally Theorem, Duato Escape VC Deadlock-Free Routing, Topology Optimization Variants, Cerebras Color Mechanism, Interconnection Network Design Space, Cerebras WSE — 流控/VC/流水线交叉引用。
- Ingest: 体系结构 30 天学习笔记 Day 27–30 →
raw/articles/arch-study-30d-day-27.md…day-30.md(源:openclawdata/.../arch-study-30d/day-27..30.md)。 - Creation: LLM Distributed Training Collectives, Architecture Paper Reading Methodology, Post-Moore Architecture Frontiers, Arch-Study 30d Knowledge Map。
- Update: AllReduce Algorithms, WSE Reduce Algorithms(补 FRED/FREDR 笔记命名对照), Parallelism Transition Point, WSE Quantitative Architecture Analysis, Cerebras WSE, Architecture Benchmark Methodology, DNN Accelerator Systolic Dataflow, Quantitative Architecture Fundamentals, Interconnection Network Design Space — Day 27–30 交叉引用。
2026-07-22
- Ingest: 4 篇 layout/NoC paper PDF →
raw/papers/MAERI_*,SIGMA_*,SmartMem_*,Venus_*。 - Creation: MAERI, SIGMA, SmartMem, Venus, Layout-Aware NoC and Flexible Dataflow Accelerators。
- Update: WaferLLM Compiler Research Gaps — 新增 Gap 7 (layout-aware mesh GEMV) + 阶段 B 加 3 个新 pass。
- Update: Cerebras WSE, FEATHER Accelerator — 反向链接到新概念页。
- Validation:
validate_bundle.py通过;11 个 index 自动重生成。
2026-07-09
- Ingest: Dally & Towles 互连网络 Day 13–14 →
raw/articles/interconn-study-21d-day-13.md、day-14.md(源:openclawdata/.../interconn-study-21d/day-13..14.md)。 - Creation: Deadlock-Free Routing CDG and Dally Theorem, Duato Escape VC Deadlock-Free Routing.
- Update: Adaptive Routing for NoC, Deterministic Routing and DOR, Interconnection Network Design Space, Mesh and Torus Topology, NoC Router 微架构, Switching Principles, Cerebras WSE — CDG/Dally、逃逸 VC、协议层死锁、Mesh vs Torus 交叉引用。
- Ingest: 体系结构 30 天学习笔记 Day 25–26 →
raw/articles/arch-study-30d-day-25.md、day-26.md(源:openclawdata/.../arch-study-30d/day-25..26.md)。 - Creation: DNN Accelerator Systolic Dataflow, WSE Quantitative Architecture Analysis.
- Update: GPU SIMT Architecture, DSA Processor Design Tradeoffs, Eyeriss Accelerator, Plasticine Accelerator, Cerebras WSE, GEMM vs GEMV in LLM Inference, WaferLLM System — 脉动/WS·OS·RS、Amdahl/Roofline/Mesh 量化、SLA vs TPU 交叉引用。
2026-07-07
-
Creation: GEMM vs GEMV in LLM Inference — 算子基础概念页:算术强度公式、Roofline、H100 decode <1% 峰值 FLOPS、prefill/decode 对应、编译器优化空间。
-
Update: Prefill-Decode Resource Divergence, WaferLLM System, FlashDecoding++, WaferLLM Compiler Research Gaps — 反向链接到新 GEMM vs GEMV 页。
-
Ingest: 4 篇 paper PDF →
raw/papers/LoopLynx_*,SambaNova_SN40L_*,LLM_Inference_Acceleration_*,AI_Accelerators_LLM_*(arXiv: 2504.09561, 2405.07518, 2410.04466, 2506.00008)。 -
Creation: LoopLynx, SambaNova SN40L, LLM Inference Hardware Survey, AI Accelerators Cross-Architecture, vLLM, WaferLLM Compiler Research Gaps。
-
Update: WaferLLM System — 补 §7.5/§8 作者承认的 3 个未解瓶颈,Cerebras WSE — 新增 compiler research gaps 链接。
-
Validation:
validate_bundle.py通过(4 个新 raw + 1 个新 analyses + 4 个新 paper 摘要 + 1 个新 concept);11 个 index 自动重生成。 -
Ingest: WaferLLM PDF →
raw/papers/WaferLLM_LLM_Inference_at_Wafer_Scale_2025.pdf(Zotero: arXiv:2502.04563v3)。 -
Creation: WaferLLM System, waferllm-wafer-scale-llm-inference.md,
raw/papers/waferllm-wafer-scale-llm-inference.md. -
Update: Cerebras WSE, Prefill-Decode Resource Divergence, SpaDA Programming Language, DSA Processor Design Tradeoffs — PLMR/MeshGEMM/V、KV shift、WSE LLM serving 交叉引用。
-
Ingest: 体系结构 30 天学习笔记 Day 24 →
raw/articles/arch-study-30d-day-24.md(源:openclawdata/.../arch-study-30d/day-24.md)。 -
Creation: GPU SIMT Architecture.
-
Update: Multicore SMT and NUCA, Instruction-Level Parallelism, Branch Prediction, DRAM and Memory System, DSA Processor Design Tradeoffs — SIMT/Warp/Tensor Core、Roofline、WSE 对照交叉引用。
-
Ingest: Dally & Towles 互连网络 Day 12 →
raw/articles/interconn-study-21d-day-12.md(源:openclawdata/.../interconn-study-21d/day-12.md)。 -
Creation: Adaptive Routing for NoC.
-
Update: Deterministic Routing and DOR, Interconnection Network Design Space, NoC Router 微架构, Cerebras WSE — 最小/VRR/VC、Duato 预告、DOR 选型交叉引用。
2026-07-06
- Ingest: Dally & Towles 互连网络 21 天学习笔记 Day 9–11 →
raw/articles/interconn-study-21d-day-09.md、day-10.md、day-11.md(源:openclawdata/.../interconn-study-21d/day-09..11.md)。 - Creation: Topology Optimization Variants, Deterministic Routing and DOR.
- Update: Interconnection Topology Metrics — 六拓扑统一比较、选型决策树(Day 10);Mesh and Torus Topology, Interconnection Network Design Space, Butterfly and MIN Topology, Flattened Butterfly Topology, Collective-Capable NoC, NoC Router 微架构, Cerebras WSE — Folding/CMesh/Express、Dally 1990、XY/e-cube 路由交叉引用。
- Ingest: 体系结构 30 天学习笔记 Day 21–23 →
raw/articles/arch-study-30d-day-21.md、day-22.md、day-23.md(源:openclawdata/.../arch-study-30d/day-21..23.md)。 - Creation: NoC Fundamentals (H&P Appendix F), End-to-End Memory Data Path, Multicore SMT and NUCA.
- Update: Interconnection Network Cost Model, DRAM and Memory System, Memory Hierarchy and Cache, Cache Coherence, Memory Consistency Model, SSD and NVMe Storage System, Quantitative Architecture Fundamentals, NoC Router 微架构, Cerebras WSE — 存储篇综合、NoC 五问、SMT/NUCA/Amdahl 交叉引用。
- Ingest: TileLoom PDF →
raw/papers/TileLoom_Automatic_Dataflow_Planning_2026.pdf(Zotero: arXiv:2512.22168v2)。 - Creation: TileLoom Compiler, tileloom-automatic-dataflow-planning.md,
raw/papers/tileloom-automatic-dataflow-planning.md. - Update: SpaDA Programming Language, Plasticine Accelerator, Collective-Capable NoC, DSA Processor Design Tradeoffs — Triton/Helion dataflow planning vs WSE SpaDA 交叉引用。
2026-07-03
- Ingest: Constable 精读笔记 →
raw/reports/constable-deepdive.md(源:openclawdata/.../superscalar-cpu/constable-deepdive.md)。 - Creation: Constable Load Elimination, constable-load-elimination.md.
- Update: Superscalar CPU Research (2023-2026), Out-of-Order Execution, Instruction-Level Parallelism, Prefill-Decode Resource Divergence — SLD/RMT/AMT、ISCA’24 Best Paper 交叉引用。
- Ingest: OpenClaw 超标量 CPU 研究综述 →
raw/reports/superscalar-cpu-final-report.md、raw/reports/superscalar-cpu-report.md(源:openclawdata/.../superscalar-cpu/FINAL-report.md)。 - Creation: Superscalar CPU Research (2023-2026), superscalar-cpu-research-2023-2026.md.
- Update: Branch Prediction, Out-of-Order Execution, Instruction-Level Parallelism, DSA Processor Design Tradeoffs, Cerebras WSE, Prefill-Decode Resource Divergence, Memory Consistency Model — Constable/Bullseye/Prophet/CVA6S+ 与 WSE/LLM 交叉引用。
- Ingest: Dally & Towles 互连网络 21 天学习笔记 Day 8 →
raw/articles/interconn-study-21d-day-08.md(源:openclawdata/.../interconn-study-21d/day-08.md)。 - Creation: Butterfly and MIN Topology.
- Update: Clos and Fat-Tree Topology, Switching Networks, Flattened Butterfly Topology, Mesh and Torus Topology — Butterfly/Omega/Banyan/Batcher-Banyan 与 Clos/Mesh 交叉引用。
- Ingest: 体系结构 30 天学习笔记 Day 20 →
raw/articles/arch-study-30d-day-20.md(源:openclawdata/.../arch-study-30d/day-20.md)。 - Creation: SSD and NVMe Storage System.
- Update: DRAM and Memory System, Memory Hierarchy and Cache, Cerebras WSE, CMX & STX, Prefill-Decode Resource Divergence, Inference Capacity Trap — FTL/RAID/NVMe/io_uring、memoryX、KV tier 交叉引用。
- Ingest: 郑启航 知乎「分布式存储架构下的矩阵乘与编译器」→
raw/articles/分布式存储架构下的矩阵乘与编译器.md(已有 clippings;补 OKF frontmatter)。 - Creation: Distributed GEMM Algorithms, distributed-gemm-and-compiler.md.
- Update: Mesh and Torus Topology, Linear and Ring Topology, AllReduce Algorithms, Parallelism Transition Point, Graphcore IPU, SpaDA Programming Language — Cannon/SUMMA/2.5D/3D GEMM 与 T10 rTensor 交叉引用。
2026-06-24
- Ingest: Rabenseifner 2004 MPI collective reduction ICCS PDF →
raw/papers/Rabenseifner_Collective_Reduction_Operations_2004.pdf(Zotero: ICCS 2004, LNCS 3036)。 - Creation: AllReduce Algorithms, rabenseifner-collective-reduction-operations.md,
raw/papers/rabenseifner-collective-reduction-operations.md. - Update: WSE Reduce Algorithms, Linear and Ring Topology, Interconnection Network Cost Model, Parallelism Transition Point, Near-Optimal Wafer-Scale Reduce — MPI Ring/RHD 与 WSE collective 谱系交叉引用。
- Ingest: Aimuyo et al. 2025 FlashMoE NeurIPS PDF →
raw/papers/FlashMoE_Fast_Distributed_MoE_Single_Kernel_2025.pdf(Zotero: arXiv:2506.04667)。 - Creation: FlashMoE Kernel, flashmoe-fast-distributed-moe-single-kernel.md,
raw/papers/flashmoe-fast-distributed-moe-single-kernel.md. - Update: MegaMoE Kernel, M2N Communication, Disaggregated Inference, Parallelism Transition Point, MegaScale-Infer — MoE EP kernel 栈交叉引用。
- Ingest: Shah et al. 2024 FlashAttention-3 PDF →
raw/papers/FlashAttention3_Asynchrony_Low_Precision_2024.pdf(Zotero: arXiv:2407.08608)。 - Creation: FlashAttention-3, flashattention-3-asynchrony-low-precision.md,
raw/papers/flashattention-3-asynchrony-low-precision.md. - Update: FlashAttention, FlashAttention-2, FlashDecoding++, Prefill-Decode Resource Divergence, flashattention-2-faster-attention.md, flashdecoding-plus-plus-llm-gpu-inference.md — FA→FA2→FA3 谱系补全。
- Ingest: Dao et al. 2022 FlashAttention NeurIPS PDF →
raw/papers/FlashAttention_Fast_IO_Aware_Attention_2022.pdf(Zotero: arXiv:2205.14135, NeurIPS 2022)。 - Creation: FlashAttention, flashattention-io-aware-exact-attention.md,
raw/papers/flashattention-io-aware-exact-attention.md. - Update: FlashAttention-2, FlashDecoding++, Prefill-Decode Resource Divergence, flashattention-2-faster-attention.md, flashdecoding-plus-plus-llm-gpu-inference.md — FA→FA2→FlashDecoding 谱系交叉引用。
- Ingest: Dao 2023 FlashAttention-2 PDF →
raw/papers/FlashAttention2_Faster_Attention_2023.pdf(Zotero: arXiv:2307.08691)。 - Creation: FlashAttention-2, flashattention-2-faster-attention.md,
raw/papers/flashattention-2-faster-attention.md. - Update: FlashDecoding++, Prefill-Decode Resource Divergence, dspark-speculative-decoding.md — prefill attention vs decode kernel 栈交叉引用。
- Ingest: Hong et al. 2024 FlashDecoding++ PDF →
raw/papers/FlashDecoding_PlusPlus_LLM_Inference_GPUs_2024.pdf(Zotero: arXiv:2311.01282)。 - Creation: FlashDecoding++, flashdecoding-plus-plus-llm-gpu-inference.md,
raw/papers/flashdecoding-plus-plus-llm-gpu-inference.md. - Update: Prefill-Decode Resource Divergence, DSpark Speculative Decoding, Heterogeneous Inference, dspark-speculative-decoding.md — decode kernel vs speculative/异构推理交叉引用。
- Ingest: Prabhakar et al. 2017 Plasticine ISCA PDF →
raw/papers/Plasticine_Reconfigurable_Parallel_Patterns_2017.pdf(Zotero: ISCA 2017, DOI 10.1145/3079856.3080256)。 - Creation: Plasticine Accelerator, plasticine-reconfigurable-parallel-patterns.md,
raw/papers/plasticine-reconfigurable-parallel-patterns.md. - Update: Basic Data-Flow Processor, SpaDA Programming Language, DSA Processor Design Tradeoffs, FEATHER Accelerator, Eyeriss Accelerator — parallel patterns CGRA、dataflow 谱系交叉引用。
- Ingest: Chen et al. 2017 Eyeriss JSSC PDF →
raw/papers/Eyeriss_Energy_Efficient_CNN_Accelerator_2017.pdf(Zotero: JSSC 2017, DOI 10.1109/JSSC.2016.2616357)。 - Creation: Eyeriss Accelerator, eyeriss-energy-efficient-cnn-accelerator.md,
raw/papers/eyeriss-energy-efficient-cnn-accelerator.md. - Update: FEATHER Accelerator, feather-reconfigurable-accelerator.md, DSA Processor Design Tradeoffs, Collective-Capable NoC, NoC Router 微架构 — RS dataflow、GIN 组播 NoC、FEATHER 固定基线交叉引用。
- Ingest: Dennis & Misunas 1975 basic data-flow processor PDF →
raw/papers/Dennis_Misunas_Basic_Data_Flow_Processor_1975.pdf(Zotero: ISCA 1975, ACM 641675.642111)。 - Creation: Basic Data-Flow Processor, dennis-misunas-basic-data-flow-processor.md,
raw/papers/dennis-misunas-basic-data-flow-processor.md. - Update: Deterministic Execution, DSA Processor Design Tradeoffs, CPU Pipeline Fundamentals, SpaDA Programming Language, Cerebras WSE — 数据流架构历史交叉引用。
- Update: Collective-Capable NoC, collective-capable-noc-ml-accelerators.md — 扩充 DCA 范式。
- Ingest: Colagrande et al. 2026 collective-capable NoC PDF →
raw/papers/Collective_Capable_NoC_ML_Accelerators_2026.pdf(Zotero: MLSys 2026, arXiv:2603.26438)。 - Creation: Collective-Capable NoC, collective-capable-noc-ml-accelerators.md,
raw/papers/collective-capable-noc-ml-accelerators.md. - Update: NoC Router 微架构, Mesh and Torus Topology, WSE Reduce Algorithms, Memory Consistency Model, Cerebras WSE — FlooNoC multicast/reduction/DCA/barrier 交叉引用。
- Ingest: 体系结构 30 天学习笔记 Day 19 →
raw/articles/arch-study-30d-day-19.md. - Creation: Memory Consistency Model.
- Update: Cache Coherence, Memory Fence and Barrier, Out-of-Order Execution, Deterministic Execution, DSA Processor Design Tradeoffs, Cerebras WSE, WSE Reduce Algorithms — SC/TSO/ARM、fence、CAS/MCS 锁、WSE barrier 交叉引用。
- Ingest: Dally & Towles 互连网络 21 天学习笔记 Day 7 →
raw/articles/interconn-study-21d-day-07.md. - Creation: Clos and Fat-Tree Topology.
- Update: Mesh and Torus Topology, Linear and Ring Topology, Interconnection Topology Metrics, Interconnection Network Design Space, Interconnection Network Cost Model, Switching Networks, Multi-plane Clos Topology for AI Training — Clos 定理、Fat-Tree、间接网络交叉引用。
- Ingest: FEATHER 论文 PDF →
raw/papers/FEATHER_Reconfigurable_Accelerator_Dataflow_Switching_2024.pdf(Zotero: Tong et al. 2024, arXiv:2405.13170)。 - Creation: FEATHER Accelerator, feather-reconfigurable-accelerator.md,
raw/papers/feather-reconfigurable-accelerator.md. - Update: 3D-Stacked AI Chip, DSA Processor Design Tradeoffs, SpaDA Programming Language — dataflow/layout 可重构交叉引用。
- Ingest: SpaDA 论文 PDF →
raw/papers/SpaDA_Spatial_Dataflow_Architecture_Programming_Language_2026.pdf(Zotero: Gianinazzi et al. 2026, arXiv:2511.09447)。 - Creation: SpaDA Programming Language, spada-spatial-dataflow-architecture.md,
raw/papers/spada-spatial-dataflow-architecture.md. - Update: Cerebras WSE, Deterministic Execution, Cerebras Color Mechanism, WSE Reduce Algorithms, Cache Coherence — SpaDA/CSL 编程模型交叉引用。
- Ingest: Dally & Towles 互连网络 21 天学习笔记 Day 6 →
raw/articles/interconn-study-21d-day-06.md. - Creation: Mesh and Torus Topology.
- Update: Interconnection Topology Metrics, Linear and Ring Topology, Interconnection Network Design Space, Interconnection Network Cost Model — 2-D Mesh/Torus、k-ary n-cube、Dally d_opt 交叉引用。
- Ingest: 体系结构 30 天学习笔记 Day 18 →
raw/articles/arch-study-30d-day-18.md. - Creation: Cache Coherence.
- Update: Memory Hierarchy and Cache, DSA Processor Design Tradeoffs, Memory Fence and Barrier, Virtual Memory and TLB, Cerebras WSE — MESI/Snooping/Directory/False Sharing 交叉引用。
- Ingest: 体系结构 30 天学习笔记 Day 17 →
raw/articles/arch-study-30d-day-17.md. - Creation: DRAM and Memory System.
- Update: Memory Hierarchy and Cache, Cerebras WSE — DRAM/HBM/内存墙交叉引用。
- Ingest: Dally & Towles 互连网络 21 天学习笔记 Day 5 →
raw/articles/interconn-study-21d-day-05.md. - Creation: Linear and Ring Topology.
- Update: Interconnection Topology Metrics, Interconnection Network Design Space — 1-D 基线、TileLink Ring 交叉引用。
- Ingest: DSpark 论文 PDF →
raw/papers/DSpark_Confidence-Scheduled_Speculative_Decoding_2026.pdf(Zotero: Cheng et al. 2026)。 - Creation: DSpark Speculative Decoding, dspark-speculative-decoding.md,
raw/papers/dspark-speculative-decoding.md. - Update: DeepSeek-V4, Prefill-Decode Resource Divergence, Heterogeneous Inference — speculative decode 交叉引用。
- Ingest: Memory Fence 深度研究报告 →
raw/articles/memory-fence-hardware-2026-06-28.md(源:openclawdata/.../notes/reports/)。 - Creation: Memory Fence and Barrier.
- Update: Out-of-Order Execution, Deterministic Execution, Virtual Memory and TLB, DSA Processor Design Tradeoffs, ISA Design Principles, Cerebras WSE — fence/coherence 交叉引用。
- Ingest: Dally & Towles 互连网络 21 天学习笔记 Day 3–4 →
raw/articles/interconn-study-21d-day-03.md,interconn-study-21d-day-04.md. - Creation: Interconnection Topology Metrics, Interconnection Network Cost Model.
- Update: Interconnection Network Design Space, Cerebras WSE — 拓扑度量、延迟/B_b 模型、Mesh vs Torus 权衡。
- Ingest: 体系结构 30 天学习笔记 Day 15–16 →
raw/articles/arch-study-30d-day-15.md,arch-study-30d-day-16.md. - Creation: Virtual Memory and TLB, DSA Processor Design Tradeoffs.
- Update: Memory Hierarchy and Cache, Cerebras WSE, Deterministic Execution — TLB/核心篇总结交叉引用。
- Ingest: Hennessy & Patterson 30 天体系结构学习笔记 Day 1–14 →
raw/articles/arch-study-30d-day-*.md(14 文件)。 - Creation: Quantitative Architecture Fundamentals, ISA Design Principles, Numeric Formats for AI Hardware, Architecture Benchmark Methodology, CPU Pipeline Fundamentals, Instruction-Level Parallelism, Out-of-Order Execution, Branch Prediction, Memory Hierarchy and Cache.
- Update: Cerebras WSE, Deterministic Execution, FP4 Quantization-Aware Training — 与 CPU 体系结构概念交叉引用。
- Schema: 标签 taxonomy 新增
isa,pipeline,cache,power。 - Ingest: Dally & Towles 互连网络 21 天学习笔记 Day 1–2 →
raw/articles/interconn-study-21d-day-01.md,interconn-study-21d-day-02.md. - Creation: Interconnection Network Design Space, Interconnection Network Protocol Stack.
- Update: Switching Principles — 报文/虫孔交换、历史里程碑;Cerebras WSE — Mesh 度量与虫孔选型。
- Schema: 标签 taxonomy 新增
interconnect。 - Cleanup: 删除重复的
references/raw/(OKF 转换副本);唯一原始资料目录为raw/。megascale-infer-2504.02263.pdf、cassini-network-aware-scheduling-2308.00852.pdf本就位于raw/papers/(与references/raw/papers/为同内容副本),无需迁移。 - Docs: README 与 OKF skill 统一为仅使用
raw/。 - Creation: Graphcore IPU, Core Group (DRAM Access Synchronization).
- Update: 3D-Stacked AI Chip, Voxel Simulator, Voxel 3D-Stacked AI Chip LLM Inference — 交叉引用拆分页。
- Schema: 标签 taxonomy 新增
graphcore。 - Ingest: Voxel 3D-Stacked AI Chip LLM Inference from
raw/papers/Exploring the efficiency of 3D-stacked AI chip architecture for LLM inference with voxel.pdf(arXiv:2604.26821). - Creation: 3D-Stacked AI Chip, Voxel Simulator.
- Update: Prefill-Decode Resource Divergence — 3D chip prefill/decode 设计空间差异。
- Creation: Converted LLM wiki at
/home/luke/wikito OKF v0.1 bundle (54 work pages + raw sources). - Source: Karpathy-style LLM wiki (entities, concepts, papers, summaries, analyses).
- Update: Generated interactive
viz.html(74 concepts, 237 cross-links).