AI Infra Wiki
搜索
Search
暗色模式
亮色模式
探索
标签: expert-parallelism
此标签下有14条笔记。
2026年10月06日
MegaScale-Infer
moe
disaggregated-inference
expert-parallelism
serving-system
bytedance
2026年10月06日
AI Accelerators for Large Language Model Inference: Architecture Analysis and Scaling Strategies
accelerator
survey
gpu
tpu
wse
lpu
inference
scaling
expert-parallelism
2026年10月06日
FlashMoE: Fast Distributed MoE in a Single Kernel
moe
gpu
kernel
expert-parallelism
communication
neurips
2026年10月06日
FlashMoE Kernel
moe
gpu
kernel
expert-parallelism
communication
llm
training
inference
2026年10月06日
AFORE: Attention–FFN Disaggregation with Overlapped Reconfiguration of Experts
moe
expert-parallelism
disaggregated-inference
serving-system
inference
decode
scheduling
latency
throughput
gpu
nvidia
2026年10月06日
RailWave: Adaptive Spatial and Temporal Scheduling for Expert-Parallel Communication
moe
expert-parallelism
communication
collective
scale-out
rdma
congestion-control
scheduling
training
gpu
nvidia
networking
2026年10月05日
GPU-Initiated Communication: Dissecting Down to the Bone
architecture
interconnect
rdma
scale-out
networking
communication
collective
moe
expert-parallelism
gpu
nvidia
latency
benchmark
2026年10月05日
MegaFlux: Skew-Resilient MoE Megakernels via Pipelined Expert Replication
architecture
llm
moe
training
inference
expert-parallelism
kernel
gpu
nvidia
communication
scheduling
throughput
2026年10月02日
HAPMoE: Heterogeneity-Aware Automatic Parallelism for MoE Training
architecture
llm
moe
training
training-system
expert-parallelism
collective
distributed
parallelism
scheduling
throughput
2026年10月02日
ThunderEP: Expert-Parallel Communication on PCIe Consumer GPUs
architecture
llm
moe
inference
serving
expert-parallelism
collective
interconnect
gpu
nvidia
distributed
kernel
throughput
2026年10月01日
Mixture-of-Kittens: MoE Megakernel for NVL72s
architecture
llm
moe
training
training-system
expert-parallelism
collective
gpu
nvidia
scale-up
interconnect
kernel
throughput
distributed
2026年9月22日
Weave: MoE Megakernel 内的细粒度动态 SM 调度
architecture
inference
llm
moe
expert-parallelism
interconnect
gpu
kernel
scheduling
2026年9月10日
HDA-MoE: Hybrid Parallelism for MoE on 3D Near-Memory Processing
moe
inference
3d
hybrid-bonding
noc
mesh
scheduling
expert-parallelism
llm
memory
interconnect
accelerator
architecture
parallelism
2026年9月03日
DynaNDE: Dynamic Near-Data Expert Scheduling for Batched MoE Inference
moe
inference
cxl
scheduling
memory
accelerator
llm
throughput
interconnect
batching
expert-parallelism
serving
architecture