AI Infra Wiki
搜索
Search
暗色模式
亮色模式
探索
标签: scale-out
此标签下有10条笔记。
2026年10月08日
NCCL M2N: A Layout- and Topology-Aware Collective for Distributed Tensor Resharding
collective
communication
training
distributed
scale-up
scale-out
nvidia
moe
2026年10月06日
RailWave: Adaptive Spatial and Temporal Scheduling for Expert-Parallel Communication
moe
expert-parallelism
communication
collective
scale-out
rdma
congestion-control
scheduling
training
gpu
nvidia
networking
2026年10月05日
GPU-Initiated Communication: Dissecting Down to the Bone
architecture
interconnect
rdma
scale-out
networking
communication
collective
moe
expert-parallelism
gpu
nvidia
latency
benchmark
2026年9月14日
AI Infra Book Ch.7 Datacenter Network
book
networking
scale-out
fabric
rdma
topology
congestion-control
datacenter
2026年9月10日
Sharing a Fabric with Collective Communication: Two Storage Penalties
fabric
communication
collective
training
storage
rdma
scale-out
networking
infrastructure
distributed
2026年9月08日
REACT: Tuning Collective Patterns to Alleviate Congestion in Shared AI Clusters
llm
training
collective
communication
congestion-control
fabric
scale-out
gpu
nvidia
training-system
2026年8月26日
Hot Chips 2026: Broadcom Thor Ultra
broadcom
rdma
scale-out
congestion-control
protocol
interconnect
switch
fabric
architecture
2026年8月26日
Hot Chips 2026: NVIDIA BlueField-4
nvidia
scale-out
interconnect
fabric
rack
rdma
storage
protocol
architecture
2026年8月26日
Hot Chips 2026: NVIDIA Spectrum-X Multiplane
nvidia
scale-out
fabric
switch
topology
congestion-control
cpo
optical
interconnect
architecture
2026年8月26日
HCCL: Collective Communication for Meta MTIA 300
chiplet
interconnect
scale-up
scale-out
communication
rdma
accelerator
training
inference
fabric
protocol
rack
hbm
architecture
moe
meta