AI Infra Wiki
搜索
Search
暗色模式
亮色模式
探索
标签: batching
此标签下有8条笔记。
2026年9月29日
EAServe: Encode-Aware Disaggregated MLLM Serving
architecture
llm
inference
serving
serving-system
disaggregated-inference
prefill
decode
batching
scheduling
throughput
latency
gpu
2026年9月15日
Dissecting GPU Utilization for LLM Inference on Nvidia Hopper
gpu
nvidia
inference
decode
prefill
attention
kernel
llm
throughput
latency
batching
serving
architecture
benchmark
memory-bandwidth
moe
2026年9月15日
RoofLang: Enabling AI-Driven Architecting of LLM Inference Systems
inference
serving-system
llm
decode
prefill
kv-cache
moe
deepseek
nvidia
architecture
optimization
parallelism
throughput
latency
batching
distributed
scale-up
2026年9月14日
AI Infra Book Ch.8 Inference Optimization
book
inference
serving
kv-cache
batching
quantization
speculative-decoding
2026年9月04日
Scaling Inference Prefill with High-Radix Photonic Interconnects
photonic
optical
cpo
lightmatter
interconnect
scale-up
fabric
moe
llm
inference
prefill
disaggregated-inference
serving
architecture
3d
packaging
latency
throughput
batching
2026年9月03日
DynaNDE: Dynamic Near-Data Expert Scheduling for Batched MoE Inference
moe
inference
cxl
scheduling
memory
accelerator
llm
throughput
interconnect
batching
expert-parallelism
serving
architecture
2026年9月03日
LEAP: LLM Inference on IMC-NoC with Balanced Dataflow and Fine-Grained Parallelism
noc
interconnect
mesh
accelerator
memory
llm
inference
kv-cache
serving
throughput
latency
architecture
pipeline
power
batching
prefill
decode
disaggregated-inference
dataflow
2026年9月02日
CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference
chiplet
noc
interconnect
mesh
accelerator
sram
memory
llm
inference
kv-cache
serving
throughput
latency
packaging
architecture
pipeline
power
batching