AI Infra Wiki
搜索
Search
暗色模式
亮色模式
探索
标签: pipeline
此标签下有13条笔记。
2026年10月06日
MOCAP: Wafer-Scale Chunked Pipelining for Prefill-Only LLM Inference
wse
prefill
inference
kv-cache
pipeline
parallelism
throughput
latency
llm
2026年10月06日
Out-of-Order Execution
architecture
cpu
pipeline
scheduling
2026年10月06日
NoC Router Pipeline and Allocators
noc
router
pipeline
crossbar
arbitration
islip
virtual-channel
2026年10月06日
NoC Router Pipeline Optimizations
noc
router
pipeline
speculation
look-ahead
high-radix
cmesh
buffer
2026年10月06日
Instruction-Level Parallelism
architecture
cpu
pipeline
parallelism
2026年10月06日
CPU Pipeline Fundamentals
architecture
cpu
pipeline
2026年10月06日
Branch Prediction
architecture
cpu
pipeline
2026年9月23日
SPECTRA: 运行时可重构瓦片上的 Speculative Decoding
architecture
accelerator
inference
llm
speculative-decoding
dataflow
pipeline
noc
throughput
latency
2026年9月17日
PipeSwift: Pipeline Parallelism for Completion-Oriented Agentic Serving
agentic-ai
ai-agent
serving
serving-system
disaggregated-inference
moe
llm
inference
prefill
decode
scheduling
parallelism
pipeline
speculative-decoding
throughput
latency
architecture
2026年9月17日
The World Model Hardware Accelerator (WMHA)
accelerator
architecture
pipeline
attention
inference
llm
agentic-ai
dataflow
isa
sram
hardware
2026年9月15日
Vortex: Bridging Extreme Compression and Efficient LLM Inference
accelerator
dataflow
quantization
compression
sparse
llm
inference
kv-cache
prefill
decode
throughput
latency
architecture
memory
pipeline
2026年9月03日
LEAP: LLM Inference on IMC-NoC with Balanced Dataflow and Fine-Grained Parallelism
noc
interconnect
mesh
accelerator
memory
llm
inference
kv-cache
serving
throughput
latency
architecture
pipeline
power
batching
prefill
decode
disaggregated-inference
dataflow
2026年9月02日
CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference
chiplet
noc
interconnect
mesh
accelerator
sram
memory
llm
inference
kv-cache
serving
throughput
latency
packaging
architecture
pipeline
power
batching