AI Infra Wiki
搜索
Search
暗色模式
亮色模式
探索
标签: benchmark
此标签下有13条笔记。
2026年10月07日
Beyond LLM Serving: Characterizing Vision-Language-Action Workloads for Embodied AI System Design
inference
latency
benchmark
hardware
gpu
memory-bandwidth
power
agentic-ai
architecture
2026年10月06日
Understanding Silent Data Corruptions in a Large Production CPU Population
cpu
infrastructure
benchmark
optimization
2026年10月06日
Understanding Inference Scaling for LLMs
inference
reasoning
parallelism
kv-cache
decode
prefill
latency
throughput
serving
benchmark
2026年10月06日
Voxel: 3D-Stacked AI Chip Efficiency for LLM Inference
accelerator
chiplet
memory-bandwidth
inference
prefill
decode
noc
interconnect
architecture
benchmark
2026年10月06日
Optimizing the Parallelism of Communication and Computation in Distributed Training Platform
training
training-system
parallelism
communication
optimization
benchmark
2026年10月06日
Graphcore IPU
graphcore
accelerator
inference
mesh
sram
memory-bandwidth
benchmark
2026年10月06日
Quantitative Architecture Fundamentals
architecture
cpu
benchmark
power
wse
2026年10月06日
Voxel Simulator
architecture
benchmark
inference
accelerator
compiler
2026年10月06日
Architecture Benchmark Methodology
architecture
benchmark
methodology
formal-analysis
2026年10月06日
Architecture Paper Reading Methodology
methodology
paper
research
architecture
benchmark
wse
reduce
2026年10月06日
Divide and Conquer: Scalable Performance and Energy in MCM GPUs
architecture
gpu
chiplet
noc
interconnect
topology
mesh
routing
packaging
power
benchmark
2026年10月05日
GPU-Initiated Communication: Dissecting Down to the Bone
architecture
interconnect
rdma
scale-out
networking
communication
collective
moe
expert-parallelism
gpu
nvidia
latency
benchmark
2026年9月15日
Dissecting GPU Utilization for LLM Inference on Nvidia Hopper
gpu
nvidia
inference
decode
prefill
attention
kernel
llm
throughput
latency
batching
serving
architecture
benchmark
memory-bandwidth
moe