AI Infra Wiki
搜索
Search
暗色模式
亮色模式
探索
标签: cpu
此标签下有26条笔记。
2026年10月07日
HiNa-MoE: High-Performance, Non-Intrusive MoE Inference on CPUs with Matrix Engines
moe
inference
decode
cpu
kernel
memory-bandwidth
memory
intel
hardware
2026年10月06日
Superscalar CPU Research (2023-2026)
architecture
cpu
superscalar
isca
research-survey
2026年10月06日
Understanding Silent Data Corruptions in a Large Production CPU Population
cpu
infrastructure
benchmark
optimization
2026年10月06日
Superscalar CPU Research (2023-2026)
architecture
cpu
superscalar
branch-prediction
load-elimination
risc-v
llm
wse
2026年10月06日
Vera ETL256
nvidia
cpu
storage
hardware
networking
2026年10月06日
Quantitative Architecture Fundamentals
architecture
cpu
benchmark
power
wse
2026年10月06日
Constable: Safely Eliminating Load Instruction Execution
architecture
cpu
load-elimination
isca
power
ooo
2026年10月06日
Cache-Resident LLM Inference in GB-Scale LLCs
inference
cache
cpu
memory
llm
kv-cache
parallelism
optimization
2026年10月06日
Multicore SMT and NUCA
architecture
cpu
multicore
smt
numa
nuca
amdahl
wse
2026年10月06日
Out-of-Order Execution
architecture
cpu
pipeline
scheduling
2026年10月06日
ISA Design Principles
architecture
cpu
compiler
programming-model
2026年10月06日
Virtual Memory and TLB
architecture
memory
cpu
virtualization
wse
2026年10月06日
Memory Consistency Model
architecture
memory
cpu
multicore
wse
deterministic
2026年10月06日
Memory Fence and Barrier
architecture
cpu
noc
protocol
deterministic
compiler
2026年10月06日
Instruction-Level Parallelism
architecture
cpu
pipeline
parallelism
2026年10月06日
Cache Coherence
architecture
memory
cache
cpu
multicore
wse
2026年10月06日
CPU Pipeline Fundamentals
architecture
cpu
pipeline
2026年10月06日
Branch Prediction
architecture
cpu
pipeline
2026年10月06日
Constable Load Elimination
architecture
cpu
load-elimination
ooo
power
llm
isca
2026年10月06日
EdgeAgent: On-Device LLM Inference for End-User Multi-Agent Systems on CPU-GPU UMA
agentic-ai
ai-agent
inference
decode
speculative-decoding
memory-bandwidth
memory
cpu
gpu
scheduling
hardware
2026年10月05日
RapidMoE: Adaptive Residual Offloading for Large-Scale MoE Inference
architecture
llm
moe
inference
decode
prefill
quantization
memory
cpu
gpu
serving
throughput
2026年9月23日
DSA Processor Design Tradeoffs
architecture
accelerator
wse
cpu
deterministic
compiler
2026年9月16日
BOOST: Concurrent Host+HBM Access for LLM Inference
gpu
hbm
memory
memory-bandwidth
llm
inference
kv-cache
serving
serving-system
throughput
latency
architecture
nvidia
cpu
2026年9月08日
Huawei’s τ Chip Was Supposed to Melt?(LogicFolding / Hybrid Bonding)
3d
hybrid-bonding
power
huawei
interconnect
packaging
architecture
cpu
chiplet
2026年8月26日
Hot Chips 2026: NVIDIA RISC-V CPUs and NVLink Fusion
nvidia
scale-up
interconnect
fabric
gpu
cpu
isa
architecture
chiplet
switch
protocol
2026年8月26日
Hot Chips 2026: NVIDIA Vera CPU
nvidia
cpu
scale-up
cxl
memory
inference
rack
interconnect
architecture
agentic-ai
memory-bandwidth