AI Infra Wiki
搜索
Search
暗色模式
亮色模式
探索
标签: hardware
此标签下有15条笔记。
2026年10月07日
HiNa-MoE: High-Performance, Non-Intrusive MoE Inference on CPUs with Matrix Engines
moe
inference
decode
cpu
kernel
memory-bandwidth
memory
intel
hardware
2026年10月07日
Terracotta: Enabling the Adoption of New DRAM Techniques via a Flexible DRAM Interface and Memory Controller
memory
pim
hardware
architecture
programming-model
2026年10月07日
Beyond LLM Serving: Characterizing Vision-Language-Action Workloads for Embodied AI System Design
inference
latency
benchmark
hardware
gpu
memory-bandwidth
power
agentic-ai
architecture
2026年10月06日
HyperMR: Efficient Hypergraph-enhanced Matrix Storage on Compute-in-Memory Architecture
memory
accelerator
hardware
sparse
optimization
kernel
2026年10月06日
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
survey
llm
inference
hardware
quantization
speculative-decoding
sparse-attention
fpga
2026年10月06日
Vera ETL256
nvidia
cpu
storage
hardware
networking
2026年10月06日
DynaX: Dynamic X:M Sparse Attention Acceleration
attention
sparse
accelerator
optimization
transformer
llm
kernel
hardware
2026年10月06日
UB 事务层
interconnect
scale-up
fabric
protocol
hardware
2026年10月06日
UB 内存管理
interconnect
scale-up
memory
hardware
virtualization
2026年10月06日
UB 资源管理
interconnect
scale-up
fabric
virtualization
hardware
2026年10月06日
CMX & STX
nvidia
storage
inference
kv-cache
hardware
2026年10月06日
EdgeAgent: On-Device LLM Inference for End-User Multi-Agent Systems on CPU-GPU UMA
agentic-ai
ai-agent
inference
decode
speculative-decoding
memory-bandwidth
memory
cpu
gpu
scheduling
hardware
2026年9月22日
COMET: 丢包 WAN 上纠删码 RDMA 的 FPGA 包追踪
architecture
networking
interconnect
accelerator
memory
rdma
transport
hardware
2026年9月17日
The World Model Hardware Accelerator (WMHA)
accelerator
architecture
pipeline
attention
inference
llm
agentic-ai
dataflow
isa
sram
hardware
2026年9月16日
UNISON: Near-Memory Session KV Scheduler for LLM Agents
agentic-ai
ai-agent
kv-cache
memory
hbm
sram
accelerator
scheduling
llm
inference
serving
latency
architecture
hardware