AI Infra Wiki
搜索
Search
暗色模式
亮色模式
探索
标签: ai-agent
此标签下有11条笔记。
2026年10月06日
智能体辅助编程的信息论价值模型
ai-agent
information-theory
knowledge-management
formal-analysis
2026年10月06日
Heterogeneous Computing for AI Agent Inference
agentic-ai
ai-agent
inference
llm
memory
accelerator
memory-bandwidth
optimization
2026年10月06日
EdgeAgent: On-Device LLM Inference for End-User Multi-Agent Systems on CPU-GPU UMA
agentic-ai
ai-agent
inference
decode
speculative-decoding
memory-bandwidth
memory
cpu
gpu
scheduling
hardware
2026年9月30日
RR-Evict: Round-Robin Prefix Cache Eviction for Agentic Serving
architecture
llm
inference
serving
serving-system
kv-cache
agentic-ai
ai-agent
cache
memory
scheduling
disaggregated-inference
throughput
latency
gpu
2026年9月29日
DynBranch: Speculative Subgraph Reuse for Agentic Serving
architecture
llm
inference
serving
serving-system
agentic-ai
ai-agent
kv-cache
scheduling
latency
throughput
speculative-decoding
2026年9月23日
AHRR: 更高抽象能否让 Agent 设计更好的芯片
architecture
agentic-ai
ai-agent
accelerator
compiler
methodology
llm
2026年9月18日
Ask the Tool: Progress-Aware KV for Agentic Serving
agentic-ai
ai-agent
serving
serving-system
kv-cache
llm
inference
scheduling
latency
hbm
memory
architecture
2026年9月18日
Fathom: Per-Query Bit-Plane Scan for Offloaded KV
kv-cache
inference
serving
llm
memory
memory-bandwidth
agentic-ai
ai-agent
quantization
sparse
latency
throughput
architecture
hbm
2026年9月17日
PipeSwift: Pipeline Parallelism for Completion-Oriented Agentic Serving
agentic-ai
ai-agent
serving
serving-system
disaggregated-inference
moe
llm
inference
prefill
decode
scheduling
parallelism
pipeline
speculative-decoding
throughput
latency
architecture
2026年9月16日
PDD: Cross-Datacenter Prefill-Decode Disaggregation
disaggregated-inference
prefill
decode
kv-cache
serving
serving-system
llm
inference
latency
throughput
rdma
datacenter
interconnect
agentic-ai
ai-agent
architecture
2026年9月16日
UNISON: Near-Memory Session KV Scheduler for LLM Agents
agentic-ai
ai-agent
kv-cache
memory
hbm
sram
accelerator
scheduling
llm
inference
serving
latency
architecture
hardware