AI Infra Wiki

标签: transformer

此标签下有4条笔记。

  • 2026年8月25日

    Multi-Branch Self-Drafting for LLM Inference Acceleration

    • inference
    • decode
    • llm
    • optimization
    • serving
    • transformer
  • 2026年8月25日

    DynaX: Dynamic X:M Sparse Attention Acceleration

    • attention
    • sparse
    • accelerator
    • optimization
    • transformer
    • llm
    • kernel
    • hardware
  • 2026年7月30日

    Linear Attention Evolution

    • attention
    • llm
    • transformer
    • architecture
    • model
    • moonshot
    • inference
    • optimization
  • 2026年7月30日

    22580: From GPT-2 to Kimi K3, Explained

    • attention
    • moonshot
    • llm
    • transformer
    • architecture
    • model
    • inference

Created with Quartz v4.5.1 © 2026

  • Source Wiki
  • Quartz