AI Infra Wiki

标签: transformer

此标签下有8条笔记。

  • 2026年10月06日

    Multi-Branch Self-Drafting for LLM Inference Acceleration

    • inference
    • decode
    • llm
    • optimization
    • serving
    • transformer
  • 2026年10月06日

    DynaX: Dynamic X:M Sparse Attention Acceleration

    • attention
    • sparse
    • accelerator
    • optimization
    • transformer
    • llm
    • kernel
    • hardware
  • 2026年9月14日

    AI Infra Book Ch.2 Model Architecture

    • book
    • transformer
    • llm
    • moe
    • kv-cache
    • attention
    • memory
  • 2026年9月08日

    FlexPosit: Tunable Fractional Precision for LLM Inference Accelerators

    • llm
    • accelerator
    • quantization
    • inference
    • throughput
    • dataflow
    • transformer
    • inference-system
  • 2026年9月07日

    BASP: Communication-Efficient Batch-Aware Sequence Parallelism for LLM Training

    • llm
    • training
    • parallelism
    • communication
    • attention
    • transformer
    • collective
    • gpu
    • nvidia
    • training-system
  • 2026年9月07日

    Einsummable: Automatic Multi-GPU Parallelism via Join-Agg Specs

    • llm
    • parallelism
    • communication
    • transformer
    • attention
    • gpu
    • nvidia
    • training-system
    • inference-system
    • optimization
    • collective
  • 2026年7月30日

    Linear Attention Evolution

    • attention
    • llm
    • transformer
    • architecture
    • model
    • moonshot
    • inference
    • optimization
  • 2026年7月30日

    22580: From GPT-2 to Kimi K3, Explained

    • attention
    • moonshot
    • llm
    • transformer
    • architecture
    • model
    • inference

Created with Quartz v4.5.1 © 2026

  • Source Wiki
  • Quartz