AI Infra Wiki

标签: gemm

此标签下有7条笔记。

  • 2026年10月06日

    分布式存储架构下的矩阵乘与编译器

    • gemm
    • distributed-memory
    • mpi
    • compiler
    • tensor-parallelism
    • zhihu
  • 2026年10月06日

    SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN Training

    • accelerator
    • gemm
    • sparse
    • training
    • noc
    • flexible
    • interconnect
    • hpca
    • krishna
  • 2026年10月06日

    WaferLLM: Large Language Model Inference at Wafer Scale

    • cerebras
    • wse
    • llm
    • inference
    • gemm
    • gemv
    • mesh
    • noc
    • plmr
  • 2026年10月06日

    A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators

    • noc
    • collective
    • multicast
    • reduction
    • accelerator
    • mesh
    • gemm
  • 2026年10月06日

    Distributed GEMM Algorithms

    • gemm
    • distributed-memory
    • mpi
    • tensor-parallelism
    • mesh
    • compiler
    • communication
  • 2026年9月23日

    GEMM vs GEMV in LLM Inference

    • gemm
    • gemv
    • arithmetic-intensity
    • roofline
    • compute-bound
    • bandwidth-bound
    • prefill
    • decode
    • matmul
    • kernel
  • 2026年7月30日

    WaferLLM System

    • cerebras
    • wse
    • llm
    • inference
    • gemm
    • gemv
    • mesh
    • noc
    • plmr
    • kv-cache

Created with Quartz v4.5.1 © 2026

  • Source Wiki
  • Quartz