Skip to main content
buildradar
Sign in
Topic · flash-attention

flash-attention

Tracked open-source repos tagged flash-attention, sorted by stars.

Repos
10
Total stars
58,567
Avg. stars
5,857
Share
0.00%

Topics that frequently appear alongside flash-attention on the same repo.

Recent risers

Repos created in the last 90 days, tagged flash-attention.

  • tessera@zengxiao-he

    From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.

    559
  • Qwen@QwenLM

    The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.

    21,674+47Star change over the last 7 days
  • LeetCUDA@xlite-dev

    Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.

    11,850+47Star change over the last 7 days
  • InternLM@InternLM

    Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).

    7,273+6Star change over the last 7 days
  • 中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)

    7,120-2Star change over the last 7 days
  • Awesome-LLM-Inference@xlite-dev

    📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

    5,479+8Star change over the last 7 days
  • MoBA@MoonshotAI

    MoBA: Mixture of Block Attention for Long-Context LLMs

    2,174+7Star change over the last 7 days
  • cudnn-frontend@NVIDIA

    cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.

    921+5Star change over the last 7 days
  • rcm@NVlabs

    rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffusion Distillation at Scale

    795+8Star change over the last 7 days
  • Trainable fast and memory-efficient sparse attention

    754+2Star change over the last 7 days
  • tessera@zengxiao-he

    From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.

    559+35Star change over the last 7 days
← Back to topics