flash-attention
Tracked open-source repos tagged flash-attention, sorted by stars.
Related topics
Topics that frequently appear alongside flash-attention on the same repo.
Recent risers
Repos created in the last 90 days, tagged flash-attention.
- #1
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
★ 21,674+47Star change over the last 7 days - #2
Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
★ 11,850+47Star change over the last 7 days - #3
Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).
★ 7,273+6Star change over the last 7 days - #4
中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)
★ 7,120-2Star change over the last 7 days - #5
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
★ 5,479+8Star change over the last 7 days - #6★ 2,174+7Star change over the last 7 days
- #7
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
★ 921+5Star change over the last 7 days - #8
rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffusion Distillation at Scale
★ 795+8Star change over the last 7 days - #9
Trainable fast and memory-efficient sparse attention
★ 754+2Star change over the last 7 days - #10
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.
★ 559+35Star change over the last 7 days