attention
Tracked open-source repos tagged attention, sorted by stars.
Related topics
Topics that frequently appear alongside attention on the same repo.
Recent risers
Repos created in the last 90 days, tagged attention.
No new repos tagged with this topic in the last 90 days.
- #1
🧑🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), gans(cyclegan, stylegan2, ...), 🎮 reinforcement learning (ppo, dqn), capsnet, distillation, ... 🧠
★ 67,389+21Star change over the last 7 days - #2
SGLang is a high-performance serving framework for large language models and multimodal models.
★ 33,538+791Star change over the last 7 days - #3
Natural Language Processing Tutorial for Deep Learning Researchers
★ 14,927+3Star change over the last 7 days - #4
A theoretical reconstruction of the Claude Mythos architecture, built from first principles using the available research literature.
★ 14,867+14Star change over the last 7 days - #5
🍀 Pytorch implementation of various Attention Mechanisms, MLP, Re-parameter, Convolution, which is helpful to further understand papers.⭐⭐⭐
★ 12,185+1Star change over the last 7 days - #6
A PyTorch implementation of the Transformer model in "Attention is All You Need".
★ 9,782-1Star change over the last 7 days - #7
FlashInfer: Kernel Library for LLM Serving
★ 6,321+38Star change over the last 7 days - #8
Tutorials on implementing a few sequence-to-sequence (seq2seq) models with PyTorch and TorchText.
★ 5,708+1Star change over the last 7 days - #9
Transformer: PyTorch Implementation of "Attention Is All You Need"
★ 4,651+3Star change over the last 7 days - #10★ 3,821+0Star change over the last 7 days
- #11
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
★ 3,699+14Star change over the last 7 days - #12
This is a pytorch repository of YOLOv4, attentive YOLOv4 and mobilenet YOLOv4 with PASCAL VOC and COCO
★ 1,676+0Star change over the last 7 days - #13
A compilation of the best multi-agent papers
★ 1,664+6Star change over the last 7 days - #14
Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAI
★ 1,398+3Star change over the last 7 days - #15
Replication of simple CV Projects including attention, classification, detection, keypoint detection, etc.
★ 1,260-1Star change over the last 7 days - #16
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
★ 1,094+15Star change over the last 7 days - #17
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.
★ 1,045+1Star change over the last 7 days - #18
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
★ 931+11Star change over the last 7 days - #19
A PyTorch library for all things Reinforcement Learning (RL) for Combinatorial Optimization (CO)
★ 900+1Star change over the last 7 days - #20
Implementation of the sparse attention pattern proposed by the Deepseek team in their "Native Sparse Attention" paper
★ 813+1Star change over the last 7 days - #21
多标签文本分类,多标签分类,文本分类, multi-label, classifier, text classification, BERT, seq2seq,attention, multi-label-classification
★ 805+0Star change over the last 7 days - #22
[TNSRE 23] EEG Transformer 2.0. i. Convolutional Transformer for EEG Decoding. ii. Novel visualization - Class Activation Topography.
★ 760+0Star change over the last 7 days - #23
Implementation of plug in and play Attention from "LongNet: Scaling Transformers to 1,000,000,000 Tokens"
★ 726+0Star change over the last 7 days - #24
Achieve the llama3 inference step-by-step, grasp the core concepts, master the process derivation, implement the code.
★ 630+0Star change over the last 7 days - #25
Keras beit,caformer,CMT,CoAtNet,convnext,davit,dino,efficientdet,edgenext,efficientformer,efficientnet,eva,fasternet,fastervit,fastvit,flexivit,gcvit,ghostnet,gpvit,hornet,hiera,iformer,inceptionnext,lcnet,levit,maxvit,mobilevit,moganet,nat,nfnets,pvt,swin,tinynet,tinyvit,uniformer,volo,vanillanet,yolor,yolov7,yolov8,yolox,gpt2,llama2, alias kecam
★ 626+0Star change over the last 7 days - #26
Julia Implementation of Transformer models
★ 570-1Star change over the last 7 days - #27
The official PyTorch implementation of the paper "SAITS: Self-Attention-based Imputation for Time Series". A fast and state-of-the-art (SOTA) deep-learning neural network model for efficient time-series imputation (impute multivariate incomplete time series containing NaN missing data/values with machine learning). https://arxiv.org/abs/2202.08516
★ 513+1Star change over the last 7 days - #28
Design hardware-friendly model architectures and migrate existing LLMs with minimal performance loss
★ 502+1Star change over the last 7 days