vision-transformer
Tracked open-source repos tagged vision-transformer, sorted by stars.
Related topics
Topics that frequently appear alongside vision-transformer on the same repo.
Recent risers
Repos created in the last 90 days, tagged vision-transformer.
No new repos tagged with this topic in the last 90 days.
- #1
OpenMMLab Detection Toolbox and Benchmark
★ 32,901+6Star change over the last 7 days - #2★ 16,549+0Star change over the last 7 days
- #3
This repository contains demos I made with the Transformers library by HuggingFace.
★ 11,748+0Star change over the last 7 days - #4
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
★ 8,728-1Star change over the last 7 days - #5
Ingest, parse, and optimize any data format ➡️ from documents to multimedia ➡️ for enhanced compatibility with GenAI frameworks
★ 7,823+4Star change over the last 7 days - #6★ 5,586+6Star change over the last 7 days
- #7
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
★ 5,462+25Star change over the last 7 days - #8
An ultimately comprehensive paper list of Vision Transformer/Attention, including papers, codes, and related websites
★ 5,046-3Star change over the last 7 days - #9
Efficient AI Backbones including GhostNet, TNT and MLP, developed by Huawei Noah's Ark Lab.
★ 4,419+1Star change over the last 7 days - #10
OpenMMLab Pre-training Toolbox and Benchmark
★ 3,850-1Star change over the last 7 days - #11
The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。
★ 3,839+83Star change over the last 7 days - #12★ 3,821+0Star change over the last 7 days
- #13
Towhee is a framework that is dedicated to making neural data processing pipelines simple and fast.
★ 3,453-1Star change over the last 7 days - #14
Efficient vision foundation models for high-resolution generation and perception.
★ 3,356+1Star change over the last 7 days - #15
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
★ 2,925-1Star change over the last 7 days - #16★ 2,695+2Star change over the last 7 days
- #17
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
★ 2,374+5Star change over the last 7 days - #18
[CVPR 2025] Official PyTorch Implementation of MambaVision: A Hybrid Mamba-Transformer Vision Backbone
★ 2,227+2Star change over the last 7 days - #19
The official repo for [NeurIPS'22] "ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation" and [TPAMI'23] "ViTPose++: Vision Transformer for Generic Body Pose Estimation"
★ 2,139+0Star change over the last 7 days - #20
[CVPR 2021] Official PyTorch implementation for Transformer Interpretability Beyond Attention Visualization, a novel method to visualize classifications by Transformer based networks.
★ 2,014+0Star change over the last 7 days - #21★ 1,955+1Star change over the last 7 days
- #22★ 1,841+2Star change over the last 7 days
- #23
All-in-one training for vision models (YOLO, ViTs, RT-DETR, DINOv3): pretraining, fine-tuning, distillation.
★ 1,656+4Star change over the last 7 days - #24
[ICLR 2023 Spotlight] Vision Transformer Adapter for Dense Predictions
★ 1,503+1Star change over the last 7 days - #25
A curated list of foundation models for vision and language tasks
★ 1,177+0Star change over the last 7 days - #26
UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery, ISPRS. Also, including other vision transformers and CNNs for satellite, aerial image and UAV image segmentation.
★ 1,101+3Star change over the last 7 days - #27
A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
★ 1,060+0Star change over the last 7 days - #28
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.
★ 1,045+3Star change over the last 7 days - #29
:fire: :fire: :fire: A paper list of some recent Computer Vision(CV) works
★ 973+1Star change over the last 7 days - #30
Repository of Vision Transformer with Deformable Attention (CVPR2022) and DAT++: Spatially Dynamic Vision Transformerwith Deformable Attention
★ 942+0Star change over the last 7 days