Skip to main content
buildradar
Sign in
Topic · vision-transformer

vision-transformer

Tracked open-source repos tagged vision-transformer, sorted by stars.

Repos
45
Total stars
154,486
Avg. stars
3,433
Share
0.01%

Topics that frequently appear alongside vision-transformer on the same repo.

Recent risers

Repos created in the last 90 days, tagged vision-transformer.

No new repos tagged with this topic in the last 90 days.

  • mmdetection@open-mmlab

    OpenMMLab Detection Toolbox and Benchmark

    32,901+6Star change over the last 7 days
  • LaTeX-OCR@lukas-blecher

    pix2tex: Using a ViT to convert images of equations into LaTeX code.

    16,549+0Star change over the last 7 days
  • Transformers-Tutorials@NielsRogge

    This repository contains demos I made with the Transformers library by HuggingFace.

    11,748+0Star change over the last 7 days
  • VAR@FoundationVision

    [NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!

    8,728-1Star change over the last 7 days
  • omniparse@adithya-s-k

    Ingest, parse, and optimize any data format ➡️ from documents to multimedia ➡️ for enhanced compatibility with GenAI frameworks

    7,823+4Star change over the last 7 days
  • SwinIR@JingyunLiang

    SwinIR: Image Restoration Using Swin Transformer (official repository)

    5,586+6Star change over the last 7 days
  • mlx-vlm@Blaizzy

    MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.

    5,462+25Star change over the last 7 days
  • An ultimately comprehensive paper list of Vision Transformer/Attention, including papers, codes, and related websites

    5,046-3Star change over the last 7 days
  • Efficient-AI-Backbones@huawei-noah

    Efficient AI Backbones including GhostNet, TNT and MLP, developed by Huawei Noah's Ark Lab.

    4,419+1Star change over the last 7 days
  • mmpretrain@open-mmlab

    OpenMMLab Pre-training Toolbox and Benchmark

    3,850-1Star change over the last 7 days
  • modlens@liustack

    The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。

    3,839+83Star change over the last 7 days
  • scenic@google-research

    Scenic: A Jax Library for Computer Vision Research and Beyond

    3,821+0Star change over the last 7 days
  • towhee@towhee-io

    Towhee is a framework that is dedicated to making neural data processing pipelines simple and fast.

    3,453-1Star change over the last 7 days
  • efficientvit@mit-han-lab

    Efficient vision foundation models for high-resolution generation and perception.

    3,356+1Star change over the last 7 days
  • InternLM-XComposer@InternLM

    InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

    2,925-1Star change over the last 7 days
  • EVA@baaivision

    EVA Series: Visual Representation Fantasies from BAAI

    2,695+2Star change over the last 7 days
  • InternVideo@OpenGVLab

    [ECCV2024] Video Foundation Models & Data for Multimodal Understanding

    2,374+5Star change over the last 7 days
  • MambaVision@NVlabs

    [CVPR 2025] Official PyTorch Implementation of MambaVision: A Hybrid Mamba-Transformer Vision Backbone

    2,227+2Star change over the last 7 days
  • ViTPose@ViTAE-Transformer

    The official repo for [NeurIPS'22] "ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation" and [TPAMI'23] "ViTPose++: Vision Transformer for Generic Body Pose Estimation"

    2,139+0Star change over the last 7 days
  • Transformer-Explainability@hila-chefer

    [CVPR 2021] Official PyTorch implementation for Transformer Interpretability Beyond Attention Visualization, a novel method to visualize classifications by Transformer based networks.

    2,014+0Star change over the last 7 days
  • EasyCV@alibaba

    An all-in-one toolkit for computer vision

    1,955+1Star change over the last 7 days
  • Cream@microsoft

    This is a collection of our NAS and Vision Transformer work.

    1,841+2Star change over the last 7 days
  • lightly-train@lightly-ai

    All-in-one training for vision models (YOLO, ViTs, RT-DETR, DINOv3): pretraining, fine-tuning, distillation.

    1,656+4Star change over the last 7 days
  • ViT-Adapter@czczup

    [ICLR 2023 Spotlight] Vision Transformer Adapter for Dense Predictions

    1,503+1Star change over the last 7 days
  • A curated list of foundation models for vision and language tasks

    1,177+0Star change over the last 7 days
  • GeoSeg@WangLibo1995

    UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery, ISPRS. Also, including other vision transformers and CNNs for satellite, aerial image and UAV image segmentation.

    1,101+3Star change over the last 7 days
  • ONE-PEACE@OFA-Sys

    A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

    1,060+0Star change over the last 7 days
  • SpargeAttn@thu-ml

    [ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.

    1,045+3Star change over the last 7 days
  • :fire: :fire: :fire: A paper list of some recent Computer Vision(CV) works

    973+1Star change over the last 7 days
  • DAT@LeapLabTHU

    Repository of Vision Transformer with Deformable Attention (CVPR2022) and DAT++: Spatially Dynamic Vision Transformerwith Deformable Attention

    942+0Star change over the last 7 days
← Back to topics