Skip to main content
buildradar
Sign in
Topic · vision-language-model

vision-language-model

Tracked open-source repos tagged vision-language-model, sorted by stars.

59 repos
  • prismer@NVlabs

    The implementation of "Prismer: A Vision-Language Model with Multi-Task Experts".

    1,309+0Star change over the last 7 days
  • LLaVA-OneVision-2@EvolvingLMMs-Lab

    Fully Open Framework for Democratized Multimodal Training

    1,195+1Star change over the last 7 days
  • vlms-zero-to-hero@SkalskiP

    This series will take you on a journey from the fundamentals of NLP and Computer Vision to the cutting edge of Vision-Language Models.

    1,179+0Star change over the last 7 days
  • doc7@magicrew

    Turn documents into AI-ready Markdown with visual understanding

    1,153-25Star change over the last 7 days
  • 为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

    1,136+18Star change over the last 7 days
  • VisRAG@OpenBMB

    Parsing-free RAG supported by VLMs

    979-1Star change over the last 7 days
  • groundingLMM@mbzuai-oryx

    [CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with object segmentation masks.

    967+0Star change over the last 7 days
  • Chat-UniVi@PKU-YuanGroup

    [CVPR 2024 Highlight🔥] Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

    941+0Star change over the last 7 days
  • About This repository is a curated collection of the most exciting and influential CVPR 2025 papers. 🔥 [Paper + Code + Demo]

    895+1Star change over the last 7 days
  • AlphaCLIP@SunzeY

    [CVPR 2024] Alpha-CLIP: A CLIP Model Focusing on Wherever You Want

    874+0Star change over the last 7 days
  • dsh-vision-toolkit@Anionex

    [dsh]为纯文本模型设计更强大的视觉工具箱:一行安装使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.

    856+17Star change over the last 7 days
  • VoxPoser@huangwl18

    VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

    834+1Star change over the last 7 days
  • OpenCUA@xlang-ai

    [NeurIPS 2025 Spotlight] OpenCUA: Open Foundations for Computer-Use Agents

    833+4Star change over the last 7 days
  • MINT-1T@mlfoundations

    🍃 MINT-1T: A one trillion token multimodal interleaved dataset.

    832+0Star change over the last 7 days
  • A curated list of 3D Vision papers relating to Robotics domain in the era of large models i.e. LLMs/VLMs, inspired by awesome-computer-vision, including papers, codes, and related websites

    821+2Star change over the last 7 days
  • A curated list of awesome prompt/adapter learning methods for vision-language models like CLIP.

    797+1Star change over the last 7 days
  • X-VLA@2toinf

    [ICLR 2026] The offical Implementation of "Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model"

    725+5Star change over the last 7 days
  • OmniVinci@NVlabs

    OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.

    677+0Star change over the last 7 days
  • MiMo-VL@XiaomiMiMo

    MiMo-VL

    642+0Star change over the last 7 days
  • SpatialVID@NJU-3DV

    [CVPR 2026] SpatialVID: A Large-Scale Video Dataset with Spatial Annotations

    600+1Star change over the last 7 days
  • LAMDA-PILOT@LAMDA-CL

    🎉 PILOT: A Pre-trained Model-Based Continual Learning Toolbox

    600+3Star change over the last 7 days
  • t2v_metrics@linzhiqiu

    Evaluating text-to-image/video/3D models with VQAScore

    600+0Star change over the last 7 days
  • MOSS-VL@OpenMOSS

    An open-weight 11B model series for long-form and real-time video understanding

    596Star change over the last 7 days
  • Groma@FoundationVision

    [ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization

    586+0Star change over the last 7 days
  • LLaVA-Mini@ictnlp

    LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images, high-resolution images, and videos in an efficient manner.

    577+0Star change over the last 7 days
  • cambrian-s@cambrian-mllm

    Cambrian-S: Towards Spatial Supersensing in Video

    567-1Star change over the last 7 days
  • Multi-Modality-Arena@OpenGVLab

    Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4, LLaMA-Adapter V2, LLaVA, BLIP-2, and many more!

    566+0Star change over the last 7 days
  • Flame-Code-VLM@Flame-Code-VLM

    Flame is an open-source multimodal AI system designed to translate UI design mockups into high-quality React code. It leverages vision-language modeling, automated data synthesis, and structured training workflows to bridge the gap between design and front-end development.

    562+0Star change over the last 7 days
  • count-anything@Mengqi-Lei

    Code and implementation guidelines for the paper ✨Counting Anything. Project Page: https://mengqi-lei.github.io/count-anything-projectpage/

    539+2Star change over the last 7 days
← Back to topics