vision-language-model
Tracked open-source repos tagged vision-language-model, sorted by stars.
- #31★ 1,309+0Star change over the last 7 days
- #32
Fully Open Framework for Democratized Multimodal Training
★ 1,195+1Star change over the last 7 days - #33
This series will take you on a journey from the fundamentals of NLP and Computer Vision to the cutting edge of Vision-Language Models.
★ 1,179+0Star change over the last 7 days - #34★ 1,153-25Star change over the last 7 days
- #35
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
★ 1,136+18Star change over the last 7 days - #36★ 979-1Star change over the last 7 days
- #37
[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with object segmentation masks.
★ 967+0Star change over the last 7 days - #38
[CVPR 2024 Highlight🔥] Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
★ 941+0Star change over the last 7 days - #39
About This repository is a curated collection of the most exciting and influential CVPR 2025 papers. 🔥 [Paper + Code + Demo]
★ 895+1Star change over the last 7 days - #40★ 874+0Star change over the last 7 days
- #41
[dsh]为纯文本模型设计更强大的视觉工具箱:一行安装使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.
★ 856+17Star change over the last 7 days - #42★ 834+1Star change over the last 7 days
- #43★ 833+4Star change over the last 7 days
- #44★ 832+0Star change over the last 7 days
- #45
A curated list of 3D Vision papers relating to Robotics domain in the era of large models i.e. LLMs/VLMs, inspired by awesome-computer-vision, including papers, codes, and related websites
★ 821+2Star change over the last 7 days - #46
A curated list of awesome prompt/adapter learning methods for vision-language models like CLIP.
★ 797+1Star change over the last 7 days - #47
[ICLR 2026] The offical Implementation of "Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model"
★ 725+5Star change over the last 7 days - #48
OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.
★ 677+0Star change over the last 7 days - #49★ 642+0Star change over the last 7 days
- #50
[CVPR 2026] SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
★ 600+1Star change over the last 7 days - #51
🎉 PILOT: A Pre-trained Model-Based Continual Learning Toolbox
★ 600+3Star change over the last 7 days - #52
Evaluating text-to-image/video/3D models with VQAScore
★ 600+0Star change over the last 7 days - #53★ 596—Star change over the last 7 days
- #54
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
★ 586+0Star change over the last 7 days - #55
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images, high-resolution images, and videos in an efficient manner.
★ 577+0Star change over the last 7 days - #56
Cambrian-S: Towards Spatial Supersensing in Video
★ 567-1Star change over the last 7 days - #57
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4, LLaMA-Adapter V2, LLaVA, BLIP-2, and many more!
★ 566+0Star change over the last 7 days - #58
Flame is an open-source multimodal AI system designed to translate UI design mockups into high-quality React code. It leverages vision-language modeling, automated data synthesis, and structured training workflows to bridge the gap between design and front-end development.
★ 562+0Star change over the last 7 days - #59
Code and implementation guidelines for the paper ✨Counting Anything. Project Page: https://mengqi-lei.github.io/count-anything-projectpage/
★ 539+2Star change over the last 7 days