vision-language-model
Tracked open-source repos tagged vision-language-model, sorted by stars.
Related topics
Topics that frequently appear alongside vision-language-model on the same repo.
Recent risers
Repos created in the last 90 days, tagged vision-language-model.
- #1★ 1,153
- #2
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
★ 1,136 - #3
[dsh]为纯文本模型设计更强大的视觉工具箱:一行安装使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.
★ 854 - #4
🎬 OraRL — Annotations as Rollouts for efficient, scalable reinforcement learning of unified video MLLMs.
★ 152
- #1
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25,014+9Star change over the last 7 days - #2
X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.
★ 10,313+59Star change over the last 7 days - #3
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
★ 10,150+3Star change over the last 7 days - #4
👀 Train a 65M-parameter VLM from scratch in just 2h!
★ 8,535+40Star change over the last 7 days - #5
The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed by Alibaba Cloud.
★ 6,727+1Star change over the last 7 days - #6
MineContext is your proactive context-aware AI partner(Context-Engineering+ChatGPT Pulse)
★ 5,497+1Star change over the last 7 days - #7
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
★ 5,462+25Star change over the last 7 days - #8
Align Anything: Training All-modality Model with Feedback
★ 4,671+3Star change over the last 7 days - #9
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4,390+7Star change over the last 7 days - #10
DeepSeek-VL: Towards Real-World Vision-Language Understanding
★ 4,177+2Star change over the last 7 days - #11
The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on Linear Attention
★ 3,467+0Star change over the last 7 days - #12
Official repo for "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models"
★ 3,327+0Star change over the last 7 days - #13
Collection of AWESOME vision-language models for vision tasks
★ 3,126+0Star change over the last 7 days - #14
The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers. Runs on your CPU, GPU, or NPU. No account, no API key, zero data leaves your device.
★ 3,038+17Star change over the last 7 days - #15
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
★ 2,925-1Star change over the last 7 days - #16
The code used to train and run inference with the ColVision models, e.g. ColPali, ColQwen2, and ColSmol.
★ 2,809+7Star change over the last 7 days - #17
The Cradle framework is a first attempt at General Computer Control (GCC). Cradle supports agents to ace any computer task by enabling strong reasoning abilities, self-improvment, and skill curation, in a standardized general environment with minimal requirements.
★ 2,578+2Star change over the last 7 days - #18
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
★ 1,961+1Star change over the last 7 days - #19
[CVPR 2025] Open-source, End-to-end, Vision-Language-Action model for GUI Agent & Computer Use.
★ 1,896+2Star change over the last 7 days - #20
A curated list of awesome LLM/VLM/VLA/World Model for Autonomous Driving(LLM4AD) resources (continually updated)
★ 1,890-2Star change over the last 7 days - #21
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.
★ 1,840+10Star change over the last 7 days - #22
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.
★ 1,835+1Star change over the last 7 days - #23
Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving state-of-the-art performance on 38 out of 60 public benchmarks.
★ 1,586+0Star change over the last 7 days - #24
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
★ 1,557+6Star change over the last 7 days - #25★ 1,524+0Star change over the last 7 days
- #26
[ICCV 2025] Implementation for Describe Anything: Detailed Localized Image and Video Captioning
★ 1,517+3Star change over the last 7 days - #27
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
★ 1,514+2Star change over the last 7 days - #28
日本語LLMまとめ - Overview of Japanese LLMs
★ 1,428+1Star change over the last 7 days - #29
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
★ 1,398+8Star change over the last 7 days - #30
Awesome Unified Multimodal Models
★ 1,314+1Star change over the last 7 days