Skip to main content
buildradar
Sign in
Topic · vision-language-model

vision-language-model

Tracked open-source repos tagged vision-language-model, sorted by stars.

Repos
59
Total stars
150,825
Avg. stars
2,556
Share
0.01%

Topics that frequently appear alongside vision-language-model on the same repo.

Recent risers

Repos created in the last 90 days, tagged vision-language-model.

  • doc7@magicrew

    Turn documents into AI-ready Markdown with visual understanding

    1,153
  • 为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

    1,136
  • dsh-vision-toolkit@Anionex

    [dsh]为纯文本模型设计更强大的视觉工具箱:一行安装使用、粘贴图片直接识别、多张图片问答、截图到前端UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.

    854
  • OraRL@HVision-NKU

    🎬 OraRL — Annotations as Rollouts for efficient, scalable reinforcement learning of unified video MLLMs.

    152
  • LLaVA@haotian-liu

    [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.

    25,014+9Star change over the last 7 days
  • X-AnyLabeling@CVHub520

    X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.

    10,313+59Star change over the last 7 days
  • InternVL@OpenGVLab

    [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型

    10,150+3Star change over the last 7 days
  • minimind-v@jingyaogong

    👀 Train a 65M-parameter VLM from scratch in just 2h!

    8,535+40Star change over the last 7 days
  • Qwen-VL@QwenLM

    The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed by Alibaba Cloud.

    6,727+1Star change over the last 7 days
  • MineContext@volcengine

    MineContext is your proactive context-aware AI partner(Context-Engineering+ChatGPT Pulse)

    5,497+1Star change over the last 7 days
  • mlx-vlm@Blaizzy

    MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.

    5,462+25Star change over the last 7 days
  • align-anything@PKU-Alignment

    Align Anything: Training All-modality Model with Feedback

    4,671+3Star change over the last 7 days
  • lmms-eval@EvolvingLMMs-Lab

    One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

    4,390+7Star change over the last 7 days
  • DeepSeek-VL@deepseek-ai

    DeepSeek-VL: Towards Real-World Vision-Language Understanding

    4,177+2Star change over the last 7 days
  • MiniMax-01@MiniMax-AI

    The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on Linear Attention

    3,467+0Star change over the last 7 days
  • MGM@JIA-Lab-research

    Official repo for "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models"

    3,327+0Star change over the last 7 days
  • VLM_survey@jingyi0000

    Collection of AWESOME vision-language models for vision tasks

    3,126+0Star change over the last 7 days
  • OGAM@off-grid-ai

    The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers. Runs on your CPU, GPU, or NPU. No account, no API key, zero data leaves your device.

    3,038+17Star change over the last 7 days
  • InternLM-XComposer@InternLM

    InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

    2,925-1Star change over the last 7 days
  • colpali@illuin-tech

    The code used to train and run inference with the ColVision models, e.g. ColPali, ColQwen2, and ColSmol.

    2,809+7Star change over the last 7 days
  • Cradle@BAAI-Agents

    The Cradle framework is a first attempt at General Computer Control (GCC). Cradle supports agents to ace any computer task by enabling strong reasoning abilities, self-improvment, and skill curation, in a standardized general environment with minimal requirements.

    2,578+2Star change over the last 7 days
  • An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.

    1,961+1Star change over the last 7 days
  • ShowUI@showlab

    [CVPR 2025] Open-source, End-to-end, Vision-Language-Action model for GUI Agent & Computer Use.

    1,896+2Star change over the last 7 days
  • Awesome-LLM4AD@Thinklab-SJTU

    A curated list of awesome LLM/VLM/VLA/World Model for Autonomous Driving(LLM4AD) resources (continually updated)

    1,890-2Star change over the last 7 days
  • video-search-and-summarization@NVIDIA-AI-Blueprints

    NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.

    1,840+10Star change over the last 7 days
  • AdvancedLiterateMachinery@AlibabaResearch

    A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.

    1,835+1Star change over the last 7 days
  • Seed1.5-VL@ByteDance-Seed

    Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving state-of-the-art performance on 38 out of 60 public benchmarks.

    1,586+0Star change over the last 7 days
  • vllm-mlx@waybarrios

    High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

    1,557+6Star change over the last 7 days
  • thepipe@emcf

    Get clean data from tricky documents, powered by vision-language models ⚡

    1,524+0Star change over the last 7 days
  • [ICCV 2025] Implementation for Describe Anything: Detailed Localized Image and Video Captioning

    1,517+3Star change over the last 7 days
  • Ovis@ATH-MaaS

    A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.

    1,514+2Star change over the last 7 days
  • awesome-japanese-llm@llm-jp

    日本語LLMまとめ - Overview of Japanese LLMs

    1,428+1Star change over the last 7 days
  • mlx-tune@ARahim3

    Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.

    1,398+8Star change over the last 7 days
  • Awesome Unified Multimodal Models

    1,314+1Star change over the last 7 days
← Back to topics