Skip to main content
buildradar
Sign in
Topic · multimodal

multimodal

Tracked open-source repos tagged multimodal, sorted by stars.

148 repos
  • 为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

    1,136+18Star change over the last 7 days
  • The Self-Coding System for Your App — Alan AI SDK for Cordova

    1,132+0Star change over the last 7 days
  • sglang-omni@sgl-project

    SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.

    1,118+143Star change over the last 7 days
  • MOVA@OpenMOSS

    A foundation model that generates synchronized video and audio in a single model

    1,110+5Star change over the last 7 days
  • dsh-vision-router@ysr666

    Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.

    1,089+68Star change over the last 7 days
  • Aria@rhymes-ai

    Codebase for Aria - an Open Multimodal Native MoE

    1,088+1Star change over the last 7 days
  • XrayGLM@WangRongsheng

    🩺 首个会看胸部X光片的中文多模态医学大模型 | The first Chinese Medical Multimodal Model that Chest Radiographs Summarization.

    1,084+2Star change over the last 7 days
  • VisCPM@OpenBMB

    [ICLR'24 spotlight] Chinese and English Multimodal Large Model Series (Chat and Paint) | 基于CPM基础模型的中英双语多模态大模型系列

    1,061-1Star change over the last 7 days
  • ONE-PEACE@OFA-Sys

    A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

    1,060+0Star change over the last 7 days
  • PointLLM@InternRobotics

    [ECCV 2024 Best Paper Candidate & TPAMI 2025] PointLLM: Empowering Large Language Models to Understand Point Clouds

    1,054+3Star change over the last 7 days
  • CLIP4Clip@ArrowLuo

    An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"

    1,032+1Star change over the last 7 days
  • Awesome-MCoT@yaotingwangofficial

    Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

    1,025+2Star change over the last 7 days
  • tongflow@tong-io

    TongFlow — Multimodal GenAI Studio

    1,017+38Star change over the last 7 days
  • vectordb-recipes@lancedb

    Resource, examples & tutorials for multimodal AI, RAG and agents using vector search and LLMs

    973-1Star change over the last 7 days
  • verl-omni@verl-project

    Multimodal RL training framework for diffusion & omni models

    968+73Star change over the last 7 days
  • rag-time@microsoft

    RAG Time: A 5-week Learning Journey to Mastering RAG

    902+3Star change over the last 7 days
  • About This repository is a curated collection of the most exciting and influential CVPR 2025 papers. 🔥 [Paper + Code + Demo]

    895+1Star change over the last 7 days
  • NEO@EvolvingLMMs-Lab

    NEO Series: Native Vision-Language Models from First Principles

    889+1Star change over the last 7 days
  • AnyGPT@OpenMOSS

    A unified multimodal language model based on discrete sequence modeling

    881-1Star change over the last 7 days
  • ohmycaptcha@shenhao-stu

    ⚡ Self-hostable YesCaptcha-compatible captcha solver built with FastAPI, Playwright, and OpenAI-compatible multimodal models.

    869+4Star change over the last 7 days
  • lmms-engine@EvolvingLMMs-Lab

    A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.

    825+1Star change over the last 7 days
  • [TMLR 2025🔥] A survey for the autoregressive models in vision.

    808+1Star change over the last 7 days
  • papermage@allenai

    library supporting NLP and CV research on scientific papers

    803+0Star change over the last 7 days
  • contrastors@nomic-ai

    Train Models Contrastively in Pytorch

    802+0Star change over the last 7 days
  • 本地监控 + AI 视觉 — LAN-based smartphone-powered AI monitoring framework with structured event output for data acquisition and analysis.

    770+12Star change over the last 7 days
  • Paper list about multimodal and large language models, only used to record papers I read in the daily arxiv for personal needs.

    761+1Star change over the last 7 days
  • InstructIR@mv-lab

    [ECCV 2024] InstructIR: High-Quality Image Restoration Following Human Instructions https://huggingface.co/spaces/marcosv/InstructIR

    741+0Star change over the last 7 days
  • PaddleMIX@PaddlePaddle

    Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high performance and flexibility.

    723-1Star change over the last 7 days
  • MMRec@enoche

    A Toolbox for MultiModal Recommendation. Integrating 10+ Models...

    689+1Star change over the last 7 days
  • VLM2Vec@TIGER-AI-Lab

    This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]

    682+2Star change over the last 7 days
← Back to topics