multimodal
Tracked open-source repos tagged multimodal, sorted by stars.
- #91
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
★ 1,136+18Star change over the last 7 days - #92
The Self-Coding System for Your App — Alan AI SDK for Cordova
★ 1,132+0Star change over the last 7 days - #93
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
★ 1,118+143Star change over the last 7 days - #94★ 1,110+5Star change over the last 7 days
- #95
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
★ 1,089+68Star change over the last 7 days - #96★ 1,088+1Star change over the last 7 days
- #97
🩺 首个会看胸部X光片的中文多模态医学大模型 | The first Chinese Medical Multimodal Model that Chest Radiographs Summarization.
★ 1,084+2Star change over the last 7 days - #98
[ICLR'24 spotlight] Chinese and English Multimodal Large Model Series (Chat and Paint) | 基于CPM基础模型的中英双语多模态大模型系列
★ 1,061-1Star change over the last 7 days - #99
A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
★ 1,060+0Star change over the last 7 days - #100
[ECCV 2024 Best Paper Candidate & TPAMI 2025] PointLLM: Empowering Large Language Models to Understand Point Clouds
★ 1,054+3Star change over the last 7 days - #101
An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"
★ 1,032+1Star change over the last 7 days - #102
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
★ 1,025+2Star change over the last 7 days - #103★ 1,017+38Star change over the last 7 days
- #104
Resource, examples & tutorials for multimodal AI, RAG and agents using vector search and LLMs
★ 973-1Star change over the last 7 days - #105★ 968+73Star change over the last 7 days
- #106★ 902+3Star change over the last 7 days
- #107
About This repository is a curated collection of the most exciting and influential CVPR 2025 papers. 🔥 [Paper + Code + Demo]
★ 895+1Star change over the last 7 days - #108★ 889+1Star change over the last 7 days
- #109★ 881-1Star change over the last 7 days
- #110
⚡ Self-hostable YesCaptcha-compatible captcha solver built with FastAPI, Playwright, and OpenAI-compatible multimodal models.
★ 869+4Star change over the last 7 days - #111
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
★ 825+1Star change over the last 7 days - #112
[TMLR 2025🔥] A survey for the autoregressive models in vision.
★ 808+1Star change over the last 7 days - #113★ 803+0Star change over the last 7 days
- #114
Train Models Contrastively in Pytorch
★ 802+0Star change over the last 7 days - #115
本地监控 + AI 视觉 — LAN-based smartphone-powered AI monitoring framework with structured event output for data acquisition and analysis.
★ 770+12Star change over the last 7 days - #116
Paper list about multimodal and large language models, only used to record papers I read in the daily arxiv for personal needs.
★ 761+1Star change over the last 7 days - #117
[ECCV 2024] InstructIR: High-Quality Image Restoration Following Human Instructions https://huggingface.co/spaces/marcosv/InstructIR
★ 741+0Star change over the last 7 days - #118
Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high performance and flexibility.
★ 723-1Star change over the last 7 days - #119★ 689+1Star change over the last 7 days
- #120
This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]
★ 682+2Star change over the last 7 days