multimodal
Tracked open-source repos tagged multimodal, sorted by stars.
Related topics
Topics that frequently appear alongside multimodal on the same repo.
Recent risers
Repos created in the last 90 days, tagged multimodal.
- #1
Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.
★ 2,109 - #2
WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.
★ 1,363 - #3★ 1,153
- #4
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
★ 1,136 - #5
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
★ 1,085
- #1
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience
★ 65,534+163Star change over the last 7 days - #2
《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码
★ 44,381+1,022Star change over the last 7 days - #3
The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
★ 38,817+75Star change over the last 7 days - #4
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25,014+9Star change over the last 7 days - #5★ 22,203+7Star change over the last 7 days
- #6★ 21,859-2Star change over the last 7 days
- #7
YC (S26) | Open Computer History | Record your screen continuously locally and provide context to your agents (Claude, Codex, Openclaw, Hermes, Runner...)
★ 21,376+80Star change over the last 7 days - #8★ 17,763+1Star change over the last 7 days
- #9
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
★ 15,488+88Star change over the last 7 days - #10★ 11,387+15Star change over the last 7 days
- #11
🔍大模型应用开发实战一:RAG 技术全栈指南,在线阅读地址:https://datawhalechina.github.io/all-in-rag/
★ 10,770+94Star change over the last 7 days - #12
Production ready toolkit to run AI locally
★ 10,281+1Star change over the last 7 days - #13
A Python library for anomaly detection across tabular, time series, graph, text, image, and audio data. 60+ detectors, benchmark-backed ADEngine orchestration, and an agentic workflow for AI agents.
★ 9,978+1Star change over the last 7 days - #14
https://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
★ 9,848+60Star change over the last 7 days - #15
SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool.
★ 9,612+12Star change over the last 7 days - #16
Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training.
★ 9,229-1Star change over the last 7 days - #17
Mobile-Agent: The Powerful GUI Agent Family
★ 9,157+9Star change over the last 7 days - #18
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
★ 8,817+4Star change over the last 7 days - #19
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
★ 7,826+23Star change over the last 7 days - #20★ 6,593+136Star change over the last 7 days
- #21
This repository is a curated collection of links to various courses and resources about Artificial Intelligence (AI)
★ 6,482+3Star change over the last 7 days - #22
Open-source framework for building agentic apps in JavaScript, Go, Dart, and Python, built and used in production by Google
★ 6,394+10Star change over the last 7 days - #23
notes for software engineers getting up to speed on new AI developments. Serves as datastore for https://latent.space writing, and product brainstorming, but has cleaned up canonical references under the /Resources folder.
★ 6,247+0Star change over the last 7 days - #24★ 6,018+4Star change over the last 7 days
- #25★ 5,779-1Star change over the last 7 days
- #26
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
★ 5,736+4Star change over the last 7 days - #27★ 5,678+4Star change over the last 7 days
- #28
A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)
★ 5,633-1Star change over the last 7 days - #29★ 5,188+3Star change over the last 7 days
- #30
Align Anything: Training All-modality Model with Feedback
★ 4,671+3Star change over the last 7 days