multimodal-ai
Tracked open-source repos tagged multimodal-ai, sorted by stars.
Related topics
Topics that frequently appear alongside multimodal-ai on the same repo.
Recent risers
Repos created in the last 90 days, tagged multimodal-ai.
No new repos tagged with this topic in the last 90 days.
- #1
🚀 Truly open-source AI avatar(digital human) toolkit for offline video generation and digital human cloning.
★ 14,992+65Star change over the last 7 days - #2
Official Agnes AI gateway and model catalog for OpenAI-compatible text, image, video, and agent workflows.
★ 5,044+28Star change over the last 7 days - #3
Multi-modal Generative Media Skills for AI Agents (Claude Code, Cursor, Gemini CLI). High-quality image, video, and audio generation powered by muapi.ai.
★ 4,205+19Star change over the last 7 days - #4
Mano-P: Open-source GUI-VLA agent for edge devices. #1 on OSWorld (specialized, 58.2%). Runs locally on Apple M4 Mac mini/MacBook — no data leaves your device.Mano-P 是一个开源 GUI-VLA 项目,支持在 Mac mini/MacBook 上或通过算力棒本地运行推理,实现纯视觉驱动的跨平台 GUI 自动化操作。数据完全本地处理,支持复杂多步骤任务规划与执行。
★ 2,615+15Star change over the last 7 days - #5
EVA OS — A real-time multimodal AIOS for next-generation hardware, enabling your devices being “alive” and as intelligent as a real brain.
★ 1,869+87Star change over the last 7 days - #6
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.
★ 1,840+10Star change over the last 7 days - #7
OpenGUI is an Android GUI agent framework for phone-use AI that can see, plan, and operate real mobile apps through the GUI.
★ 1,624+38Star change over the last 7 days - #8
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
★ 1,557+6Star change over the last 7 days - #9
Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.
★ 1,310+2Star change over the last 7 days - #10
小遥搜索,听懂你的话、看懂你的图,用AI找到本地任何文件。让搜索像聊天一样简单。XiaoyaoSearch: Understands your words, reads your images, finds any local file with AI. Making search as easy as chatting.
★ 1,084+1Star change over the last 7 days - #11
Your Open Autonomous Android Agent — A production-ready, self-planning AI assistant powered by local/remote LLMs and accessibility-driven screen automation.
★ 1,022+44Star change over the last 7 days - #12
Resource, examples & tutorials for multimodal AI, RAG and agents using vector search and LLMs
★ 973-1Star change over the last 7 days - #13★ 596—Star change over the last 7 days