Skip to main content
buildradar
Sign in
Topic · multimodal-ai

multimodal-ai

Tracked open-source repos tagged multimodal-ai, sorted by stars.

Repos
13
Total stars
38,644
Avg. stars
2,973
Share
0.00%

Topics that frequently appear alongside multimodal-ai on the same repo.

Recent risers

Repos created in the last 90 days, tagged multimodal-ai.

No new repos tagged with this topic in the last 90 days.

  • 🚀 Truly open-source AI avatar(digital human) toolkit for offline video generation and digital human cloning.

    14,992+65Star change over the last 7 days
  • AgnesAI-Models@AgnesAI-Labs

    Official Agnes AI gateway and model catalog for OpenAI-compatible text, image, video, and agent workflows.

    5,044+28Star change over the last 7 days
  • Multi-modal Generative Media Skills for AI Agents (Claude Code, Cursor, Gemini CLI). High-quality image, video, and audio generation powered by muapi.ai.

    4,205+19Star change over the last 7 days
  • Mano-P@Mininglamp-AI

    Mano-P: Open-source GUI-VLA agent for edge devices. #1 on OSWorld (specialized, 58.2%). Runs locally on Apple M4 Mac mini/MacBook — no data leaves your device.Mano-P 是一个开源 GUI-VLA 项目,支持在 Mac mini/MacBook 上或通过算力棒本地运行推理,实现纯视觉驱动的跨平台 GUI 自动化操作。数据完全本地处理,支持复杂多步骤任务规划与执行。

    2,615+15Star change over the last 7 days
  • EVA-OS@AutoArk

    EVA OS — A real-time multimodal AIOS for next-generation hardware, enabling your devices being “alive” and as intelligent as a real brain.

    1,869+87Star change over the last 7 days
  • video-search-and-summarization@NVIDIA-AI-Blueprints

    NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.

    1,840+10Star change over the last 7 days
  • OpenGUI@Core-Mate

    OpenGUI is an Android GUI agent framework for phone-use AI that can see, plan, and operate real mobile apps through the GUI.

    1,624+38Star change over the last 7 days
  • vllm-mlx@waybarrios

    High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

    1,557+6Star change over the last 7 days
  • Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.

    1,310+2Star change over the last 7 days
  • xiaoyaosearch@dtsola

    小遥搜索,听懂你的话、看懂你的图,用AI找到本地任何文件。让搜索像聊天一样简单。XiaoyaoSearch: Understands your words, reads your images, finds any local file with AI. Making search as easy as chatting.

    1,084+1Star change over the last 7 days
  • opendroid@yashab-cyber

    Your Open Autonomous Android Agent — A production-ready, self-planning AI assistant powered by local/remote LLMs and accessibility-driven screen automation.

    1,022+44Star change over the last 7 days
  • vectordb-recipes@lancedb

    Resource, examples & tutorials for multimodal AI, RAG and agents using vector search and LLMs

    973-1Star change over the last 7 days
  • MOSS-VL@OpenMOSS

    An open-weight 11B model series for long-form and real-time video understanding

    596Star change over the last 7 days
← Back to topics