Skip to main content
buildradar
Sign in
Topic · multimodal

multimodal

Tracked open-source repos tagged multimodal, sorted by stars.

148 repos
  • stability-sdk@Stability-AI

    SDK for interacting with stability.ai APIs (e.g. stable diffusion inference)

    2,435-1Star change over the last 7 days
  • mPLUG-DocOwl@X-PLUG

    mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

    2,411+0Star change over the last 7 days
  • InternVideo@OpenGVLab

    [ECCV2024] Video Foundation Models & Data for Multimodal Understanding

    2,374+5Star change over the last 7 days
  • DataDesigner@NVIDIA-NeMo

    🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.

    2,197+7Star change over the last 7 days
  • genai-processors@google-gemini

    GenAI Processors is a lightweight Python library that enables efficient, parallel content processing.

    2,120+0Star change over the last 7 days
  • claude-real-video@HUANGCHIHHUNGLeo

    Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.

    2,109+20Star change over the last 7 days
  • parlor@fikrikarim

    On-device, real-time multimodal AI with features similar to GPT-Live

    2,048+7Star change over the last 7 days
  • Show-o@showlab

    [ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.

    1,975+1Star change over the last 7 days
  • An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.

    1,961+1Star change over the last 7 days
  • BitNet@kyegomez

    Implementation of "BitNet: Scaling 1-bit Transformers for Large Language Models" in pytorch

    1,946+0Star change over the last 7 days
  • AdvancedLiterateMachinery@AlibabaResearch

    A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.

    1,835+1Star change over the last 7 days
  • DeTikZify@potamides

    Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ.

    1,818+1Star change over the last 7 days
  • The Self-Coding System for Your App — Alan AI SDK for Android

    1,807-1Star change over the last 7 days
  • The Self-Coding System for Your App — Alan AI SDK for Flutter

    1,756-1Star change over the last 7 days
  • alan-sdk-ionic@alan-ai

    The Self-Coding System for Your App — Alan AI SDK for Ionic

    1,649-1Star change over the last 7 days
  • pixeltable@pixeltable

    Unified multimodal backend for AI data apps

    1,619+5Star change over the last 7 days
  • mllm@UbiquitousLearning

    Fast Multimodal LLM on Mobile Devices

    1,602+3Star change over the last 7 days
  • thepipe@emcf

    Get clean data from tricky documents, powered by vision-language models ⚡

    1,524+0Star change over the last 7 days
  • Ovis@ATH-MaaS

    A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.

    1,514+2Star change over the last 7 days
  • ha-llmvision@valentinfrlch

    Visual intelligence for your home.

    1,454+10Star change over the last 7 days
  • awesome-japanese-llm@llm-jp

    日本語LLMまとめ - Overview of Japanese LLMs

    1,428+1Star change over the last 7 days
  • WeMM-Embedding@Tencent

    WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.

    1,421+521Star change over the last 7 days
  • SDT@dailenson

    This repository is the official implementation of Disentangling Writer and Character Styles for Handwriting Generation (CVPR 2023)

    1,406-1Star change over the last 7 days
  • airunner@Capsize-Games

    Offline inference engine for art, real-time voice conversations, LLM powered chatbots and automated workflows

    1,314+0Star change over the last 7 days
  • Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.

    1,310+2Star change over the last 7 days
  • claude-video-vision@jordanrendric

    Give Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysis

    1,281+6Star change over the last 7 days
  • UForm@unum-cloud

    Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️

    1,247+1Star change over the last 7 days
  • xtreme1@xtreme1-io

    Xtreme1 is an all-in-one data labeling and annotation platform for multimodal data training and supports 3D LiDAR point cloud, image, and LLM.

    1,239+2Star change over the last 7 days
  • LLaMA-Mesh@nv-tlabs

    Unifying 3D Mesh Generation with Language Models

    1,167+0Star change over the last 7 days
  • doc7@magicrew

    Turn documents into AI-ready Markdown with visual understanding

    1,153-25Star change over the last 7 days
← Back to topics