multimodal
Tracked open-source repos tagged multimodal, sorted by stars.
- #61
SDK for interacting with stability.ai APIs (e.g. stable diffusion inference)
★ 2,435-1Star change over the last 7 days - #62
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
★ 2,411+0Star change over the last 7 days - #63
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
★ 2,374+5Star change over the last 7 days - #64
🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.
★ 2,197+7Star change over the last 7 days - #65
GenAI Processors is a lightweight Python library that enables efficient, parallel content processing.
★ 2,120+0Star change over the last 7 days - #66
Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.
★ 2,109+20Star change over the last 7 days - #67★ 2,048+7Star change over the last 7 days
- #68
[ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.
★ 1,975+1Star change over the last 7 days - #69
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
★ 1,961+1Star change over the last 7 days - #70
Implementation of "BitNet: Scaling 1-bit Transformers for Large Language Models" in pytorch
★ 1,946+0Star change over the last 7 days - #71
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.
★ 1,835+1Star change over the last 7 days - #72★ 1,818+1Star change over the last 7 days
- #73
The Self-Coding System for Your App — Alan AI SDK for Android
★ 1,807-1Star change over the last 7 days - #74
The Self-Coding System for Your App — Alan AI SDK for Flutter
★ 1,756-1Star change over the last 7 days - #75
The Self-Coding System for Your App — Alan AI SDK for Ionic
★ 1,649-1Star change over the last 7 days - #76
Unified multimodal backend for AI data apps
★ 1,619+5Star change over the last 7 days - #77★ 1,602+3Star change over the last 7 days
- #78★ 1,524+0Star change over the last 7 days
- #79
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
★ 1,514+2Star change over the last 7 days - #80
Visual intelligence for your home.
★ 1,454+10Star change over the last 7 days - #81
日本語LLMまとめ - Overview of Japanese LLMs
★ 1,428+1Star change over the last 7 days - #82
WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.
★ 1,421+521Star change over the last 7 days - #83
This repository is the official implementation of Disentangling Writer and Character Styles for Handwriting Generation (CVPR 2023)
★ 1,406-1Star change over the last 7 days - #84
Offline inference engine for art, real-time voice conversations, LLM powered chatbots and automated workflows
★ 1,314+0Star change over the last 7 days - #85
Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.
★ 1,310+2Star change over the last 7 days - #86
Give Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysis
★ 1,281+6Star change over the last 7 days - #87
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
★ 1,247+1Star change over the last 7 days - #88
Xtreme1 is an all-in-one data labeling and annotation platform for multimodal data training and supports 3D LiDAR point cloud, image, and LLM.
★ 1,239+2Star change over the last 7 days - #89
Unifying 3D Mesh Generation with Language Models
★ 1,167+0Star change over the last 7 days - #90★ 1,153-25Star change over the last 7 days