multimodal-llm
Tracked open-source repos tagged multimodal-llm, sorted by stars.
Related topics
Topics that frequently appear alongside multimodal-llm on the same repo.
Recent risers
Repos created in the last 90 days, tagged multimodal-llm.
- #1
WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.
★ 1,330 - #2
🎬 OraRL — Annotations as Rollouts for efficient, scalable reinforcement learning of unified video MLLMs.
★ 152
- #1
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.
★ 1,975+2Star change over the last 7 days - #2
WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.
★ 1,330+426Star change over the last 7 days - #3
Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.
★ 1,310+2Star change over the last 7 days - #4
Official implementation of paper "MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens"
★ 869+0Star change over the last 7 days - #5
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singing ASR. FireRedVAD supports speech/singing/music in 100+ langs. FireRedLID supports 100+ langs and 20+ zh dialects. FireRedPunc supports zh and en.
★ 669+12Star change over the last 7 days