multimodal-large-language-models
Tracked open-source repos tagged multimodal-large-language-models, sorted by stars.
Related topics
Topics that frequently appear alongside multimodal-large-language-models on the same repo.
Recent risers
Repos created in the last 90 days, tagged multimodal-large-language-models.
No new repos tagged with this topic in the last 90 days.
- #1
:sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 17,996+2Star change over the last 7 days - #2
Mobile-Agent: The Powerful GUI Agent Family
★ 9,157+9Star change over the last 7 days - #3
StarVector is a foundation model for SVG generation that transforms vectorization into a code generation task. Using a vision-language modeling architecture, StarVector processes both visual and textual inputs to produce high-quality SVG code with remarkable precision.
★ 4,567+6Star change over the last 7 days - #4
LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.
★ 3,147+1Star change over the last 7 days - #5
A cross-platform video structuring (video analysis) framework based on CV models & mLLM.
★ 2,933+0Star change over the last 7 days - #6
✨✨[NeurIPS 2025] VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
★ 2,532-1Star change over the last 7 days - #7
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
★ 2,411+0Star change over the last 7 days - #8★ 2,014+1Star change over the last 7 days
- #9
[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)
★ 1,844+0Star change over the last 7 days - #10
Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving state-of-the-art performance on 38 out of 60 public benchmarks.
★ 1,586+0Star change over the last 7 days - #11
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
★ 1,514+2Star change over the last 7 days - #12
Awesome Unified Multimodal Models
★ 1,314+1Star change over the last 7 days - #13
实时交互数字人,可自定义形象与音色,支持音色克隆,对话延迟低至3s。Real-time voice interactive digital human, customizable appearance and voice, supporting voice cloning, with initial package delay as low as 3s.
★ 1,303-1Star change over the last 7 days - #14
PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models
★ 1,184+2Star change over the last 7 days - #15
A Framework for Speech, Language, Audio, Music Processing with Large Language Model
★ 1,058+2Star change over the last 7 days - #16★ 1,053+0Star change over the last 7 days
- #17
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
★ 1,025+2Star change over the last 7 days - #18
A collection of resources on applications of multi-modal learning in medical imaging.
★ 975+1Star change over the last 7 days - #19★ 889+0Star change over the last 7 days
- #20
✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
★ 791+2Star change over the last 7 days - #21
LLaVA-Plus: Large Language and Vision Assistants that Plug and Learn to Use Skills
★ 770+0Star change over the last 7 days - #22
[CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
★ 706+0Star change over the last 7 days - #23★ 702+0Star change over the last 7 days
- #24
Personal Project: MPP-Qwen14B & MPP-Qwen-Next(Multimodal Pipeline Parallel based on Qwen-LM). Support [video/image/multi-image] {sft/conversations}. Don't let the poverty limit your imagination! Train your own 8B/14B LLaVA-training-like MLLM on RTX3090/4090 24GB.
★ 687+0Star change over the last 7 days - #25
OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.
★ 678+0Star change over the last 7 days - #26
✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models
★ 650+1Star change over the last 7 days - #27
(Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generators
★ 640+0Star change over the last 7 days - #28
NeurIPS 2024 Paper: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
★ 577+1Star change over the last 7 days - #29
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images, high-resolution images, and videos in an efficient manner.
★ 577+0Star change over the last 7 days - #30
Cambrian-S: Towards Spatial Supersensing in Video
★ 567-1Star change over the last 7 days