large-multimodal-models
Tracked open-source repos tagged large-multimodal-models, sorted by stars.
Related topics
Topics that frequently appear alongside large-multimodal-models on the same repo.
Recent risers
Repos created in the last 90 days, tagged large-multimodal-models.
No new repos tagged with this topic in the last 90 days.
- #1
✨✨[NeurIPS 2025] VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
★ 2,532-2Star change over the last 7 days - #2
[ICCV 2025] Implementation for Describe Anything: Detailed Localized Image and Video Captioning
★ 1,517+3Star change over the last 7 days - #3
Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.
★ 1,310+4Star change over the last 7 days - #4
[NeurIPS 2024] An official implementation of "ShareGPT4Video: Improving Video Understanding and Generation with Better Captions"
★ 1,094+0Star change over the last 7 days - #5
A Framework of Small-scale Large Multimodal Models
★ 1,004+1Star change over the last 7 days - #6
A collection of resources on applications of multi-modal learning in medical imaging.
★ 975+1Star change over the last 7 days - #7
LLaVA-Plus: Large Language and Vision Assistants that Plug and Learn to Use Skills
★ 770+0Star change over the last 7 days - #8
This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"
★ 593+1Star change over the last 7 days - #9
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images, high-resolution images, and videos in an efficient manner.
★ 577+0Star change over the last 7 days