Skip to main content
buildradar
Sign in
Topic · multimodal-large-language-models

multimodal-large-language-models

Tracked open-source repos tagged multimodal-large-language-models, sorted by stars.

Repos
32
Total stars
66,902
Avg. stars
2,091
Share
0.00%

Topics that frequently appear alongside multimodal-large-language-models on the same repo.

Recent risers

Repos created in the last 90 days, tagged multimodal-large-language-models.

No new repos tagged with this topic in the last 90 days.

  • :sparkles::sparkles:Latest Advances on Multimodal Large Language Models

    17,996+2Star change over the last 7 days
  • MobileAgent@X-PLUG

    Mobile-Agent: The Powerful GUI Agent Family

    9,157+9Star change over the last 7 days
  • star-vector@joanrod

    StarVector is a foundation model for SVG generation that transforms vectorization into a code generation task. Using a vision-language modeling architecture, StarVector processes both visual and textual inputs to produce high-quality SVG code with remarkable precision.

    4,567+6Star change over the last 7 days
  • BayLing-Speech@BayLing-Models

    LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.

    3,147+1Star change over the last 7 days
  • VideoPipe@sherlockchou86

    A cross-platform video structuring (video analysis) framework based on CV models & mLLM.

    2,933+0Star change over the last 7 days
  • VITA@VITA-MLLM

    ✨✨[NeurIPS 2025] VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

    2,532-1Star change over the last 7 days
  • mPLUG-DocOwl@X-PLUG

    mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

    2,411+0Star change over the last 7 days
  • cambrian@cambrian-mllm

    Cambrian-1 is a family of multimodal LLMs with a vision-centric design.

    2,014+1Star change over the last 7 days
  • RPG-DiffusionMaster@YangLing0818

    [ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)

    1,844+0Star change over the last 7 days
  • Seed1.5-VL@ByteDance-Seed

    Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving state-of-the-art performance on 38 out of 60 public benchmarks.

    1,586+0Star change over the last 7 days
  • Ovis@ATH-MaaS

    A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.

    1,514+2Star change over the last 7 days
  • Awesome Unified Multimodal Models

    1,314+1Star change over the last 7 days
  • VideoChat@Henry-23

    实时交互数字人,可自定义形象与音色,支持音色克隆,对话延迟低至3s。Real-time voice interactive digital human, customizable appearance and voice, supporting voice cloning, with initial package delay as low as 3s.

    1,303-1Star change over the last 7 days
  • PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models

    1,184+2Star change over the last 7 days
  • SLAM-LLM@X-LANCE

    A Framework for Speech, Language, Audio, Music Processing with Large Language Model

    1,058+2Star change over the last 7 days
  • Bunny@BAAI-DCAI

    A family of lightweight multimodal models.

    1,053+0Star change over the last 7 days
  • Awesome-MCoT@yaotingwangofficial

    Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

    1,025+2Star change over the last 7 days
  • A collection of resources on applications of multi-modal learning in medical imaging.

    975+1Star change over the last 7 days
  • NEO@EvolvingLMMs-Lab

    NEO Series: Native Vision-Language Models from First Principles

    889+0Star change over the last 7 days
  • Video-MME@MME-Benchmarks

    ✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

    791+2Star change over the last 7 days
  • LLaVA-Plus-Codebase@LLaVA-VL

    LLaVA-Plus: Large Language and Vision Assistants that Plug and Learn to Use Skills

    770+0Star change over the last 7 days
  • MovieChat@wenhaochai

    [CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

    706+0Star change over the last 7 days
  • unicom@deepglint

    Large-Scale Visual Representation Model

    702+0Star change over the last 7 days
  • MPP-LLaVA@Coobiw

    Personal Project: MPP-Qwen14B & MPP-Qwen-Next(Multimodal Pipeline Parallel based on Qwen-LM). Support [video/image/multi-image] {sft/conversations}. Don't let the poverty limit your imagination! Train your own 8B/14B LLaVA-training-like MLLM on RTX3090/4090 24GB.

    687+0Star change over the last 7 days
  • OmniVinci@NVlabs

    OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.

    678+0Star change over the last 7 days
  • Woodpecker@VITA-MLLM

    ✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models

    650+1Star change over the last 7 days
  • Liquid@FoundationVision

    (Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generators

    640+0Star change over the last 7 days
  • Vitron@SkyworkAI

    NeurIPS 2024 Paper: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing

    577+1Star change over the last 7 days
  • LLaVA-Mini@ictnlp

    LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images, high-resolution images, and videos in an efficient manner.

    577+0Star change over the last 7 days
  • cambrian-s@cambrian-mllm

    Cambrian-S: Towards Spatial Supersensing in Video

    567-1Star change over the last 7 days
← Back to topics