Skip to main content
buildradar
Sign in
Topic · mllm

mllm

Tracked open-source repos tagged mllm, sorted by stars.

Repos
36
Total stars
92,155
Avg. stars
2,560
Share
0.00%

Topics that frequently appear alongside mllm on the same repo.

Recent risers

Repos created in the last 90 days, tagged mllm.

  • unilm@microsoft

    Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities

    22,203+7Star change over the last 7 days
  • Agent-S@simular-ai

    Agent S: an open agentic framework that uses computers like a human

    12,219+17Star change over the last 7 days
  • MobileAgent@X-PLUG

    Mobile-Agent: The Powerful GUI Agent Family

    9,157+9Star change over the last 7 days
  • SpatialLM@manycore-research

    [NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modeling

    4,726+3Star change over the last 7 days
  • MagicQuill@robbyant-research

    [CVPR'25] Official Implementations for Paper - MagicQuill: An Intelligent Interactive Image Editing System

    3,691+2Star change over the last 7 days
  • From Chain-of-Thought prompting to OpenAI o1 and DeepSeek-R1 🍓

    3,681+3Star change over the last 7 days
  • NExT-GPT@NExT-GPT

    Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model

    3,634-2Star change over the last 7 days
  • Eagle@NVlabs

    Eagle: Frontier Vision-Language Models with Data-Centric Strategies

    3,511+32Star change over the last 7 days
  • InternLM-XComposer@InternLM

    InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

    2,925-1Star change over the last 7 days
  • mPLUG-DocOwl@X-PLUG

    mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

    2,411+0Star change over the last 7 days
  • cambrian@cambrian-mllm

    Cambrian-1 is a family of multimodal LLMs with a vision-centric design.

    2,014+1Star change over the last 7 days
  • 🚀🚀🚀 A collection of some awesome public YOLO object detection series projects and the related object detection datasets.

    1,783-2Star change over the last 7 days
  • Sa2VA@bytedance

    Official Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)

    1,666+0Star change over the last 7 days
  • Rex-Omni@IDEA-Research

    [CVPR2026] Detect Anything via Next Point Prediction

    1,570+8Star change over the last 7 days
  • Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.

    1,310+2Star change over the last 7 days
  • LLaVA-OneVision-2@EvolvingLMMs-Lab

    Fully Open Framework for Democratized Multimodal Training

    1,195+1Star change over the last 7 days
  • Bunny@BAAI-DCAI

    A family of lightweight multimodal models.

    1,053+0Star change over the last 7 days
  • OpenEMMA@taco-group

    OpenEMMA, a permissively licensed open source "reproduction" of Waymo’s EMMA model.

    951+0Star change over the last 7 days
  • NEO@EvolvingLMMs-Lab

    NEO Series: Native Vision-Language Models from First Principles

    888+0Star change over the last 7 days
  • JarvisArt@LYL1015

    [NeurIPS' 2025] JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent

    862+1Star change over the last 7 days
  • Osprey@CircleRadon

    [CVPR2024] The code for "Osprey: Pixel Understanding with Visual Instruction Tuning"

    843+0Star change over the last 7 days
  • FSDrive@MIV-XJTU

    [NeurIPS 2025 spotlight] Official implementation for "FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving"

    831+9Star change over the last 7 days
  • 4KAgent@taco-group

    [NeurIPS 2025] 4KAgent: Agentic Any Image to 4K Super-Resolution. An intelligent computer vision agent that can magically restore any image to perfect-4K!

    821+2Star change over the last 7 days
  • 🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language Model(VLM), Vision Language Action(VLA), AI Generated Content(AIGC), the related Datasets and Applications.

    815+1Star change over the last 7 days
  • This is a repository for listing papers on scene graph generation and application.

    724+5Star change over the last 7 days
  • MPP-LLaVA@Coobiw

    Personal Project: MPP-Qwen14B & MPP-Qwen-Next(Multimodal Pipeline Parallel based on Qwen-LM). Support [video/image/multi-image] {sft/conversations}. Don't let the poverty limit your imagination! Train your own 8B/14B LLaVA-training-like MLLM on RTX3090/4090 24GB.

    686+0Star change over the last 7 days
  • SenseNova-Vision@OpenSenseNova

    Vision as Unified Multimodal Generation

    680+25Star change over the last 7 days
  • Woodpecker@VITA-MLLM

    ✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models

    650+1Star change over the last 7 days
  • FakeShield@zhipeixu

    [ICLR 2025] FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models

    623+2Star change over the last 7 days
  • Emotion-LLaMA@ZebangCheng

    Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning

    617+0Star change over the last 7 days
← Back to topics