Skip to main content
buildradar
Sign in
Topic · multimodal

multimodal

Tracked open-source repos tagged multimodal, sorted by stars.

148 repos
  • ✨✨Latest Papers and Benchmarks in Reasoning with Foundation Models

    657+1Star change over the last 7 days
  • MOSS-Audio@OpenMOSS

    An open-source model for understanding speech, environmental sounds, and music through captioning, question answering, and reasoning

    656+5Star change over the last 7 days
  • SEED@AILab-CVC

    Official implementation of SEED-LLaMA (ICLR 2024).

    641-1Star change over the last 7 days
  • Seg-Zero@JIA-Lab-research

    Project Page For "Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement"

    639+0Star change over the last 7 days
  • 👁️ + 💬 + 🎧 = 🤖 Curated list of top foundation and multimodal models! [Paper + Code + Examples + Tutorials]

    637+0Star change over the last 7 days
  • blended-latent-diffusion@omriav

    Official implementation for "Blended Latent Diffusion" [SIGGRAPH 2023]

    632+0Star change over the last 7 days
  • second-brain@henrydaum

    Second Brain is an agentic framework that acts as an operating system, using local file intelligence, workflow automation, and LLMs to complete tasks and communicate over multiple modalities and messaging platforms.

    630+12Star change over the last 7 days
  • RAG-Driven-Generative-AI@Denis2054

    This repository provides programs to build Retrieval Augmented Generation (RAG) code for Generative AI with LlamaIndex, Deep Lake, and Pinecone leveraging the power of OpenAI and Hugging Face models for generation and evaluation.

    624+2Star change over the last 7 days
  • VisualThinker-R1-Zero@turningpoint-ai

    Explore the Multimodal “Aha Moment” on 2B Model

    623+0Star change over the last 7 days
  • Hunyuan3D-Omni@Tencent-Hunyuan

    Hunyuan3D-Omni: A Unified Framework for Controllable Generation of 3D Assets

    616+3Star change over the last 7 days
  • tokenize-anything@baaivision

    [ECCV 2024] Tokenize Anything via Prompting

    600+0Star change over the last 7 days
  • MOSS-VL@OpenMOSS

    An open-weight 11B model series for long-form and real-time video understanding

    596Star change over the last 7 days
  • MMMU@MMMU-Benchmark

    This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"

    594+1Star change over the last 7 days
  • Relax@redai-studio

    An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

    592+11Star change over the last 7 days
  • blended-diffusion@omriav

    Official implementation for "Blended Diffusion for Text-driven Editing of Natural Images" [CVPR 2022]

    589+0Star change over the last 7 days
  • Groma@FoundationVision

    [ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization

    586+0Star change over the last 7 days
  • AI-Employe@vignshwarar

    Create browser automation as if you were teaching a human using GPT-4 Vision.

    586+0Star change over the last 7 days
  • rai@RobotecAI

    RAI is a vendor agnostic agentic framework for Physical AI robotics, utilizing ROS 2 tools to perform complex actions, defined scenarios, free interface execution, log summaries, voice interaction and more.

    582+6Star change over the last 7 days
  • motis@motis-project

    multimodal routing, geocoding, and map tiles

    579+13Star change over the last 7 days
  • LLaVA-Mini@ictnlp

    LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images, high-resolution images, and videos in an efficient manner.

    577+0Star change over the last 7 days
  • The Self-Coding System for Your App — Alan AI SDK for React Native

    575+0Star change over the last 7 days
  • multimodal-agents-course@the-ai-merge

    An MCP Multimodal AI Agent with eyes and ears!

    575-1Star change over the last 7 days
  • psi@microsoft

    Platform for Situated Intelligence

    572+1Star change over the last 7 days
  • clip.cpp@monatis

    CLIP inference in plain C/C++ with no extra dependencies

    568+0Star change over the last 7 days
  • Flame-Code-VLM@Flame-Code-VLM

    Flame is an open-source multimodal AI system designed to translate UI design mockups into high-quality React code. It leverages vision-language modeling, automated data synthesis, and structured training workflows to bridge the gap between design and front-end development.

    562+0Star change over the last 7 days
  • vlmrun-hub@vlm-run

    A hub for various industry-specific schemas to be used with VLMs.

    555+0Star change over the last 7 days
  • Awesome Multimodal Modeling [Covers MLLM, UMM, and NMM]

    543+2Star change over the last 7 days
  • EVF-SAM@hustvl

    Official code of "EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model"

    508+1Star change over the last 7 days
← Back to topics