multimodal
Tracked open-source repos tagged multimodal, sorted by stars.
- #121
✨✨Latest Papers and Benchmarks in Reasoning with Foundation Models
★ 657+1Star change over the last 7 days - #122
An open-source model for understanding speech, environmental sounds, and music through captioning, question answering, and reasoning
★ 656+5Star change over the last 7 days - #123★ 641-1Star change over the last 7 days
- #124
Project Page For "Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement"
★ 639+0Star change over the last 7 days - #125
👁️ + 💬 + 🎧 = 🤖 Curated list of top foundation and multimodal models! [Paper + Code + Examples + Tutorials]
★ 637+0Star change over the last 7 days - #126
Official implementation for "Blended Latent Diffusion" [SIGGRAPH 2023]
★ 632+0Star change over the last 7 days - #127
Second Brain is an agentic framework that acts as an operating system, using local file intelligence, workflow automation, and LLMs to complete tasks and communicate over multiple modalities and messaging platforms.
★ 630+12Star change over the last 7 days - #128
This repository provides programs to build Retrieval Augmented Generation (RAG) code for Generative AI with LlamaIndex, Deep Lake, and Pinecone leveraging the power of OpenAI and Hugging Face models for generation and evaluation.
★ 624+2Star change over the last 7 days - #129
Explore the Multimodal “Aha Moment” on 2B Model
★ 623+0Star change over the last 7 days - #130
Hunyuan3D-Omni: A Unified Framework for Controllable Generation of 3D Assets
★ 616+3Star change over the last 7 days - #131
[ECCV 2024] Tokenize Anything via Prompting
★ 600+0Star change over the last 7 days - #132★ 596—Star change over the last 7 days
- #133
This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"
★ 594+1Star change over the last 7 days - #134
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
★ 592+11Star change over the last 7 days - #135
Official implementation for "Blended Diffusion for Text-driven Editing of Natural Images" [CVPR 2022]
★ 589+0Star change over the last 7 days - #136
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
★ 586+0Star change over the last 7 days - #137
Create browser automation as if you were teaching a human using GPT-4 Vision.
★ 586+0Star change over the last 7 days - #138
RAI is a vendor agnostic agentic framework for Physical AI robotics, utilizing ROS 2 tools to perform complex actions, defined scenarios, free interface execution, log summaries, voice interaction and more.
★ 582+6Star change over the last 7 days - #139★ 579+13Star change over the last 7 days
- #140
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images, high-resolution images, and videos in an efficient manner.
★ 577+0Star change over the last 7 days - #141
The Self-Coding System for Your App — Alan AI SDK for React Native
★ 575+0Star change over the last 7 days - #142
An MCP Multimodal AI Agent with eyes and ears!
★ 575-1Star change over the last 7 days - #143★ 572+1Star change over the last 7 days
- #144★ 568+0Star change over the last 7 days
- #145
Flame is an open-source multimodal AI system designed to translate UI design mockups into high-quality React code. It leverages vision-language modeling, automated data synthesis, and structured training workflows to bridge the gap between design and front-end development.
★ 562+0Star change over the last 7 days - #146
A hub for various industry-specific schemas to be used with VLMs.
★ 555+0Star change over the last 7 days - #147
Awesome Multimodal Modeling [Covers MLLM, UMM, and NMM]
★ 543+2Star change over the last 7 days - #148
Official code of "EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model"
★ 508+1Star change over the last 7 days