Skip to main content
buildradar
Sign in
Topic · multi-modal

multi-modal

Tracked open-source repos tagged multi-modal, sorted by stars.

Repos
42
Total stars
189,535
Avg. stars
4,513
Share
0.01%

Topics that frequently appear alongside multi-modal on the same repo.

Recent risers

Repos created in the last 90 days, tagged multi-modal.

No new repos tagged with this topic in the last 90 days.

  • agentscope@agentscope-ai

    Build and run agents you can see, understand and trust.

    30,463+420Star change over the last 7 days
  • MiniCPM-V@OpenBMB

    A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone

    26,283+27Star change over the last 7 days
  • ten-framework@TEN-framework

    Open-source framework for conversational voice AI agents

    11,102+10Star change over the last 7 days
  • InternVL@OpenGVLab

    [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型

    10,150+3Star change over the last 7 days
  • modelscope@modelscope

    ModelScope: bring the notion of Model-as-a-Service to life.

    9,121+4Star change over the last 7 days
  • big-AGI@enricoros

    AI suite powered by state-of-the-art models and providing advanced AI/AGI functions. Includes AI personas, AGI functions, world-class Beam multi-model chats, text-to-image, voice, response streaming, code highlighting and execution, PDF import, presets for developers, much more. Deploy on-prem or in the cloud.

    7,108+1Star change over the last 7 days
  • data-juicer@datajuicer

    Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷

    6,975+26Star change over the last 7 days
  • CogVLM@zai-org

    a state-of-the-art-level open visual language model | 多模态预训练模型

    6,744+0Star change over the last 7 days
  • valhalla@valhalla

    Open Source Routing Engine for OpenStreetMap

    6,154+24Star change over the last 7 days
  • Chinese-CLIP@OFA-Sys

    Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.

    6,001+1Star change over the last 7 days
  • DALLE-pytorch@lucidrains

    Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch

    5,626-2Star change over the last 7 days
  • marqo@marqo-ai

    Ecommerce Search and Discovery - marqo.ai

    5,031-1Star change over the last 7 days
  • DeepKE@zjunlp

    [EMNLP 2022] An Open Toolkit for Knowledge Graph Extraction and Construction

    4,476+1Star change over the last 7 days
  • VLMEvalKit@open-compass

    Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

    4,375+13Star change over the last 7 days
  • OmniGen@VectorSpaceLab

    OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340

    4,342+1Star change over the last 7 days
  • VisualGLM-6B@zai-org

    Chinese and English multimodal conversational language model | 多模态中英双语对话语言模型

    4,155+0Star change over the last 7 days
  • LLamaSharp@SciSharp

    A C#/.NET library to run LLM (🦙LLaMA/LLaVA) on your local device efficiently.

    3,790+2Star change over the last 7 days
  • Video-LLaVA@PKU-YuanGroup

    【EMNLP 2024🔥】Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

    3,500+1Star change over the last 7 days
  • py-xiaozhi@huangjunsen0406

    Open-source AI assistant ecosystem with MCP integrations, multimodal workflows, IoT support, and cross-platform voice interaction.

    3,463+9Star change over the last 7 days
  • docarray@docarray

    Represent, send, store and search multimodal data

    3,123-1Star change over the last 7 days
  • LISA@JIA-Lab-research

    Project Page for "LISA: Reasoning Segmentation via Large Language Model"

    2,674+0Star change over the last 7 days
  • CogVLM2@zai-org

    GPT4V-level open-source multi-modal model based on Llama3-8B

    2,433+0Star change over the last 7 days
  • MoE-LLaVA@PKU-YuanGroup

    【TMM 2025🔥】 Mixture-of-Experts for Large Vision-Language Models

    2,322+0Star change over the last 7 days
  • RecSysPapers@tangxyw

    推荐/广告/搜索领域工业界经典以及最前沿论文集合。A collection of industry classics and cutting-edge papers in the field of recommendation/advertising/search.

    2,188-1Star change over the last 7 days
  • MotionGPT@OpenMotionLab

    [NeurIPS 2023] MotionGPT: Human Motion as a Foreign Language, a unified motion-language generation model using LLMs

    1,963+1Star change over the last 7 days
  • GPTDiscord@Kav-K

    A robust, all-in-one GPT interface for Discord. ChatGPT-style conversations, image generation, AI-moderation, custom indexes/knowledgebase, youtube summarizer, and more!

    1,853-2Star change over the last 7 days
  • SALMONN@bytedance

    SALMONN family: A suite of advanced multi-modal LLMs

    1,521+3Star change over the last 7 days
  • MedMNIST@MedMNIST

    [pip install medmnist] 18x Standardized Datasets for 2D and 3D Biomedical Image Classification

    1,399+0Star change over the last 7 days
  • transfusion-pytorch@lucidrains

    Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAI

    1,398+3Star change over the last 7 days
  • This repository collects papers for "A Survey on Knowledge Distillation of Large Language Models". We break down KD into Knowledge Elicitation and Distillation Algorithms, and explore the Skill & Vertical Distillation of LLMs.

    1,306+1Star change over the last 7 days
← Back to topics