Skip to main content
buildradar
Sign in
Topic · multimodal

multimodal

Tracked open-source repos tagged multimodal, sorted by stars.

147 repos
  • tree-of-thoughts@kyegomez

    Plug in and Play Implementation of Tree of Thoughts: Deliberate Problem Solving with Large Language Models that Elevates Model Reasoning by atleast 70%

    4,592+1Star change over the last 7 days
  • Curated tutorials and resources for Large Language Models, AI Painting, and more.

    4,537+2Star change over the last 7 days
  • img2dataset@rom1504

    Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.

    4,445+2Star change over the last 7 days
  • lmms-eval@EvolvingLMMs-Lab

    One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

    4,390+12Star change over the last 7 days
  • Fengshenbang-LM@IDEA-CCNL

    Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系,成为中文AIGC和认知智能的基础设施。

    4,122-1Star change over the last 7 days
  • MOSS-TTS@OpenMOSS

    An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS

    4,058+26Star change over the last 7 days
  • mmpretrain@open-mmlab

    OpenMMLab Pre-training Toolbox and Benchmark

    3,850+0Star change over the last 7 days
  • modlens@liustack

    The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。

    3,839+134Star change over the last 7 days
  • SimpleMem@aiming-lab

    SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal

    3,735+15Star change over the last 7 days
  • morphik-core@morphik-org

    Open-source multimodal retrieval engine (Morphik Core). By Morphik — AI back office for skilled nursing & senior living (morphik.ai).

    3,710+2Star change over the last 7 days
  • From Chain-of-Thought prompting to OpenAI o1 and DeepSeek-R1 🍓

    3,681+6Star change over the last 7 days
  • NExT-GPT@NExT-GPT

    Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model

    3,634-3Star change over the last 7 days
  • mteb@embeddings-benchmark

    MTEB: State-of-the-art evaluation of embeddings across languages and modalities

    3,414+8Star change over the last 7 days
  • InternGPT@OpenGVLab

    InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)

    3,205+0Star change over the last 7 days
  • vortex@vortex-data

    An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux Foundation.

    3,171+13Star change over the last 7 days
  • torchscale@microsoft

    Foundation Architecture for (M)LLMs

    3,139+1Star change over the last 7 days
  • docarray@docarray

    Represent, send, store and search multimodal data

    3,123-1Star change over the last 7 days
  • OSWorld@xlang-ai

    [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

    3,118+11Star change over the last 7 days
  • ai@TanStack

    🤖 Type-safe, provider-agnostic TypeScript AI SDK for streaming chat, tool calling, agents, and multimodal apps across OpenAI, Anthropic, Gemini, React, Vue, Svelte, and Solid.

    3,057+29Star change over the last 7 days
  • InternLM-XComposer@InternLM

    InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

    2,925+0Star change over the last 7 days
  • datachain@datachain-ai

    The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure

    2,816+4Star change over the last 7 days
  • clip-retrieval@rom1504

    Easily compute clip embeddings and build a clip retrieval system with them

    2,796+1Star change over the last 7 days
  • autodistill@autodistill

    Images to inference with no labeling (use foundation models to train supervised models).

    2,770+6Star change over the last 7 days
  • maestro@roboflow

    streamline the fine-tuning process for multimodal models: PaliGemma 2, Florence-2, and Qwen2.5-VL

    2,695+1Star change over the last 7 days
  • OmAgent@om-ai-lab

    [EMNLP-2024] Build multimodal language agents for fast prototype and production

    2,666+1Star change over the last 7 days
  • generative-ai@genieincodebottle

    Comprehensive resources on Generative AI, including a detailed roadmap, projects, use cases, interview preparation, and coding preparation.

    2,610+2Star change over the last 7 days
  • OFA@OFA-Sys

    Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

    2,557+0Star change over the last 7 days
  • mPLUG-Owl@X-PLUG

    mPLUG-Owl: The Powerful Multi-modal Large Language Model Family

    2,539+0Star change over the last 7 days
  • HuixiangDou@InternLM

    HuixiangDou: Overcoming Group Chat Scenarios with LLM-based Technical Assistance

    2,502+2Star change over the last 7 days
  • (ෆ`꒳´ෆ) A Survey on Text-to-Image Generation/Synthesis.

    2,444+0Star change over the last 7 days
← Back to topics