Skip to main content
buildradar
Sign in
Topic · multi-modality

multi-modality

Tracked open-source repos tagged multi-modality, sorted by stars.

Repos
11
Total stars
70,438
Avg. stars
6,403
Share
0.00%

Topics that frequently appear alongside multi-modality on the same repo.

Recent risers

Repos created in the last 90 days, tagged multi-modality.

No new repos tagged with this topic in the last 90 days.

  • LLaVA@haotian-liu

    [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.

    25,005+9Star change over the last 7 days
  • :sparkles::sparkles:Latest Advances on Multimodal Large Language Models

    17,994+8Star change over the last 7 days
  • clip-as-service@jina-ai

    🏄 Scalable embedding, reasoning, ranking for images and sentences with CLIP

    12,835+0Star change over the last 7 days
  • MOSS-TTS-Nano@OpenMOSS

    MOSS-TTS-Nano is an open-source multilingual tiny speech generation model from MOSI.AI and the OpenMOSS team. With only 0.1B parameters, it is designed for realtime speech generation, can run directly on CPU without a GPU, and keeps the deployment stack simple enough for local demos, web serving, and lightweight product integration.

    4,271+51Star change over the last 7 days
  • Otter@EvolvingLMMs-Lab

    🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.

    3,436+2Star change over the last 7 days
  • InternLM-XComposer@InternLM

    InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

    2,926+1Star change over the last 7 days
  • Algorithms and Publications on 3D Object Tracking

    1,027+2Star change over the last 7 days
  • VisRAG@OpenBMB

    Parsing-free RAG supported by VLMs

    980+1Star change over the last 7 days
  • Long-RL@NVlabs

    Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)

    729+2Star change over the last 7 days
  • MINIMA@LSXI7

    [CVPR 2025] MINIMA: Modality Invariant Image Matching

    671+2Star change over the last 7 days
  • Multi-Modality-Arena@OpenGVLab

    Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4, LLaMA-Adapter V2, LLaVA, BLIP-2, and many more!

    566+0Star change over the last 7 days
← Back to topics