multi-modality
Tracked open-source repos tagged multi-modality, sorted by stars.
Related topics
Topics that frequently appear alongside multi-modality on the same repo.
Recent risers
Repos created in the last 90 days, tagged multi-modality.
No new repos tagged with this topic in the last 90 days.
- #1
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25,005+9Star change over the last 7 days - #2
:sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 17,994+8Star change over the last 7 days - #3
🏄 Scalable embedding, reasoning, ranking for images and sentences with CLIP
★ 12,835+0Star change over the last 7 days - #4
MOSS-TTS-Nano is an open-source multilingual tiny speech generation model from MOSI.AI and the OpenMOSS team. With only 0.1B parameters, it is designed for realtime speech generation, can run directly on CPU without a GPU, and keeps the deployment stack simple enough for local demos, web serving, and lightweight product integration.
★ 4,271+51Star change over the last 7 days - #5
🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.
★ 3,436+2Star change over the last 7 days - #6
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
★ 2,926+1Star change over the last 7 days - #7
Algorithms and Publications on 3D Object Tracking
★ 1,027+2Star change over the last 7 days - #8★ 980+1Star change over the last 7 days
- #9★ 729+2Star change over the last 7 days
- #10★ 671+2Star change over the last 7 days
- #11
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4, LLaMA-Adapter V2, LLaVA, BLIP-2, and many more!
★ 566+0Star change over the last 7 days