multi-modal
Tracked open-source repos tagged multi-modal, sorted by stars.
Related topics
Topics that frequently appear alongside multi-modal on the same repo.
Recent risers
Repos created in the last 90 days, tagged multi-modal.
No new repos tagged with this topic in the last 90 days.
- #1
Build and run agents you can see, understand and trust.
★ 30,463+420Star change over the last 7 days - #2
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
★ 26,283+27Star change over the last 7 days - #3
Open-source framework for conversational voice AI agents
★ 11,102+10Star change over the last 7 days - #4
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
★ 10,150+3Star change over the last 7 days - #5
ModelScope: bring the notion of Model-as-a-Service to life.
★ 9,121+4Star change over the last 7 days - #6
AI suite powered by state-of-the-art models and providing advanced AI/AGI functions. Includes AI personas, AGI functions, world-class Beam multi-model chats, text-to-image, voice, response streaming, code highlighting and execution, PDF import, presets for developers, much more. Deploy on-prem or in the cloud.
★ 7,108+1Star change over the last 7 days - #7
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
★ 6,975+26Star change over the last 7 days - #8★ 6,744+0Star change over the last 7 days
- #9★ 6,154+24Star change over the last 7 days
- #10
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
★ 6,001+1Star change over the last 7 days - #11
Implementation / replication of DALL-E, OpenAI's Text to Image Transformer, in Pytorch
★ 5,626-2Star change over the last 7 days - #12★ 5,031-1Star change over the last 7 days
- #13★ 4,476+1Star change over the last 7 days
- #14
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4,375+13Star change over the last 7 days - #15★ 4,342+1Star change over the last 7 days
- #16
Chinese and English multimodal conversational language model | 多模态中英双语对话语言模型
★ 4,155+0Star change over the last 7 days - #17
A C#/.NET library to run LLM (🦙LLaMA/LLaVA) on your local device efficiently.
★ 3,790+2Star change over the last 7 days - #18
【EMNLP 2024🔥】Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
★ 3,500+1Star change over the last 7 days - #19
Open-source AI assistant ecosystem with MCP integrations, multimodal workflows, IoT support, and cross-platform voice interaction.
★ 3,463+9Star change over the last 7 days - #20★ 3,123-1Star change over the last 7 days
- #21★ 2,674+0Star change over the last 7 days
- #22★ 2,433+0Star change over the last 7 days
- #23★ 2,322+0Star change over the last 7 days
- #24
推荐/广告/搜索领域工业界经典以及最前沿论文集合。A collection of industry classics and cutting-edge papers in the field of recommendation/advertising/search.
★ 2,188-1Star change over the last 7 days - #25
[NeurIPS 2023] MotionGPT: Human Motion as a Foreign Language, a unified motion-language generation model using LLMs
★ 1,963+1Star change over the last 7 days - #26
A robust, all-in-one GPT interface for Discord. ChatGPT-style conversations, image generation, AI-moderation, custom indexes/knowledgebase, youtube summarizer, and more!
★ 1,853-2Star change over the last 7 days - #27★ 1,521+3Star change over the last 7 days
- #28
[pip install medmnist] 18x Standardized Datasets for 2D and 3D Biomedical Image Classification
★ 1,399+0Star change over the last 7 days - #29
Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAI
★ 1,398+3Star change over the last 7 days - #30
This repository collects papers for "A Survey on Knowledge Distillation of Large Language Models". We break down KD into Knowledge Elicitation and Distillation Algorithms, and explore the Skill & Vertical Distillation of LLMs.
★ 1,306+1Star change over the last 7 days