multimodal-llm
标记 multimodal-llm 主题、收录中的开源项目,按星标数排序。
相关主题
常跟 multimodal-llm 一起出现在同一个项目上的主题。
近期新秀
近 90 天内创建、标记 multimodal-llm 主题的项目。
- #1
WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.
★ 1,461 - #2
🎬 OraRL — Annotations as Rollouts for efficient, scalable reinforcement learning of unified video MLLMs.
★ 152
- #1
支援普通话、中文方言与英语的开源工业级 ASR 模型,在公开的普通话 ASR 基准测试中达到全新的 SOTA,同时提供出色的歌唱歌词辨识能力。
★ 1,975+0近 7 天星标变化 - #2
WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.
★ 1,461+363近 7 天星标变化 - #3
包含 155 种以上 VLM/MLLM 架构的精选视觉化目录:包含论文、图表、训练配方、数据集,以及多模态 AI 代理的发布时间轴。
★ 1,310+0近 7 天星标变化 - #4
论文「MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens」的官方实现
★ 869+0近 7 天星标变化 - #5
一个 SOTA 工业级一体化 ASR 系统,包含 ASR、VAD、LID 和 Punc 模块。FireRedASR2 支持中文(普通话、20+ 种方言/口音)、英文、中英夹杂(code-switching),以及语音和歌唱的 ASR。FireRedVAD 支持 100+ 种语言的语音/歌唱/音乐。FireRedLID 支持 100+ 种语言和 20+ 种中文方言。FireRedPunc 支持中英文。
★ 676+8近 7 天星标变化