multimodal-llm
標記 multimodal-llm 主題、收錄中的開源專案,依星數排序。
相關主題
常跟 multimodal-llm 一起出現在同一個專案上的主題。
近期新秀
近 90 天內建立、標記 multimodal-llm 主題的專案。
- #1
WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.
★ 1,325 - #2
🎬 OraRL — Annotations as Rollouts for efficient, scalable reinforcement learning of unified video MLLMs.
★ 152
- #1
支援普通話、中文方言與英語的開源工業級 ASR 模型,在公開的普通話 ASR 基準測試中達到全新的 SOTA,同時提供出色的歌唱歌詞辨識能力。
★ 1,975+3近 7 天星數變化 - #2
WeMM-Embedding is a family of universal multimodal embedding models by the WeChat Vision Team at Tencent, supporting multimodal understanding and retrieval.
★ 1,325+892近 7 天星數變化 - #3
包含 155 種以上 VLM/MLLM 架構的精選視覺化目錄:包含論文、圖表、訓練配方、資料集,以及多模態 AI 代理的發布時間軸。
★ 1,310+4近 7 天星數變化 - #4
論文「MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens」的官方實作
★ 869+0近 7 天星數變化 - #5
一個 SOTA 工業級一體化 ASR 系統,包含 ASR、VAD、LID 與 Punc 模組。FireRedASR2 支援中文(普通話、20+ 種方言/口音)、英文、中英夾雜(code-switching),以及語音和歌唱的 ASR。FireRedVAD 支援 100+ 種語言的語音/歌唱/音樂。FireRedLID 支援 100+ 種語言與 20+ 種中文方言。FireRedPunc 支援中英文。
★ 669+10近 7 天星數變化