multi-modal
Tracked open-source repos tagged multi-modal, sorted by stars.
- #31★ 979-1Star change over the last 7 days
- #32
FarmVibes.AI: Multi-Modal GeoSpatial ML Models for Agriculture and Sustainability
★ 895+0Star change over the last 7 days - #33
【ICLR 2024🔥】 Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
★ 885+0Star change over the last 7 days - #34
Use late-interaction multi-modal models such as ColPali in just a few lines of code.
★ 852+0Star change over the last 7 days - #35
[CVPR 2026🔥] 🧑🎨 OmniLottie, an open-sourced multi-modal instructed vector animation generator that produces Lottie JSONs.
★ 774+3Star change over the last 7 days - #36
Multi-modal OCR pipeline optimized for ML training (text, figure, math, tables, diagrams)
★ 676+0Star change over the last 7 days - #37
Unified Controllable Visual Generation Model
★ 662+0Star change over the last 7 days - #38
[ICCV 2025 & ICCV 2025 RIWM Outstanding Paper] Aether: Geometric-Aware Unified World Modeling
★ 609+0Star change over the last 7 days - #39★ 583-1Star change over the last 7 days
- #40
FashionCLIP is a CLIP-like model fine-tuned for the fashion domain.
★ 537+3Star change over the last 7 days - #41
Multi-model DAG-driven parallel AI film generation — parallel speedup scales with scene independence; Generate film scenes simultaneously instead of one by one; "把影视生成的执行图从拓扑序变成关键路径最优调度" ; 唯一把场景叙事依赖建模为 DAG、以 CPM 算法驱动并行调度的影视生成引擎
★ 536+0Star change over the last 7 days - #42
🔥🔥🔥[AAAI 2026 Oral] Official Implementation of Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding
★ 511+0Star change over the last 7 days