multimodal
Tracked open-source repos tagged multimodal, sorted by stars.
- #31
Plug in and Play Implementation of Tree of Thoughts: Deliberate Problem Solving with Large Language Models that Elevates Model Reasoning by atleast 70%
★ 4,592+1Star change over the last 7 days - #32
Curated tutorials and resources for Large Language Models, AI Painting, and more.
★ 4,537+2Star change over the last 7 days - #33
Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
★ 4,445+2Star change over the last 7 days - #34
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4,390+12Star change over the last 7 days - #35
Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系,成为中文AIGC和认知智能的基础设施。
★ 4,122-1Star change over the last 7 days - #36
An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS
★ 4,058+26Star change over the last 7 days - #37
OpenMMLab Pre-training Toolbox and Benchmark
★ 3,850+0Star change over the last 7 days - #38
The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。
★ 3,839+134Star change over the last 7 days - #39★ 3,735+15Star change over the last 7 days
- #40
Open-source multimodal retrieval engine (Morphik Core). By Morphik — AI back office for skilled nursing & senior living (morphik.ai).
★ 3,710+2Star change over the last 7 days - #41
From Chain-of-Thought prompting to OpenAI o1 and DeepSeek-R1 🍓
★ 3,681+6Star change over the last 7 days - #42
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
★ 3,634-3Star change over the last 7 days - #43
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
★ 3,414+8Star change over the last 7 days - #44
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
★ 3,205+0Star change over the last 7 days - #45
An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux Foundation.
★ 3,171+13Star change over the last 7 days - #46
Foundation Architecture for (M)LLMs
★ 3,139+1Star change over the last 7 days - #47★ 3,123-1Star change over the last 7 days
- #48
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
★ 3,118+11Star change over the last 7 days - #49
🤖 Type-safe, provider-agnostic TypeScript AI SDK for streaming chat, tool calling, agents, and multimodal apps across OpenAI, Anthropic, Gemini, React, Vue, Svelte, and Solid.
★ 3,057+29Star change over the last 7 days - #50
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
★ 2,925+0Star change over the last 7 days - #51
The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure
★ 2,816+4Star change over the last 7 days - #52
Easily compute clip embeddings and build a clip retrieval system with them
★ 2,796+1Star change over the last 7 days - #53
Images to inference with no labeling (use foundation models to train supervised models).
★ 2,770+6Star change over the last 7 days - #54
streamline the fine-tuning process for multimodal models: PaliGemma 2, Florence-2, and Qwen2.5-VL
★ 2,695+1Star change over the last 7 days - #55★ 2,666+1Star change over the last 7 days
- #56
Comprehensive resources on Generative AI, including a detailed roadmap, projects, use cases, interview preparation, and coding preparation.
★ 2,610+2Star change over the last 7 days - #57
Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
★ 2,557+0Star change over the last 7 days - #58★ 2,539+0Star change over the last 7 days
- #59
HuixiangDou: Overcoming Group Chat Scenarios with LLM-based Technical Assistance
★ 2,502+2Star change over the last 7 days - #60
(ෆ`꒳´ෆ) A Survey on Text-to-Image Generation/Synthesis.
★ 2,444+0Star change over the last 7 days