llava
Tracked open-source repos tagged llava, sorted by stars.
Related topics
Topics that frequently appear alongside llava on the same repo.
Recent risers
Repos created in the last 90 days, tagged llava.
No new repos tagged with this topic in the last 90 days.
- #1
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25,014+12Star change over the last 7 days - #2
SUPIR aims at developing Practical Algorithms for Photo-Realistic Image Restoration In the Wild. Our new online demo is also released at suppixel.ai.
★ 5,649+0Star change over the last 7 days - #3
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
★ 5,462+41Star change over the last 7 days - #4
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4,375+15Star change over the last 7 days - #5★ 3,836+2Star change over the last 7 days
- #6
A C#/.NET library to run LLM (🦙LLaMA/LLaVA) on your local device efficiently.
★ 3,790+3Star change over the last 7 days - #7★ 3,511+46Star change over the last 7 days
- #8★ 2,666+1Star change over the last 7 days
- #9
[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.
★ 1,507+1Star change over the last 7 days - #10★ 1,348+2Star change over the last 7 days
- #11
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
★ 1,247+2Star change over the last 7 days - #12
Fully Open Framework for Democratized Multimodal Training
★ 1,195+3Star change over the last 7 days - #13
A Framework of Small-scale Large Multimodal Models
★ 1,004+1Star change over the last 7 days - #14
🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)
★ 842+0Star change over the last 7 days - #15★ 794+1Star change over the last 7 days
- #16
Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high performance and flexibility.
★ 724+1Star change over the last 7 days - #17
A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.
★ 705+5Star change over the last 7 days - #18
👁️ + 💬 + 🎧 = 🤖 Curated list of top foundation and multimodal models! [Paper + Code + Examples + Tutorials]
★ 637+0Star change over the last 7 days - #19
ComfyUI nodes for vision-language models: Qwen3-VL, Moondream 3, Florence-2, SmolVLM2, InternVL, Gemma 3, MiniCPM-V. Plus open-vocabulary detection, SAM2/SAM3 segmentation, video temporal reasoning, GGUF via llama.cpp, and hosted LLM/VLM APIs.
★ 589+1Star change over the last 7 days - #20
LLaVA-Mini is a unified large multimodal model (LMM) that can support the understanding of images, high-resolution images, and videos in an efficient manner.
★ 577+0Star change over the last 7 days - #21
AI-powered assistant to help you with your daily tasks, powered by Llama 3, DeepSeek R1, and many more models on HuggingFace.
★ 530+0Star change over the last 7 days