vlm
Tracked open-source repos tagged vlm, sorted by stars.
- #31
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.
★ 1,840+10Star change over the last 7 days - #32
🚀🚀🚀 A collection of some awesome public YOLO object detection series projects and the related object detection datasets.
★ 1,783-2Star change over the last 7 days - #33
Declarative way to run AI models in React Native on device, powered by ExecuTorch.
★ 1,707+5Star change over the last 7 days - #34
OpenGUI is an Android GUI agent framework for phone-use AI that can see, plan, and operate real mobile apps through the GUI.
★ 1,624+38Star change over the last 7 days - #35
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
★ 1,619+58Star change over the last 7 days - #36
Awesome-Jailbreak-on-LLMs is a collection of state-of-the-art, novel, exciting jailbreak methods on LLMs. It contains papers, codes, datasets, evaluations, and analyses.
★ 1,610+5Star change over the last 7 days - #37★ 1,507+2Star change over the last 7 days
- #38
Aircraft design optimization made fast through computational graph transformations (e.g., automatic differentiation). Composable analysis tools for aerodynamics, propulsion, structures, trajectory design, and much more.
★ 1,323+4Star change over the last 7 days - #39
Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.
★ 1,310+2Star change over the last 7 days - #40
JoyCaption is an image captioning Visual Language Model (VLM) being built from the ground up as a free, open, and uncensored model for the community to use in training Diffusion models.
★ 1,248+1Star change over the last 7 days - #41★ 1,195+1Star change over the last 7 days
- #42
InternRobotics' open platform for building generalized navigation foundation models.
★ 1,090+18Star change over the last 7 days - #43★ 1,053+0Star change over the last 7 days
- #44
MindSpore + 🤗Huggingface: Run any Transformers/Diffusers model on MindSpore with seamless compatibility and acceleration.
★ 920+0Star change over the last 7 days - #45
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
★ 905+24Star change over the last 7 days - #46
【新增智能体模式】安卓端全场景GPT助手,可用音量键唤起并进行语音交流,支持联网、拍照、模板、附件解析、智能体模式等 | GPT assistant for Android, activated via volume keys for voice interaction, supporting features such as networking, taking photos, templates, parsing PDF and Office documents, and agent mode.
★ 900+0Star change over the last 7 days - #47
UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
★ 890+0Star change over the last 7 days - #48★ 888+0Star change over the last 7 days
- #49★ 874+1Star change over the last 7 days
- #50
Unified Codebase for Advanced World Models.
★ 866+2Star change over the last 7 days - #51
Official Repository of "LLM × DATA" Survey Paper
★ 821+2Star change over the last 7 days - #52
A curated list of 3D Vision papers relating to Robotics domain in the era of large models i.e. LLMs/VLMs, inspired by awesome-computer-vision, including papers, codes, and related websites
★ 820+1Star change over the last 7 days - #53
🚀🚀🚀A collection of some awesome public projects about Large Language Model(LLM), Vision Language Model(VLM), Vision Language Action(VLA), AI Generated Content(AIGC), the related Datasets and Applications.
★ 815+1Star change over the last 7 days - #54
[EMNLP 2026] MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research · 浏览器里运行的安卓模拟器 · Browser-hosted Android Simulator · Verifiable Evaluation · Scalable Online RL Training
★ 785+9Star change over the last 7 days - #55
[CVPR 2026🔥] 🧑🎨 OmniLottie, an open-sourced multi-modal instructed vector animation generator that produces Lottie JSONs.
★ 774+3Star change over the last 7 days - #56★ 753+3Star change over the last 7 days
- #57
[CVPR 2024 🔥] GeoChat, the first grounded Large Vision Language Model for Remote Sensing
★ 750+6Star change over the last 7 days - #58
[ECCV 2026] SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
★ 702+3Star change over the last 7 days - #59
This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]
★ 682+2Star change over the last 7 days - #60
Production-grade C++ edge AI engine for video analytics and on-device VLM across Sophon, Rockchip RKNN, and x86, with visual orchestration, real-time OSD, events, and reproducible benchmarks.
★ 656+45Star change over the last 7 days