video-understanding
Tracked open-source repos tagged video-understanding, sorted by stars.
Related topics
Topics that frequently appear alongside video-understanding on the same repo.
Recent risers
Repos created in the last 90 days, tagged video-understanding.
- #1★ 5,150+3Star change over the last 7 days
- #2
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4,390+7Star change over the last 7 days - #3
[CVPR2024 Highlight][VideoChatGPT] ChatGPT with video understanding! And many more supported LMs such as miniGPT4, StableLM, and MOSS.
★ 3,347+0Star change over the last 7 days - #4
GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
★ 2,379+2Star change over the last 7 days - #5
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
★ 2,374+5Star change over the last 7 days - #6
[ICCV 2019] TSM: Temporal Shift Module for Efficient Video Understanding
★ 2,222+1Star change over the last 7 days - #7
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.
★ 1,840+10Star change over the last 7 days - #8
[EMNLP2026] "VideoAgent: All-in-One Agentic Framework for Video Understanding and Editing, and Remaking"
★ 1,834+65Star change over the last 7 days - #9
Awesome video understanding toolkits based on PaddlePaddle. It supports video data annotation tools, lightweight RGB and skeleton based action recognition model, practical applications for video tagging and sport action detection.
★ 1,702+0Star change over the last 7 days - #10★ 1,521+3Star change over the last 7 days
- #11
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
★ 1,334+3Star change over the last 7 days - #12
awesome grounding: A curated list of research papers in visual grounding
★ 1,127+0Star change over the last 7 days - #13
:fire: :fire: :fire: A paper list of some recent Computer Vision(CV) works
★ 973+1Star change over the last 7 days - #14
[CVPR 2024 Highlight🔥] Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
★ 941+0Star change over the last 7 days - #15
[CVPR 2023] VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking
★ 819+3Star change over the last 7 days - #16
Official code for Goldfish model for long video understanding and MiniGPT4-video for short video understanding
★ 639+1Star change over the last 7 days - #17★ 596—Star change over the last 7 days
- #18
ComfyUI nodes for vision-language models: Qwen3-VL, Moondream 3, Florence-2, SmolVLM2, InternVL, Gemma 3, MiniCPM-V. Plus open-vocabulary detection, SAM2/SAM3 segmentation, video temporal reasoning, GGUF via llama.cpp, and hosted LLM/VLM APIs.
★ 588-1Star change over the last 7 days - #19
[ICCV 2023 & TPAMI 2025] MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions
★ 525+0Star change over the last 7 days