Skip to main content
buildradar
Sign in
Topic · video-understanding

video-understanding

Tracked open-source repos tagged video-understanding, sorted by stars.

Repos
19
Total stars
34,235
Avg. stars
1,802
Share
0.00%

Topics that frequently appear alongside video-understanding on the same repo.

Recent risers

Repos created in the last 90 days, tagged video-understanding.

  • OraRL@HVision-NKU

    🎬 OraRL — Annotations as Rollouts for efficient, scalable reinforcement learning of unified video MLLMs.

    152
  • mmaction2@open-mmlab

    OpenMMLab's Next Generation Video Understanding Toolbox and Benchmark

    5,150+3Star change over the last 7 days
  • lmms-eval@EvolvingLMMs-Lab

    One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

    4,390+7Star change over the last 7 days
  • Ask-Anything@OpenGVLab

    [CVPR2024 Highlight][VideoChatGPT] ChatGPT with video understanding! And many more supported LMs such as miniGPT4, StableLM, and MOSS.

    3,347+0Star change over the last 7 days
  • GLM-V@zai-org

    GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

    2,379+2Star change over the last 7 days
  • InternVideo@OpenGVLab

    [ECCV2024] Video Foundation Models & Data for Multimodal Understanding

    2,374+5Star change over the last 7 days
  • temporal-shift-module@mit-han-lab

    [ICCV 2019] TSM: Temporal Shift Module for Efficient Video Understanding

    2,222+1Star change over the last 7 days
  • video-search-and-summarization@NVIDIA-AI-Blueprints

    NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.

    1,840+10Star change over the last 7 days
  • VideoAgent@HKUDS

    [EMNLP2026] "VideoAgent: All-in-One Agentic Framework for Video Understanding and Editing, and Remaking"

    1,834+65Star change over the last 7 days
  • PaddleVideo@PaddlePaddle

    Awesome video understanding toolkits based on PaddlePaddle. It supports video data annotation tools, lightweight RGB and skeleton based action recognition model, practical applications for video tagging and sport action detection.

    1,702+0Star change over the last 7 days
  • SALMONN@bytedance

    SALMONN family: A suite of advanced multi-modal LLMs

    1,521+3Star change over the last 7 days
  • Lance@bytedance

    A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.

    1,334+3Star change over the last 7 days
  • awesome grounding: A curated list of research papers in visual grounding

    1,127+0Star change over the last 7 days
  • :fire: :fire: :fire: A paper list of some recent Computer Vision(CV) works

    973+1Star change over the last 7 days
  • Chat-UniVi@PKU-YuanGroup

    [CVPR 2024 Highlight🔥] Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

    941+0Star change over the last 7 days
  • VideoMAEv2@OpenGVLab

    [CVPR 2023] VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking

    819+3Star change over the last 7 days
  • MiniGPT4-video@Vision-CAIR

    Official code for Goldfish model for long video understanding and MiniGPT4-video for short video understanding

    639+1Star change over the last 7 days
  • MOSS-VL@OpenMOSS

    An open-weight 11B model series for long-form and real-time video understanding

    596Star change over the last 7 days
  • ComfyUI_VLM_nodes@gokayfem

    ComfyUI nodes for vision-language models: Qwen3-VL, Moondream 3, Florence-2, SmolVLM2, InternVL, Gemma 3, MiniCPM-V. Plus open-vocabulary detection, SAM2/SAM3 segmentation, video temporal reasoning, GGUF via llama.cpp, and hosted LLM/VLM APIs.

    588-1Star change over the last 7 days
  • MeViS@henghuiding

    [ICCV 2023 & TPAMI 2025] MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions

    525+0Star change over the last 7 days
← Back to topics