video-generation
Tracked open-source repos tagged video-generation, sorted by stars.
- #91
Video generation from text&image, 1st-gen
★ 919-2Star change over the last 7 days - #92
[AAAI 2025] Follow-Your-Click: This repo is the official implementation of "Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts"
★ 908-1Star change over the last 7 days - #93★ 896+1Star change over the last 7 days
- #94
[ACM MM 2025] Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
★ 881+10Star change over the last 7 days - #95
[ICLR2026] SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training
★ 867+13Star change over the last 7 days - #96
[CVPR 2025 Highlight🔥] Identity-Preserving Text-to-Video Generation by Frequency Decomposition
★ 857+2Star change over the last 7 days - #97
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
★ 850-2Star change over the last 7 days - #98
Flatkey media generation CLI for images, videos, audio, text, credits, and model discovery.
★ 845+211Star change over the last 7 days - #99
[SIGGRAPH 2025] Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
★ 834+3Star change over the last 7 days - #100
Official implementation for "RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers" (ICML 2025) , UltraViCo (ICLR 2026) and UltraImage
★ 827+2Star change over the last 7 days - #101
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
★ 825+1Star change over the last 7 days - #102
Kandinsky 5.0: A family of diffusion models for Video & Image generation
★ 822+6Star change over the last 7 days - #103
Generate video from text using AI
★ 821+7Star change over the last 7 days - #104
Generate large-scale explorable 3D scenes with high-quality panorama videos from a single image or text prompt.
★ 816+12Star change over the last 7 days - #105
[TMLR 2025🔥] A survey for the autoregressive models in vision.
★ 808+1Star change over the last 7 days - #106
Open‑WebUI Tools is a modular toolkit designed to extend and enrich your Open WebUI instance, turning it into a powerful AI workstation. With a suite of over 15 specialized tools, function pipelines, and filters, this project supports academic research, agentic autonomy, multimodal creativity, workflows, and more
★ 807+6Star change over the last 7 days - #107
[CVPR 2024 Highlight] GenAD: Generalized Predictive Model for Autonomous Driving
★ 805+0Star change over the last 7 days - #108
rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffusion Distillation at Scale
★ 800+6Star change over the last 7 days - #109
[NeurIPS 2025 Oral]Infinity⭐️: Unified Spacetime AutoRegressive Modeling for Visual Generation
★ 784+5Star change over the last 7 days - #110
A collection of awesome video generation studies.
★ 783+1Star change over the last 7 days - #111
[ICML 2024] MagicPose(also known as MagicDance): Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion
★ 777+1Star change over the last 7 days - #112
Local-first visual generation runtime and studio for people and coding agents, with reproducible image and video workflows across multiple providers.
★ 751+32Star change over the last 7 days - #113
A Survey on Text-to-Video Generation/Synthesis.
★ 744+3Star change over the last 7 days - #114
Cosmos-Transfer2.5, built on top of Cosmos-Predict2.5, produces high-quality world simulations conditioned on multiple spatial control inputs.
★ 736+7Star change over the last 7 days - #115
[ICCV 2025] Official implementation of the paper “MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control”
★ 729+1Star change over the last 7 days - #116
A Minimalist, Batteries-included Repository for Advancing World Model Science.
★ 729+10Star change over the last 7 days - #117
Local-first, agent-native control plane for ComfyUI — MCP server + sidebar agent that generates images, video & audio, authors and runs workflows, and edits your live graph in natural language on ANY LLM (Claude, ChatGPT, Gemini, offline Ollama, or any hosted model). 178 tools, 36 AI skills, 55 installer packs. Local, LAN, VPS, or Comfy Cloud.
★ 729+35Star change over the last 7 days - #118
[ICML 2025] Official PyTorch Implementation of "History-Guided Video Diffusion"
★ 709+1Star change over the last 7 days - #119
[ICLR 2026] ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
★ 706+1Star change over the last 7 days - #120
A Curated List of Awesome Video World Models with AR Diffusion: Covering Algorithms, Applications, and Infrastructure, Aimed at Serving as a Comprehensive Resource for Researchers, Practitioners, and Enthusiasts.
★ 702+6Star change over the last 7 days