FoundationVision
FoundationVision's tracked open-source repos, sorted by stars.
- #1
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
★ 8,728-1Star change over the last 7 days - #2
[ECCV 2022] ByteTrack: Multi-Object Tracking by Associating Every Detection Box
★ 6,662+8Star change over the last 7 days - #3
Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation
★ 1,966+0Star change over the last 7 days - #4
[CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
★ 1,587+0Star change over the last 7 days - #5
[CVPR2024 Highlight]GLEE: General Object Foundation Model for Images and Videos at Scale
★ 1,170+0Star change over the last 7 days - #6
Industry-level video foundation model for unified Text-to-Video (T2V) and Image-to-Video (I2V) generation.
★ 952+0Star change over the last 7 days - #7
[NeurIPS 2025 Oral]Infinity⭐️: Unified Spacetime AutoRegressive Modeling for Visual Generation
★ 783+3Star change over the last 7 days - #8
(Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generators
★ 640+0Star change over the last 7 days - #9
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
★ 586+0Star change over the last 7 days - #10
[NeurIPS 2025 Spotlight] A Unified Tokenizer for Visual Generation and Understanding
★ 530+0Star change over the last 7 days