OpenGVLab
OpenGVLab's tracked open-source repos, sorted by stars.
- #1
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
★ 10,147+6Star change over the last 7 days - #2
[ICLR 2024] Fine-tuning LLaMA to follow Instructions within 1 Hour and 1.2M Parameters
★ 5,914+0Star change over the last 7 days - #3
[CVPR2024 Highlight][VideoChatGPT] ChatGPT with video understanding! And many more supported LMs such as miniGPT4, StableLM, and MOSS.
★ 3,347+0Star change over the last 7 days - #4
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
★ 3,205+1Star change over the last 7 days - #5
[CVPR 2023 Highlight] InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions
★ 2,841+2Star change over the last 7 days - #6
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
★ 2,370+4Star change over the last 7 days - #7★ 1,154+1Star change over the last 7 days
- #8★ 1,134+1Star change over the last 7 days
- #9
[ICLR 2026 Oral] ScaleCUA is the open-sourced computer use agents that can operate on cross-platform environments (Windows, macOS, Ubuntu, Android).
★ 1,133+3Star change over the last 7 days - #10
[ECCV2024] VideoMamba: State Space Model for Efficient Video Understanding
★ 1,125+1Star change over the last 7 days - #11
[ICLR2024 spotlight] OmniQuant is a simple and powerful quantization technique for LLMs.
★ 911+3Star change over the last 7 days - #12
[CVPR 2023] VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking
★ 816+2Star change over the last 7 days - #13★ 747+0Star change over the last 7 days
- #14
Chatbot Arena meets multi-modality! Multi-Modality Arena allows you to benchmark vision-language models side-by-side while providing images as inputs. Supports MiniGPT-4, LLaMA-Adapter V2, LLaVA, BLIP-2, and many more!
★ 566+0Star change over the last 7 days - #15
[ICLR 2025 Spotlight] Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
★ 556+0Star change over the last 7 days - #16
[ICLR2026] VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
★ 528+0Star change over the last 7 days - #17
[ICLR 2024 & ECCV 2024] The All-Seeing Projects: Towards Panoptic Visual Recognition&Understanding and General Relation Comprehension of the Open World"
★ 506+0Star change over the last 7 days