vision-language
Tracked open-source repos tagged vision-language, sorted by stars.
Related topics
Topics that frequently appear alongside vision-language on the same repo.
Recent risers
Repos created in the last 90 days, tagged vision-language.
No new repos tagged with this topic in the last 90 days.
- #1
[ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
★ 10,522+13Star change over the last 7 days - #2
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
★ 6,001+7Star change over the last 7 days - #3
Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
★ 2,557+0Star change over the last 7 days - #4
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
★ 1,961+1Star change over the last 7 days - #5
A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.
★ 1,835-1Star change over the last 7 days - #6
[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.
★ 1,506+0Star change over the last 7 days - #7
日本語LLMまとめ - Overview of Japanese LLMs
★ 1,428+1Star change over the last 7 days - #8★ 1,340+0Star change over the last 7 days
- #9
A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
★ 1,060+0Star change over the last 7 days - #10
A Framework of Small-scale Large Multimodal Models
★ 1,004-1Star change over the last 7 days - #11
CALVIN - A benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
★ 981+5Star change over the last 7 days - #12★ 874-1Star change over the last 7 days
- #13
This is the third party implementation of the paper Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.
★ 846+0Star change over the last 7 days - #14
🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)
★ 842-1Star change over the last 7 days - #15
[ICLR 2024] Controlling Vision-Language Models for Universal Image Restoration. 5th place in the NTIRE 2024 Restore Any Image Model in the Wild Challenge.
★ 817+0Star change over the last 7 days - #16★ 642+0Star change over the last 7 days
- #17
🛰️ Official repository of paper "RemoteCLIP: A Vision Language Foundation Model for Remote Sensing" (IEEE TGRS)
★ 590+0Star change over the last 7 days