Skip to main content
buildradar
Sign in
Topic · vision-language

vision-language

Tracked open-source repos tagged vision-language, sorted by stars.

Repos
17
Total stars
34,788
Avg. stars
2,046
Share
0.00%

Topics that frequently appear alongside vision-language on the same repo.

Recent risers

Repos created in the last 90 days, tagged vision-language.

No new repos tagged with this topic in the last 90 days.

  • GroundingDINO@IDEA-Research

    [ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"

    10,522+13Star change over the last 7 days
  • Chinese-CLIP@OFA-Sys

    Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.

    6,001+7Star change over the last 7 days
  • OFA@OFA-Sys

    Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

    2,557+0Star change over the last 7 days
  • An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.

    1,961+1Star change over the last 7 days
  • AdvancedLiterateMachinery@AlibabaResearch

    A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.

    1,835-1Star change over the last 7 days
  • Video-ChatGPT@mbzuai-oryx

    [ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

    1,506+0Star change over the last 7 days
  • awesome-japanese-llm@llm-jp

    日本語LLMまとめ - Overview of Japanese LLMs

    1,428+1Star change over the last 7 days
  • DriveLM@OpenDriveLab

    [ECCV 2024 Oral] DriveLM: Driving with Graph Visual Question Answering

    1,340+0Star change over the last 7 days
  • ONE-PEACE@OFA-Sys

    A general representation model across vision, audio, language modalities. Paper: ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

    1,060+0Star change over the last 7 days
  • TinyLLaVA_Factory@TinyLLaVA

    A Framework of Small-scale Large Multimodal Models

    1,004-1Star change over the last 7 days
  • calvin@mees

    CALVIN - A benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks

    981+5Star change over the last 7 days
  • AlphaCLIP@SunzeY

    [CVPR 2024] Alpha-CLIP: A CLIP Model Focusing on Wherever You Want

    874-1Star change over the last 7 days
  • Open-GroundingDino@longzw1997

    This is the third party implementation of the paper Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

    846+0Star change over the last 7 days
  • LLaVA-pp@mbzuai-oryx

    🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)

    842-1Star change over the last 7 days
  • daclip-uir@Algolzw

    [ICLR 2024] Controlling Vision-Language Models for Universal Image Restoration. 5th place in the NTIRE 2024 Restore Any Image Model in the Wild Challenge.

    817+0Star change over the last 7 days
  • SEED@AILab-CVC

    Official implementation of SEED-LLaMA (ICLR 2024).

    642+0Star change over the last 7 days
  • RemoteCLIP@ChenDelong1999

    🛰️ Official repository of paper "RemoteCLIP: A Vision Language Foundation Model for Remote Sensing" (IEEE TGRS)

    590+0Star change over the last 7 days
← Back to topics