Skip to main content
buildradar
Sign in
Topic · image-captioning

image-captioning

Tracked open-source repos tagged image-captioning, sorted by stars.

Repos
12
Total stars
29,780
Avg. stars
2,482
Share
0.00%

Topics that frequently appear alongside image-captioning on the same repo.

Recent risers

Repos created in the last 90 days, tagged image-captioning.

No new repos tagged with this topic in the last 90 days.

  • LAVIS@salesforce

    LAVIS - A One-stop Library for Language-Vision Intelligence

    11,261-1Star change over the last 7 days
  • sketch-code@ashnkumar

    Keras model to generate HTML code from hand-drawn website mockups. Implements an image captioning architecture to drawn source images.

    5,143+0Star change over the last 7 days
  • InternGPT@OpenGVLab

    InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)

    3,205+0Star change over the last 7 days
  • OFA@OFA-Sys

    Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

    2,557+0Star change over the last 7 days
  • CameraManager@imaginary-cloud

    Simple Swift class to provide all the configurations you need to create custom camera view in your app

    1,395-1Star change over the last 7 days
  • taggui@jhc13

    Tag manager and captioner for image datasets

    1,348+2Star change over the last 7 days
  • prismer@NVlabs

    The implementation of "Prismer: A Vision-Language Model with Multi-Task Experts".

    1,309+0Star change over the last 7 days
  • :fire: :fire: :fire: A paper list of some recent Computer Vision(CV) works

    973+1Star change over the last 7 days
  • ZerolanLiveRobot@AkagawaTsurunaki

    AI VTuber with LLM, ASR, TTS, OCR, CV and more technologies to live stream or play Minecraft with you.

    803+4Star change over the last 7 days
  • 👁️ + 💬 + 🎧 = 🤖 Curated list of top foundation and multimodal models! [Paper + Code + Examples + Tutorials]

    637+0Star change over the last 7 days
  • ComfyUI_VLM_nodes@gokayfem

    ComfyUI nodes for vision-language models: Qwen3-VL, Moondream 3, Florence-2, SmolVLM2, InternVL, Gemma 3, MiniCPM-V. Plus open-vocabulary detection, SAM2/SAM3 segmentation, video temporal reasoning, GGUF via llama.cpp, and hosted LLM/VLM APIs.

    588-1Star change over the last 7 days
  • virtex@kdexd

    [CVPR 2021] VirTex: Learning Visual Representations from Textual Annotations

    563+0Star change over the last 7 days
← Back to topics