Skip to main content
buildradar
Sign in
Topic · clip

clip

Tracked open-source repos tagged clip, sorted by stars.

Repos
36
Total stars
69,310
Avg. stars
1,925
Share
0.00%

Topics that frequently appear alongside clip on the same repo.

Recent risers

Repos created in the last 90 days, tagged clip.

No new repos tagged with this topic in the last 90 days.

  • X-AnyLabeling@CVHub520

    X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.

    10,313+59Star change over the last 7 days
  • Chinese-CLIP@OFA-Sys

    Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.

    6,001+1Star change over the last 7 days
  • pushdeer@easychen

    开放源码的无App推送服务,iOS14+扫码即用。亦支持快应用/iOS和Mac客户端、Android客户端、自制设备

    5,022-1Star change over the last 7 days
  • VLMEvalKit@open-compass

    Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

    4,375+13Star change over the last 7 days
  • mmpretrain@open-mmlab

    OpenMMLab Pre-training Toolbox and Benchmark

    3,850-1Star change over the last 7 days
  • zero_nlp@yuanzhoulvpi2017

    中文nlp解决方案(大模型、数据、模型、训练、推理)

    3,836+2Star change over the last 7 days
  • VLM_survey@jingyi0000

    Collection of AWESOME vision-language models for vision tasks

    3,126+0Star change over the last 7 days
  • clip-interrogator@pharmapsychotic

    Image to prompt with BLIP and CLIP

    2,984+1Star change over the last 7 days
  • clip-retrieval@rom1504

    Easily compute clip embeddings and build a clip retrieval system with them

    2,796+1Star change over the last 7 days
  • 🥂 Gracefully face hCaptcha challenge with multimodal large language model.

    2,483-1Star change over the last 7 days
  • cambrian@cambrian-mllm

    Cambrian-1 is a family of multimodal LLMs with a vision-centric design.

    2,014+1Star change over the last 7 days
  • RWidgetHelper@RuffianZhong

    Android UI 快速开发,专治原生控件各种不服

    1,958+0Star change over the last 7 days
  • Must-have resource for anyone who wants to experiment with and build on the OpenAI vision API 🔥

    1,690+0Star change over the last 7 days
  • clearcam@roryclear

    Add object detection, tracking, mobile notifications, and search to any security camera.

    1,557+255Star change over the last 7 days
  • Video-ChatGPT@mbzuai-oryx

    [ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.

    1,507+1Star change over the last 7 days
  • UForm@unum-cloud

    Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️

    1,247+1Star change over the last 7 days
  • Awesome-CLIP@yzhuoning

    Awesome list for research on CLIP (Contrastive Language-Image Pre-Training).

    1,227+0Star change over the last 7 days
  • vlms-zero-to-hero@SkalskiP

    This series will take you on a journey from the fundamentals of NLP and Computer Vision to the cutting edge of Vision-Language Models.

    1,179+0Star change over the last 7 days
  • CLIP4Clip@ArrowLuo

    An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"

    1,031+0Star change over the last 7 days
  • SmartJavaAI@geekwenjie

    🔥🔥🔥Java免费离线AI算法工具箱,支持人脸识别,活体检测,表情识别、目标检测、实例分割、行人检测、OCR文字识别、车牌识别、表格识别、ASR+TTS、机器翻译等功能,Maven引用即可使用。支持PyTorch、Tensorflow,已集成 Mtcnn、InsightFace、SeetaFace6、YOLOv8~v12、PaddleOCR(PPOCRv5)、Whisper等主流模型

    908+3Star change over the last 7 days
  • aphantasia@eps696

    CLIP + FFT/DWT/RGB = text to image/video

    789+0Star change over the last 7 days
  • FG-CLIP@360CVGroup

    New generation of CLIP with strong fine grained discrimination capability, ICML2026 and ICML2025

    733+2Star change over the last 7 days
  • PaddleMIX@PaddlePaddle

    Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high performance and flexibility.

    723-1Star change over the last 7 days
  • TrailSnap@LC044

    TrailSnap (行影集) | AI-Powered open-source photo album for travel & life memories.(AI赋能的开源相册工具,珍藏旅行与生活点滴)

    709+13Star change over the last 7 days
  • A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.

    706+3Star change over the last 7 days
  • LLM2CLIP@microsoft

    LLM2CLIP significantly improves already state-of-the-art CLIP models.

    688+1Star change over the last 7 days
  • TwitchLink@devhotteok

    Twitch Stream & Video & Clip Downloader/Recorder. This GUI downloader helps you download and record Twitch videos, including broadcasts and VODs.

    681+4Star change over the last 7 days
  • video_features@v-iashin

    Extract video features from raw videos using multiple GPUs. We support RAFT flow frames as well as S3D, I3D, R(2+1)D, VGGish, CLIP, and TIMM models.

    655+0Star change over the last 7 days
  • 👁️ + 💬 + 🎧 = 🤖 Curated list of top foundation and multimodal models! [Paper + Code + Examples + Tutorials]

    637+0Star change over the last 7 days
  • Keras beit,caformer,CMT,CoAtNet,convnext,davit,dino,efficientdet,edgenext,efficientformer,efficientnet,eva,fasternet,fastervit,fastvit,flexivit,gcvit,ghostnet,gpvit,hornet,hiera,iformer,inceptionnext,lcnet,levit,maxvit,mobilevit,moganet,nat,nfnets,pvt,swin,tinynet,tinyvit,uniformer,volo,vanillanet,yolor,yolov7,yolov8,yolox,gpt2,llama2, alias kecam

    626+0Star change over the last 7 days
← Back to topics