clip
Tracked open-source repos tagged clip, sorted by stars.
Related topics
Topics that frequently appear alongside clip on the same repo.
Recent risers
Repos created in the last 90 days, tagged clip.
No new repos tagged with this topic in the last 90 days.
- #1
X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.
★ 10,313+59Star change over the last 7 days - #2
Chinese version of CLIP which achieves Chinese cross-modal retrieval and representation generation.
★ 6,001+1Star change over the last 7 days - #3★ 5,022-1Star change over the last 7 days
- #4
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4,375+13Star change over the last 7 days - #5
OpenMMLab Pre-training Toolbox and Benchmark
★ 3,850-1Star change over the last 7 days - #6★ 3,836+2Star change over the last 7 days
- #7
Collection of AWESOME vision-language models for vision tasks
★ 3,126+0Star change over the last 7 days - #8
Image to prompt with BLIP and CLIP
★ 2,984+1Star change over the last 7 days - #9
Easily compute clip embeddings and build a clip retrieval system with them
★ 2,796+1Star change over the last 7 days - #10
🥂 Gracefully face hCaptcha challenge with multimodal large language model.
★ 2,483-1Star change over the last 7 days - #11★ 2,014+1Star change over the last 7 days
- #12
Android UI 快速开发,专治原生控件各种不服
★ 1,958+0Star change over the last 7 days - #13
Must-have resource for anyone who wants to experiment with and build on the OpenAI vision API 🔥
★ 1,690+0Star change over the last 7 days - #14
Add object detection, tracking, mobile notifications, and search to any security camera.
★ 1,557+255Star change over the last 7 days - #15
[ACL 2024 🔥] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capabilities of LLMs with a pretrained visual encoder adapted for spatiotemporal video representation. We also introduce a rigorous 'Quantitative Evaluation Benchmarking' for video-based conversational models.
★ 1,507+1Star change over the last 7 days - #16
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
★ 1,247+1Star change over the last 7 days - #17
Awesome list for research on CLIP (Contrastive Language-Image Pre-Training).
★ 1,227+0Star change over the last 7 days - #18
This series will take you on a journey from the fundamentals of NLP and Computer Vision to the cutting edge of Vision-Language Models.
★ 1,179+0Star change over the last 7 days - #19
An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"
★ 1,031+0Star change over the last 7 days - #20
🔥🔥🔥Java免费离线AI算法工具箱,支持人脸识别,活体检测,表情识别、目标检测、实例分割、行人检测、OCR文字识别、车牌识别、表格识别、ASR+TTS、机器翻译等功能,Maven引用即可使用。支持PyTorch、Tensorflow,已集成 Mtcnn、InsightFace、SeetaFace6、YOLOv8~v12、PaddleOCR(PPOCRv5)、Whisper等主流模型
★ 908+3Star change over the last 7 days - #21
CLIP + FFT/DWT/RGB = text to image/video
★ 789+0Star change over the last 7 days - #22
New generation of CLIP with strong fine grained discrimination capability, ICML2026 and ICML2025
★ 733+2Star change over the last 7 days - #23
Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high performance and flexibility.
★ 723-1Star change over the last 7 days - #24
TrailSnap (行影集) | AI-Powered open-source photo album for travel & life memories.(AI赋能的开源相册工具,珍藏旅行与生活点滴)
★ 709+13Star change over the last 7 days - #25
A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.
★ 706+3Star change over the last 7 days - #26★ 688+1Star change over the last 7 days
- #27
Twitch Stream & Video & Clip Downloader/Recorder. This GUI downloader helps you download and record Twitch videos, including broadcasts and VODs.
★ 681+4Star change over the last 7 days - #28
Extract video features from raw videos using multiple GPUs. We support RAFT flow frames as well as S3D, I3D, R(2+1)D, VGGish, CLIP, and TIMM models.
★ 655+0Star change over the last 7 days - #29
👁️ + 💬 + 🎧 = 🤖 Curated list of top foundation and multimodal models! [Paper + Code + Examples + Tutorials]
★ 637+0Star change over the last 7 days - #30
Keras beit,caformer,CMT,CoAtNet,convnext,davit,dino,efficientdet,edgenext,efficientformer,efficientnet,eva,fasternet,fastervit,fastvit,flexivit,gcvit,ghostnet,gpvit,hornet,hiera,iformer,inceptionnext,lcnet,levit,maxvit,mobilevit,moganet,nat,nfnets,pvt,swin,tinynet,tinyvit,uniformer,volo,vanillanet,yolor,yolov7,yolov8,yolox,gpt2,llama2, alias kecam
★ 626+0Star change over the last 7 days