Skip to main content
buildradar
Sign in
Topic · vision

vision

Tracked open-source repos tagged vision, sorted by stars.

Repos
33
Total stars
238,709
Avg. stars
7,234
Share
0.01%

Topics that frequently appear alongside vision on the same repo.

Recent risers

Repos created in the last 90 days, tagged vision.

  • 为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

    1,136
  • dsh-vision-router@ysr666

    Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.

    1,077
  • LibreChat@danny-avila

    Enhanced ChatGPT Clone: Features Agents, MCP, Skills, DeepSeek, Anthropic, AWS, OpenAI, Responses API, Azure, Groq, o1, GPT-5, Mistral, OpenRouter, Vertex AI, Gemini, Artifacts, AI model switching, message search, Code Interpreter, langchain, DALL-E-3, OpenAPI Actions, Functions, Secure Multi-User Auth, Presets, open-source for self-hosting. Active

    42,743+146Star change over the last 7 days
  • Xray, Penetrates Everything. Also the best v2ray-core. Where the magic happens. An open platform for various uses.

    41,371+89Star change over the last 7 days
  • UI-TARS-desktop@bytedance

    The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra

    38,817+75Star change over the last 7 days
  • caffe@BVLC

    Caffe: a fast open framework for deep learning.

    34,553-2Star change over the last 7 days
  • skyvern@Skyvern-AI

    Automate browser based workflows with AI

    22,915+37Star change over the last 7 days
  • PixelRAG@StarTrail-org

    https://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/

    9,848+60Star change over the last 7 days
  • 📸 A powerful, high-performance React Native Camera library.

    9,599+13Star change over the last 7 days
  • modlens@liustack

    The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。

    3,839+83Star change over the last 7 days
  • SimpleMem@aiming-lab

    SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal

    3,735+9Star change over the last 7 days
  • donkeycar@autorope

    Open source hardware and software platform to build a small scale self driving car.

    3,500+4Star change over the last 7 days
  • Torch-Pruning@VainF

    [CVPR 2023] DepGraph: Towards Any Structural Pruning; LLMs, Vision Foundation Models, etc.

    3,351+3Star change over the last 7 days
  • SimpleCV@sightmachine

    The Open Source Framework for Machine Vision

    2,730+0Star change over the last 7 days
  • NextLevel@NextLevel

    ⬆️ Media Capture in Swift

    2,331-1Star change over the last 7 days
  • java-docs-samples@GoogleCloudPlatform

    Java and Kotlin Code samples used on cloud.google.com

    1,907+1Star change over the last 7 days
  • ha-llmvision@valentinfrlch

    Visual intelligence for your home.

    1,454+10Star change over the last 7 days
  • aravis@AravisProject

    A vision library for genicam based cameras

    1,272+3Star change over the last 7 days
  • Deep-Learning-Experiments@roatienza

    Videos, notes and experiments to understand deep learning

    1,198+0Star change over the last 7 days
  • MLKit@jenly1314

    🌝 MLKit是一个强大易用的工具包。通过ML Kit您可以很轻松的实现文字识别、条码识别、图像标记、人脸检测、对象检测等功能。

    1,171+3Star change over the last 7 days
  • 为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

    1,136+18Star change over the last 7 days
  • dsh-vision-router@ysr666

    Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.

    1,077+54Star change over the last 7 days
  • mlp-mixer-pytorch@lucidrains

    An All-MLP solution for Vision, from Google AI

    1,064+0Star change over the last 7 days
  • calvin@mees

    CALVIN - A benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks

    983+8Star change over the last 7 days
  • CudaSift@Celebrandil

    A CUDA implementation of SIFT for NVidia GPUs (1.2 ms on a GTX 1060)

    953+0Star change over the last 7 days
  • iowncode@anupamchugh

    A curated collection of iOS, ML, AR resources sprinkled with some UI additions

    911+0Star change over the last 7 days
  • 3dmatch-toolbox@andyzeng

    3DMatch - a 3D ConvNet-based local geometric descriptor for aligning 3D meshes and point clouds.

    905+0Star change over the last 7 days
  • caer@jasmcaus

    High-performance Vision library in Python. Scale your research, not boilerplate.

    816+0Star change over the last 7 days
  • 🎰 A curated list of machine learning resources, preferably CoreML

    813+0Star change over the last 7 days
  • interpreter@bquenin

    This application can translate text captured from any application running on your computer.

    769+18Star change over the last 7 days
  • design2code@mostafasadeghi97

    Convert any web design screenshot to clean HTML/CSS code

    681+1Star change over the last 7 days
  • myvision@OvidijusParsiunas

    Computer vision based ML training data generation tool :rocket:

    612+0Star change over the last 7 days
← Back to topics