ocr
Tracked open-source repos tagged ocr, sorted by stars.
- #61
A Repo For Document AI
★ 3,257+9Star change over the last 7 days - #62
An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.
★ 3,229+3Star change over the last 7 days - #63
[验证码识别-训练] This project is based on CNN/ResNet/DenseNet+GRU/LSTM+CTC/CrossEntropy to realize verification code identification. This project is only for training the model.
★ 3,213+0Star change over the last 7 days - #64
Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown
★ 3,177+0Star change over the last 7 days - #65
Go package for OCR (Optical Character Recognition), by using Tesseract C++ library
★ 3,134+2Star change over the last 7 days - #66★ 3,065+1Star change over the last 7 days
- #67
A wrapper to work with Tesseract OCR inside PHP.
★ 3,039-2Star change over the last 7 days - #68
A curated collection of practical AI projects implementing OCR systems, RAG, AI agents, and other AI use cases.
★ 3,030+58Star change over the last 7 days - #69
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs
★ 2,996+3Star change over the last 7 days - #70
A Model Context Protocol server for converting almost anything to Markdown
★ 2,983-1Star change over the last 7 days - #71
End-to-end Chinese scene-text detection and recognition with CTPN, CRNN, and CTC (legacy project).
★ 2,955+0Star change over the last 7 days - #72
OpenRecall is a fully open-source, privacy-first alternative to proprietary solutions like Microsoft's Windows Recall. With OpenRecall, you can easily access your digital history, enhancing your memory and productivity without compromising your privacy.
★ 2,937+2Star change over the last 7 days - #73
Open Source Document Management System for Digital Archives (Scanned Documents)
★ 2,935+1Star change over the last 7 days - #74
AI comic and manga translator app/browser extension for automatically translating comics, manga, manhwa, BDs, fumetti, and more in multiple languages and formats (Images, PDF, EPUB, CBR, CBZ etc).
★ 2,928+12Star change over the last 7 days - #75
Optical character recognition for Japanese text, with the main focus being Japanese manga
★ 2,770+7Star change over the last 7 days - #76
基于manga-image-translator 实现的开源漫画AI翻译桌面工具。支持日、韩、英文漫画自动处理,集成OpenAl、Gemini等多翻译引擎;实现OCR文字检测、原文擦除、AI翻译、图像修复、译文排版完整链路,自带可视化编辑器,支持自定义文本样式,一键部署开箱即用。
★ 2,711+35Star change over the last 7 days - #77★ 2,705+4Star change over the last 7 days
- #78
Use LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI
★ 2,662+8Star change over the last 7 days - #79
Lightweight document management system packed with all the features you can expect from big expensive solutions
★ 2,561+0Star change over the last 7 days - #80
Open-source infrastructure and data orchestration platform for risk decisioning
★ 2,427+1Star change over the last 7 days - #81
The official CLI for Agentic Document Extraction (ADE) by LandingAI — parse documents and extract schema-shaped data from your terminal
★ 2,410+1Star change over the last 7 days - #82
LLM Agent Framework in ComfyUI includes MCP sever, Omost,GPT-sovits, ChatTTS,GOT-OCR2.0, and FLUX prompt nodes,access to Feishu,discord,and adapts to all llms with similar openai / aisuite interfaces, such as o1,ollama, gemini, grok, qwen, GLM, deepseek, kimi,doubao. Adapted to local llms, vlm, gguf such as llama-3.3 Janus-Pro, Linkage graphRAG
★ 2,348+5Star change over the last 7 days - #83
Assist in organizing your piles of documents, resulting from scanners, e-mails and other sources with miminal effort.
★ 2,317+2Star change over the last 7 days - #84
A tool to convert a Wallpaper's color scheme / palette, OCR with VLM's Traditional & Hybrid, Image Compression ,color palette extraction, image upsacling with Adversarial Networks and more image processing features.
★ 2,303+1Star change over the last 7 days - #85
在保留版面、公式与结构的前提下进行 PDF 翻译,适用于科研与技术文档
★ 2,234+8Star change over the last 7 days - #86★ 2,184+0Star change over the last 7 days
- #87
Rust multi‑backend OCR/VLM engine (DeepSeek‑OCR-1/2, PaddleOCR‑VL, DotsOCR) with DSQ quantization and an OpenAI‑compatible server & CLI – run locally without Python.
★ 2,184+1Star change over the last 7 days - #88★ 2,171+0Star change over the last 7 days
- #89
A search engine that "just works" for Obsidian. Supports OCR and PDF indexing.
★ 2,131+2Star change over the last 7 days - #90
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
★ 2,086-1Star change over the last 7 days