Skip to main content
buildradar
Sign in
Topic · document-processing

document-processing

Tracked open-source repos tagged document-processing, sorted by stars.

Repos
13
Total stars
58,566
Avg. stars
4,505
Share
0.00%

Topics that frequently appear alongside document-processing on the same repo.

Recent risers

Repos created in the last 90 days, tagged document-processing.

  • yoji@wangxijie001

    有情绪的 AI 桌面伴侣 | 隐私优先 · 语音唤醒 · 协助办公 · MCP 无限扩展 | An AI desktop companion with emotions — privacy-first, voice wake, office assistance, and MCP extensibility

    704
  • book-to-skill@virgiliojr94

    Turn any technical book PDF into a Claude Code skill — ready to study, reference, and use while you work.

    27,016+3,296Star change over the last 7 days
  • liteparse@run-llama

    A fast, helpful, and open-source document parser

    12,196+46Star change over the last 7 days
  • SenseNova-Skills@OpenSenseNova

    Modular SenseNova skills for building AI-powered office assistants and productivity workflows

    5,173+247Star change over the last 7 days
  • docetl@ucbepic

    A system for agentic LLM-powered data processing and ETL

    4,040+54Star change over the last 7 days
  • retain-pdf@wxyhgk

    在保留版面、公式与结构的前提下进行 PDF 翻译,适用于科研与技术文档

    2,225+20Star change over the last 7 days
  • ExtractThinker@enoch3712

    ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.

    1,595+2Star change over the last 7 days
  • OpenOCR@Topdu

    OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.

    1,440+3Star change over the last 7 days
  • pdf_oxide@yfedoseev

    The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.

    1,006+37Star change over the last 7 days
  • eclaire@eclaire-labs

    Local-first, open-source AI assistant for your data. Unify tasks, notes, docs, photos, and bookmarks. Private, self-hosted, and extensible via APIs.

    916+15Star change over the last 7 days
  • pdf-reader-mcp@SylphxAI

    Give your AI agent eyes for PDFs — structured text, tables, OCR, visual evidence, and page-level citations via MCP. Native Rust, local-first.

    906+8Star change over the last 7 days
  • docling-graph@docling-project

    Transform unstructured documents into validated, rich and queryable knowledge graphs.

    845+91Star change over the last 7 days
  • yoji@wangxijie001

    有情绪的 AI 桌面伴侣 | 隐私优先 · 语音唤醒 · 协助办公 · MCP 无限扩展 | An AI desktop companion with emotions — privacy-first, voice wake, office assistance, and MCP extensibility

    704+49Star change over the last 7 days
  • ExtractPDF4J@ExtractPDF4J

    Java PDF table extraction & OCR library. Extract structured tables from text-based and scanned PDFs using stream, lattice (OpenCV-style grid detection), and hybrid parsing.

    517-27Star change over the last 7 days
← Back to topics