Skip to main content
buildradar
Sign in
Topic · document-analysis

document-analysis

Tracked open-source repos tagged document-analysis, sorted by stars.

Repos
12
Total stars
103,717
Avg. stars
8,643
Share
0.01%

Topics that frequently appear alongside document-analysis on the same repo.

Recent risers

Repos created in the last 90 days, tagged document-analysis.

No new repos tagged with this topic in the last 90 days.

  • MinerU@opendatalab

    Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

    78,747+564Star change over the last 7 days
  • Dolphin@bytedance

    The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

    9,049-1Star change over the last 7 days
  • docetl@ucbepic

    A system for agentic LLM-powered data processing and ETL

    4,040+54Star change over the last 7 days
  • PdfPig@UglyToad

    Read and extract text and other content from PDFs in C# (port of PDFBox)

    2,552+5Star change over the last 7 days
  • docext@NanoNets

    An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)

    2,086+4Star change over the last 7 days
  • AdvancedLiterateMachinery@AlibabaResearch

    A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team in the Language Technology Lab, Tongyi Lab, Alibaba Group.

    1,834-1Star change over the last 7 days
  • documind@DocumindHQ

    Open-source platform for extracting structured data from documents using AI.

    1,520-1Star change over the last 7 days
  • OpenOCR@Topdu

    OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.

    1,440+3Star change over the last 7 days
  • dedoc@ispras

    Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical structure, tables, and meta information from textual electronic documents. (Parse document; Document content extraction; Logical structure extraction; PDF parser; Scanned document parser; DOCX parser; HTML parser

    732+13Star change over the last 7 days
  • MinerU-Diffusion@opendatalab

    [ECCV 2026] A diffusion-based framework for document OCR that replaces autoregressive decoding with block-level parallel diffusion decoding.

    615+1Star change over the last 7 days
  • PICK-pytorch@wenwenyu

    Code for the paper "PICK: Processing Key Information Extraction from Documents using Improved Graph Learning-Convolutional Networks" (ICPR 2020)

    570+0Star change over the last 7 days
  • assemblyline@CybercentreCanada

    AssemblyLine 4: File triage and malware analysis

    532+3Star change over the last 7 days
← Back to topics