text-extraction
Tracked open-source repos tagged text-extraction, sorted by stars.
Related topics
Topics that frequently appear alongside text-extraction on the same repo.
Recent risers
Repos created in the last 90 days, tagged text-extraction.
No new repos tagged with this topic in the last 90 days.
- #1
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
★ 18,547+1,664Star change over the last 7 days - #2★ 12,235+39Star change over the last 7 days
- #3
Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
★ 9,254+21Star change over the last 7 days - #4
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
★ 6,754+26Star change over the last 7 days - #5★ 3,703+1Star change over the last 7 days
- #6★ 3,109+1Star change over the last 7 days
- #7
Tika-Python is a Python binding to the Apache Tika™ REST services allowing Tika to be called natively in the Python community.
★ 1,667+0Star change over the last 7 days - #8
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
★ 1,015+11Star change over the last 7 days - #9
Zero-copy PDF text extraction library written in Zig. High-performance, memory-mapped parsing with SIMD acceleration.
★ 921+1Star change over the last 7 days - #10
PDF to markdown using vision LLMs — tables, layouts, and structure preserved
★ 902+1Star change over the last 7 days - #11
High performance and CommonMark compliant HTML to Markdown converter. Maintained by the Kreuzberg team. Kreuzberg is a fast, polyglot document intelligence engine with a Rust core. It extracts structured data from 98+ document formats using streaming parsers and built-in OCR.
★ 863+4Star change over the last 7 days - #12★ 823-1Star change over the last 7 days
- #13★ 756+1Star change over the last 7 days
- #14★ 554+0Star change over the last 7 days
- #15★ 536+0Star change over the last 7 days