Skip to main content
buildradar
Sign in
Topic · text-extraction

text-extraction

Tracked open-source repos tagged text-extraction, sorted by stars.

Repos
15
Total stars
61,632
Avg. stars
4,109
Share
0.00%

Topics that frequently appear alongside text-extraction on the same repo.

Recent risers

Repos created in the last 90 days, tagged text-extraction.

No new repos tagged with this topic in the last 90 days.

  • pdf-inspector@firecrawl

    Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.

    18,547+1,664Star change over the last 7 days
  • liteparse@run-llama

    A fast, helpful, and open-source document parser

    12,235+39Star change over the last 7 days
  • xberg@xberg-io

    Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.

    9,254+21Star change over the last 7 days
  • trafilatura@adbar

    Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML

    6,754+26Star change over the last 7 days
  • sumy@miso-belica

    Module for automatic summarization of text documents and HTML pages.

    3,703+1Star change over the last 7 days
  • unipdf@unidoc

    Golang PDF library for creating and processing PDF files (pure go)

    3,109+1Star change over the last 7 days
  • tika-python@chrismattmann

    Tika-Python is a Python binding to the Apache Tika™ REST services allowing Tika to be called natively in the Python community.

    1,667+0Star change over the last 7 days
  • pdf_oxide@yfedoseev

    The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.

    1,015+11Star change over the last 7 days
  • zpdf@Lulzx

    Zero-copy PDF text extraction library written in Zig. High-performance, memory-mapped parsing with SIMD acceleration.

    921+1Star change over the last 7 days
  • api-llm-ocr@yigitkonur

    PDF to markdown using vision LLMs — tables, layouts, and structure preserved

    902+1Star change over the last 7 days
  • html-to-markdown@xberg-io

    High performance and CommonMark compliant HTML to Markdown converter. Maintained by the Kreuzberg team. Kreuzberg is a fast, polyglot document intelligence engine with a Rust core. It extracts structured data from 98+ document formats using streaming parsers and built-in OCR.

    863+4Star change over the last 7 days
  • jusText@miso-belica

    Heuristic based boilerplate removal tool

    823-1Star change over the last 7 days
  • datashare@ICIJ

    A self‑hosted search engine for documents

    756+1Star change over the last 7 days
  • pdftools@ropensci

    Text Extraction, Rendering and Converting of PDF Documents

    554+0Star change over the last 7 days
  • srt@cdown

    A simple library and set of tools for parsing, modifying, and composing SRT files.

    536+0Star change over the last 7 days
← Back to topics