table-extraction
Tracked open-source repos tagged table-extraction, sorted by stars.
Related topics
Topics that frequently appear alongside table-extraction on the same repo.
Recent risers
Repos created in the last 90 days, tagged table-extraction.
No new repos tagged with this topic in the last 90 days.
- #1
Plumb a PDF for detailed information about each char, rectangle, line, et cetera — and easily extract text and tables.
★ 10,701+15Star change over the last 7 days - #2
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
★ 10,602+67Star change over the last 7 days - #3
Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
★ 9,233+48Star change over the last 7 days - #4
Table Transformer (TATR) is a deep learning model for extracting tables from unstructured documents (PDFs and images). This is also the official repository for the PubTables-1M dataset and GriTS evaluation metric.
★ 2,943+2Star change over the last 7 days - #5
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
★ 2,086+4Star change over the last 7 days - #6
PDF to markdown using vision LLMs — tables, layouts, and structure preserved
★ 901-2Star change over the last 7 days - #7
img2table is a table identification and extraction Python Library for PDF and images, based on OpenCV image processing
★ 894+2Star change over the last 7 days - #8
ParseBench - A Document Parsing Benchmark for AI Agents
★ 558+2Star change over the last 7 days