data-extraction
Tracked open-source repos tagged data-extraction, sorted by stars.
Related topics
Topics that frequently appear alongside data-extraction on the same repo.
Recent risers
Repos created in the last 90 days, tagged data-extraction.
- #1
Free open-source extractor for AI coding assistant chat histories. Supports Claude Code, Cursor, Windsurf, Aider, Cline/Roo Code, and more.
★ 550
- #1★ 174,034+3,406Star change over the last 7 days
- #2
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
★ 77,158+1,432Star change over the last 7 days - #3
Python scraper based on AI
★ 30,047+244Star change over the last 7 days - #4
🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥
★ 17,322+70Star change over the last 7 days - #5
Declarative data automation language and Go runtime for structured extraction workflows.
★ 6,008-1Star change over the last 7 days - #6★ 5,712+0Star change over the last 7 days
- #7
Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution, independent multi-session operation, isolated multi-account browsing.
★ 5,522+129Star change over the last 7 days - #8
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
★ 2,614+10Star change over the last 7 days - #9
The undetected self-hosted browser automation platform. Powered by Camoufox (Firefox) for 0% detection rates. Built for speed, privacy, and scalability.
★ 2,273+18Star change over the last 7 days - #10
Python package for scraping recipes data
★ 2,215+3Star change over the last 7 days - #11
ContextGem: Effortless LLM extraction from documents
★ 1,994+8Star change over the last 7 days - #12
To extract article from given URL
★ 1,910+2Star change over the last 7 days - #13
:truck: Agile Data Preparation Workflows made easy with Pandas, Dask, cuDF, Dask-cuDF, Vaex and PySpark
★ 1,536+0Star change over the last 7 days - #14
A beginner-friendly yet powerful Python toolkit for financial analysis and automation — built to make modern investing accessible to everyone
★ 1,389+10Star change over the last 7 days - #15★ 1,352+0Star change over the last 7 days
- #16
⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
★ 1,076+9Star change over the last 7 days - #17
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
★ 1,006+37Star change over the last 7 days - #18
Local-first, open-source AI assistant for your data. Unify tasks, notes, docs, photos, and bookmarks. Private, self-hosted, and extensible via APIs.
★ 916+15Star change over the last 7 days - #19
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
★ 876+230Star change over the last 7 days - #20
:newspaper: Let ChatGPT Summarize Hacker News for You
★ 756+0Star change over the last 7 days - #21
🚜 Parse text and tables from PDF files.
★ 704+0Star change over the last 7 days - #22
🤖 AI-powered web scraping editor with visual workflow builder. Build, test & deploy web scrapers using natural language. Powered by ScrapeGraphAI & LangGraph.
★ 684+2Star change over the last 7 days - #23
Google Reviews API integration and scalable Google Reviews Scraper for extracting ratings, review text, and business insights from Google Maps using a structured Google Review API.
★ 593-1Star change over the last 7 days - #24
Effortlessly extract data with our AI scraper API. Simplify data extraction, get clean JSON outputs and adapt to page changes. Try it free today!
★ 571—Star change over the last 7 days - #25
The "Missing GitHub Status Page" -- a Flat Data attempt at historically documenting GitHub statuses
★ 561+9Star change over the last 7 days - #26
Undetected web-scraping & seamless HTML parsing in Python!
★ 561+2Star change over the last 7 days - #27
Open source web infrastructure for AI. Scrape, crawl, and automate the web, clean markdown, browser sessions, ready for your agents.
★ 558+0Star change over the last 7 days - #28
Free open-source extractor for AI coding assistant chat histories. Supports Claude Code, Cursor, Windsurf, Aider, Cline/Roo Code, and more.
★ 550+11Star change over the last 7 days