Skip to main content
buildradar
Sign in
Topic · data-extraction

data-extraction

Tracked open-source repos tagged data-extraction, sorted by stars.

Repos
28
Total stars
340,458
Avg. stars
12,159
Share
0.02%

Topics that frequently appear alongside data-extraction on the same repo.

Recent risers

Repos created in the last 90 days, tagged data-extraction.

  • ai-data-extractor@bawadou

    Free open-source extractor for AI coding assistant chat histories. Supports Claude Code, Cursor, Windsurf, Aider, Cline/Roo Code, and more.

    550
  • firecrawl@firecrawl

    The context API to search, scrape, and interact with the web at scale. 🔥

    174,034+3,406Star change over the last 7 days
  • Scrapling@D4Vinci

    🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

    77,158+1,432Star change over the last 7 days
  • Scrapegraph-ai@ScrapeGraphAI

    Python scraper based on AI

    30,047+244Star change over the last 7 days
  • maxun@getmaxun

    🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥

    17,322+70Star change over the last 7 days
  • ferret@MontFerret

    Declarative data automation language and Go runtime for structured extraction workflows.

    6,008-1Star change over the last 7 days
  • flashtext@vi3k6i5

    Extract Keywords from sentence or Replace keywords in sentences.

    5,712+0Star change over the last 7 days
  • skills@browser-act

    Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution, independent multi-session operation, isolated multi-account browsing.

    5,522+129Star change over the last 7 days
  • brightdata-mcp@brightdata

    A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.

    2,614+10Star change over the last 7 days
  • HeadlessX@saifyxpro

    The undetected self-hosted browser automation platform. Powered by Camoufox (Firefox) for 0% detection rates. Built for speed, privacy, and scalability.

    2,273+18Star change over the last 7 days
  • recipe-scrapers@hhursev

    Python package for scraping recipes data

    2,215+3Star change over the last 7 days
  • contextgem@shcherbak-ai

    ContextGem: Effortless LLM extraction from documents

    1,994+8Star change over the last 7 days
  • article-extractor@extractus

    To extract article from given URL

    1,910+2Star change over the last 7 days
  • optimus@hi-primus

    :truck: Agile Data Preparation Workflows made easy with Pandas, Dask, cuDF, Dask-cuDF, Vaex and PySpark

    1,536+0Star change over the last 7 days
  • vnstock@thinh-vu

    A beginner-friendly yet powerful Python toolkit for financial analysis and automation — built to make modern investing accessible to everyone

    1,389+10Star change over the last 7 days
  • parsera@raznem

    Lightweight library for scraping web-sites with LLMs

    1,352+0Star change over the last 7 days
  • socid-extractor@soxoj

    ⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites

    1,076+9Star change over the last 7 days
  • pdf_oxide@yfedoseev

    The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.

    1,006+37Star change over the last 7 days
  • eclaire@eclaire-labs

    Local-first, open-source AI assistant for your data. Unify tasks, notes, docs, photos, and bookmarks. Private, self-hosted, and extensible via APIs.

    916+15Star change over the last 7 days
  • crw@us

    Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.

    876+230Star change over the last 7 days
  • hacker-news-digest@polyrabbit

    :newspaper: Let ChatGPT Summarize Hacker News for You

    756+0Star change over the last 7 days
  • npm-pdfreader@adrienjoly

    🚜 Parse text and tables from PDF files.

    704+0Star change over the last 7 days
  • scrapecraft@ScrapeGraphAI

    🤖 AI-powered web scraping editor with visual workflow builder. Build, test & deploy web scrapers using natural language. Powered by ScrapeGraphAI & LangGraph.

    684+2Star change over the last 7 days
  • Google Reviews API integration and scalable Google Reviews Scraper for extracting ratings, review text, and business insights from Google Maps using a structured Google Review API.

    593-1Star change over the last 7 days
  • ai-web-scraper@ScrapingBee

    Effortlessly extract data with our AI scraper API. Simplify data extraction, get clean JSON outputs and adapt to page changes. Try it free today!

    571Star change over the last 7 days
  • github-statuses@mrshu

    The "Missing GitHub Status Page" -- a Flat Data attempt at historically documenting GitHub statuses

    561+9Star change over the last 7 days
  • Stealth-Requests@jpjacobpadilla

    Undetected web-scraping & seamless HTML parsing in Python!

    561+2Star change over the last 7 days
  • reader@vakra-dev

    Open source web infrastructure for AI. Scrape, crawl, and automate the web, clean markdown, browser sessions, ready for your agents.

    558+0Star change over the last 7 days
  • ai-data-extractor@bawadou

    Free open-source extractor for AI coding assistant chat histories. Supports Claude Code, Cursor, Windsurf, Aider, Cline/Roo Code, and more.

    550+11Star change over the last 7 days
← Back to topics