Skip to main content
buildradar
Topic · scraping

scraping

Tracked open-source repos tagged scraping, sorted by stars.

Repos
100
Total stars
698,341
Avg. stars
6,983
Share
0.04%

Topics that frequently appear alongside scraping on the same repo.

Recent risers

Repos created in the last 90 days, tagged scraping.

  • trawl@germondai

    Self-hosted scraping engine — bypasses any JS challenge & captcha: Cloudflare, Turnstile, reCAPTCHA, hCaptcha, GeeTest. FlareSolverr & Byparr alternative and drop-in replacement for your *arr stack.

    742
  • browser-search@Johell1NS

    A skill for AI agents: search the web with SearXNG, browse with Camofox, bypass protections with CloakBrowser. Anti-hallucination by design. Self-hosted, free, unlimited.

    507
  • firecrawl@firecrawl

    The context API to search, scrape, and interact with the web at scale. 🔥

    174,034+3,406Star change over the last 7 days
  • Scrapling@D4Vinci

    🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

    77,158+1,432Star change over the last 7 days
  • scrapy@scrapy

    Scrapy, a fast high-level web crawling & scraping framework for Python.

    64,098+106Star change over the last 7 days
  • maigret@soxoj

    🕵️‍♂️ Collect a dossier on a person by username from 3000+ sites

    37,143+196Star change over the last 7 days
  • Open source AI job application bot in Python: browser automation and web scraping to read job postings, then auto-apply with a tailored resume and cover letter for each posting.

    30,278+58Star change over the last 7 days
  • Scrapegraph-ai@ScrapeGraphAI

    Python scraper based on AI

    30,047+244Star change over the last 7 days
  • crawlee@apify

    Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.

    25,556+98Star change over the last 7 days
  • colly@gocolly

    Elegant Scraper and Crawler Framework for Golang

    25,488+26Star change over the last 7 days
  • Pythonic HTML Parsing for Humans™

    13,814-4Star change over the last 7 days
  • undetected-chromedriver@ultrafunkamsterdam

    Custom Selenium Chromedriver | Zero-Config | Passes ALL bot mitigation systems (like Distil / Imperva/ Datadadome / CloudFlare IUAM)

    12,814+10Star change over the last 7 days
  • webmagic@code4craft

    A scalable web crawler framework for Java.

    11,680+0Star change over the last 7 days
  • camoufox@daijro

    🦊 Anti-detect browser

    11,534+210Star change over the last 7 days
  • crawlee-python@apify

    Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.

    9,477+19Star change over the last 7 days
  • camofox-browser@jo-inc

    Stealth headless browser for AI agents — bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement.

    8,960+195Star change over the last 7 days
  • awesome-web-scraping@lorien

    List of libraries, tools and APIs for web scraping and data processing.

    8,138+10Star change over the last 7 days
  • autoscraper@alirezamika

    A Smart, Automatic, Fast and Lightweight Web Scraper for Python

    7,915+25Star change over the last 7 days
  • tabula@tabulapdf

    Tabula is a tool for liberating data tables trapped inside PDF files

    7,476+5Star change over the last 7 days
  • pydoll@autoscrape-labs

    Pydoll is a library for automating chromium-based browsers without a WebDriver, offering realistic interactions.

    7,053+16Star change over the last 7 days
  • trafilatura@adbar

    Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML

    6,728+50Star change over the last 7 days
  • 🚀 Free HTTP, SOCKS4, & SOCKS5 proxy list * Updated every 5 minutes * and rotating proxy API (100+ countries)

    6,658+57Star change over the last 7 days
  • pipet@bjesus

    Swiss-army tool for scraping and extracting data from online assets, made for hackers

    4,771+1Star change over the last 7 days
  • twikit@d60

    Twitter API Scraper | Without an API key | Twitter Internal API | Free | Twitter scraper | Twitter Bot

    4,638+11Star change over the last 7 days
  • mechanize@sparklemotion

    Mechanize is a ruby library that makes automated web interaction easy.

    4,439+0Star change over the last 7 days
  • snoop@snooppr

    Snoop — инструмент разведки на основе открытых данных (OSINT world)

    4,011+1Star change over the last 7 days
  • AnyCrawl@any4ai

    AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Native multi-threading for bulk processing.

    3,415+2Star change over the last 7 days
  • Do you want to LEARN NEW STUFF for FREE? Don't worry, with the power of web-scraping and automation, this script will find the necessary Udemy coupons & enroll you for PAID UDEMY COURSES, ABSOLUTELY FREE!

    3,291+3Star change over the last 7 days
  • facebook-scraper@kevinzg

    Scrape Facebook public pages without an API key

    3,271+4Star change over the last 7 days
  • panther@symfony

    A browser testing and web crawling library for PHP and Symfony

    3,067+2Star change over the last 7 days
  • An open source fingerprint browser based on Ungoogled Chromium. 指纹浏览器 隐私浏览器

    3,003+28Star change over the last 7 days
  • Learn step-by-step how to scrape Google Trends data and make a result comparison using Python and Oxylabs SERP API. Extract keywords, their popularity, breakdown by region, related queries, and more.

    2,874+38Star change over the last 7 days
← Back to topics