Skip to main content
buildradar
Sign in
Topic · web-scraping

web-scraping

Tracked open-source repos tagged web-scraping, sorted by stars.

119 repos
  • finvizfinance@lit26

    Finviz analysis python library.

    1,624+0Star change over the last 7 days
  • DAT8@justmarkham

    General Assembly's 2015 Data Science course in Washington, DC

    1,620+0Star change over the last 7 days
  • Scweet@Altimis

    Scrape tweets, profiles, followers and following from Twitter/X, no API key needed. Python library with smart multi-account pooling, proxy support and async.

    1,607+2Star change over the last 7 days
  • ai-map-py@oxylabs

    AI Map is an AI-powered website mapping tool by Oxylabs AI Studio that uses natural language prompts to intelligently discover and extract relevant URLs from any website.

    1,598+41Star change over the last 7 days
  • single-file-cli@gildas-lormeau

    CLI tool for saving a faithful copy of a complete web page in a single HTML file (based on SingleFile)

    1,579+9Star change over the last 7 days
  • rvest@tidyverse

    Simple web scraping for R

    1,520+0Star change over the last 7 days
  • patchright-python@Kaliiiiiiiiii-Vinyzu

    Undetected Python version of the Playwright testing and automation library.

    1,492+5Star change over the last 7 days
  • fredy@orangecoding

    ❤️ Fredy - [F]ind [R]eal [E]state [D]amn Eas[y] - Fredy keeps searching for new apartments, houses, and flats in Europe on platforms like ImmoScout24, Immowelt, eBay Kleinanzeigen and instantly delivers the results to you via Slack, Telegram, Email, Discord or ntfy, so you can focus on the more important things in life ;)

    1,462+11Star change over the last 7 days
  • core@roach-php

    The complete web scraping toolkit for PHP.

    1,455+0Star change over the last 7 days
  • agentql@tinyfish-io

    AgentQL is a suite of tools for connecting your AI to the web. Featuring a query language and Playwright integrations for interacting with elements and extracting data quickly, precisely, and at scale. Includes REST API, Python and JavaScript SDKs, browser debugger.

    1,454+0Star change over the last 7 days
  • wreq-python@0x676e67

    An ergonomic, privacy-aware Python HTTP Client

    1,438+1Star change over the last 7 days
  • rebrowser-patches@rebrowser

    Collection of patches for puppeteer and playwright to avoid automation detection and leaks. Helps to avoid Cloudflare and DataDome CAPTCHA pages. Easy to patch/unpatch, can be enabled/disabled on demand.

    1,422-1Star change over the last 7 days
  • Use Web Scraper API to extract data from Google Finance, including stock titles, pricing, and price changes in percentages.

    1,403+35Star change over the last 7 days
  • open-scouts@firecrawl

    🔥 AI-powered web monitoring platform. Create automated scouts that search the web and send email alerts when they find what you're looking for.

    1,360+3Star change over the last 7 days
  • openserp@karust

    Self-hosted SERP API for AI, SEO & automation. Browser-rendered Google, Bing, Yandex, Baidu, DuckDuckGo and Ecosia search with page extraction 🎉

    1,335+16Star change over the last 7 days
  • WebAI2API@foxhui

    WebAI2API: 基于 Camoufox 的网页 AI 转 API 工具,支持 LMArena/Gemini等,多窗口并发与账号隔离。 | Web AI to OpenAI API via Camoufox. Supports LMArena/Gemini and more, multi-window concurrency & account isolation.

    1,315+16Star change over the last 7 days
  • httpcloak@sardanioss

    Go HTTP client with browser-identical TLS/HTTP2 fingerprinting. Bypass bot detection by perfectly mimicking Chrome, Firefox, and Safari at the cryptographic level (JA3/JA4, Akamai fingerprint, header order). Supports HTTP/1.1, HTTP/2, HTTP/3, sessions, cookies, and proxies.

    1,275+15Star change over the last 7 days
  • Decodo@Decodo

    HTTP(S)/SOCKS5 rotating residential proxies - code examples & general information.

    1,244+5Star change over the last 7 days
  • user-agents@intoli

    A JavaScript library for generating random user agents with data that's updated daily.

    1,188-1Star change over the last 7 days
  • miasma@austin-weeks

    Trap AI web scrapers in an endless poison pit.

    1,179+5Star change over the last 7 days
  • faster-than-requests@juancarlospaco

    Faster requests on Python 3

    1,128-1Star change over the last 7 days
  • reverse-api-engineer@nottelabs

    The agent that turns websites into APIs!

    1,121+17Star change over the last 7 days
  • kimuraframework@vifreefly

    Write web scrapers in Ruby using a clean, AI-assisted DSL. Kimurai uses AI to figure out where the data lives, then caches the selectors and scrapes with pure Ruby. Get the intelligence of an LLM without the per-request latency or token costs.

    1,102+0Star change over the last 7 days
  • scrapfly-scrapers@scrapfly

    Scalable Python web scraping scripts for +40 popular domains

    1,077+6Star change over the last 7 days
  • awesome-agent-apis@Anil-matcha

    660+ muapi-hosted generative-media models plus community-submitted third-party API tools (SEO, enrichment, social, scraping) — one YAML file per entry, browsable by capability.

    1,005+8Star change over the last 7 days
  • x-reader@runesleo

    Universal content reader MCP Server for 10+ platforms

    961+4Star change over the last 7 days
  • crw@us

    Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.

    959+94Star change over the last 7 days
  • Collect structured responses from a ChatGPT scraper by sending a prompt with valid ChatGPT scraping API credentials. Enable live search, inject HTML context, and automate intelligent scraper ChatGPT workflows in just a few parameters.

    916+181Star change over the last 7 days
  • ShardBrowser@ProxyShard

    Free, open-source anti-detect browser launcher for web scraping and multi-accounting. By the ProxyShard team. Engine-level fingerprint spoofing in Chromium 148 (WebGL / WebGPU / Client Hints / fonts / TLS), 170+ device profiles bundled, stable QUIC + WebRTC over SOCKS5.

    909+176Star change over the last 7 days
  • Official Agent skills of Oxylabs products

    870+104Star change over the last 7 days
← Back to topics