scraping
Tracked open-source repos tagged scraping, sorted by stars.
Related topics
Topics that frequently appear alongside scraping on the same repo.
Recent risers
Repos created in the last 90 days, tagged scraping.
- #1
Self-hosted scraping engine — bypasses any JS challenge & captcha: Cloudflare, Turnstile, reCAPTCHA, hCaptcha, GeeTest. FlareSolverr & Byparr alternative and drop-in replacement for your *arr stack.
★ 742 - #2
A skill for AI agents: search the web with SearXNG, browse with Camofox, bypass protections with CloakBrowser. Anti-hallucination by design. Self-hosted, free, unlimited.
★ 507
- #1★ 174,034+3,406Star change over the last 7 days
- #2
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
★ 77,158+1,432Star change over the last 7 days - #3★ 64,098+106Star change over the last 7 days
- #4★ 37,143+196Star change over the last 7 days
- #5
Open source AI job application bot in Python: browser automation and web scraping to read job postings, then auto-apply with a tailored resume and cover letter for each posting.
★ 30,278+58Star change over the last 7 days - #6
Python scraper based on AI
★ 30,047+244Star change over the last 7 days - #7
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
★ 25,556+98Star change over the last 7 days - #8★ 25,488+26Star change over the last 7 days
- #9
Pythonic HTML Parsing for Humans™
★ 13,814-4Star change over the last 7 days - #10
Custom Selenium Chromedriver | Zero-Config | Passes ALL bot mitigation systems (like Distil / Imperva/ Datadadome / CloudFlare IUAM)
★ 12,814+10Star change over the last 7 days - #11★ 11,680+0Star change over the last 7 days
- #12★ 11,534+210Star change over the last 7 days
- #13
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
★ 9,477+19Star change over the last 7 days - #14
Stealth headless browser for AI agents — bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement.
★ 8,960+195Star change over the last 7 days - #15
List of libraries, tools and APIs for web scraping and data processing.
★ 8,138+10Star change over the last 7 days - #16
A Smart, Automatic, Fast and Lightweight Web Scraper for Python
★ 7,915+25Star change over the last 7 days - #17★ 7,476+5Star change over the last 7 days
- #18
Pydoll is a library for automating chromium-based browsers without a WebDriver, offering realistic interactions.
★ 7,053+16Star change over the last 7 days - #19
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
★ 6,728+50Star change over the last 7 days - #20
🚀 Free HTTP, SOCKS4, & SOCKS5 proxy list * Updated every 5 minutes * and rotating proxy API (100+ countries)
★ 6,658+57Star change over the last 7 days - #21★ 4,771+1Star change over the last 7 days
- #22
Twitter API Scraper | Without an API key | Twitter Internal API | Free | Twitter scraper | Twitter Bot
★ 4,638+11Star change over the last 7 days - #23★ 4,439+0Star change over the last 7 days
- #24★ 4,011+1Star change over the last 7 days
- #25
AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Native multi-threading for bulk processing.
★ 3,415+2Star change over the last 7 days - #26
Do you want to LEARN NEW STUFF for FREE? Don't worry, with the power of web-scraping and automation, this script will find the necessary Udemy coupons & enroll you for PAID UDEMY COURSES, ABSOLUTELY FREE!
★ 3,291+3Star change over the last 7 days - #27
Scrape Facebook public pages without an API key
★ 3,271+4Star change over the last 7 days - #28★ 3,067+2Star change over the last 7 days
- #29
An open source fingerprint browser based on Ungoogled Chromium. 指纹浏览器 隐私浏览器
★ 3,003+28Star change over the last 7 days - #30
Learn step-by-step how to scrape Google Trends data and make a result comparison using Python and Oxylabs SERP API. Extract keywords, their popularity, breakdown by region, related queries, and more.
★ 2,874+38Star change over the last 7 days