scraper
Tracked open-source repos tagged scraper, sorted by stars.
- #61
📰 Newspaper4k a fork of the beloved Newspaper3k. Extraction of articles, titles, and metadata from news websites.
★ 1,141+1Star change over the last 7 days - #62★ 1,116+2Star change over the last 7 days
- #63
Write web scrapers in Ruby using a clean, AI-assisted DSL. Kimurai uses AI to figure out where the data lives, then caches the selectors and scrapes with pure Ruby. Get the intelligence of an LLM without the per-request latency or token costs.
★ 1,102+0Star change over the last 7 days - #64
Scalable Python web scraping scripts for +40 popular domains
★ 1,076+5Star change over the last 7 days - #65
Unofficial MyAnimeList PHP+REST API which provides functions other than the official API
★ 1,072+6Star change over the last 7 days - #66
Universal Reddit Scraper - A comprehensive Reddit scraping/archival command-line tool.
★ 1,022+1Star change over the last 7 days - #67
一个管理电脑内资源,并提供游戏化的综合管理器
★ 1,020+9Star change over the last 7 days - #68★ 1,018+13Star change over the last 7 days
- #69
Google play scraper for Python inspired by <facundoolano/google-play-scraper>
★ 1,010+3Star change over the last 7 days - #70
First ever tool to view "Instagram private posts" anonymously
★ 995+14Star change over the last 7 days - #71★ 956+0Star change over the last 7 days
- #72
Fetch X/Twitter tweets, replies, timelines, and articles without login or API keys — field tool for AI agents.
★ 955+7Star change over the last 7 days - #73★ 884+2Star change over the last 7 days
- #74
A Kotlin-based testing/scraping/parsing library providing the ability to analyze and extract data from HTML (server & client-side rendered). It places particular emphasis on ease of use and a high level of readability by providing an intuitive DSL. It aims to be a testing lib, but can also be used to scrape websites in a convenient fashion.
★ 875+0Star change over the last 7 days - #75
Command line tool to download and extract data from HTML/XML pages or JSON-APIs, using CSS, XPath 3.0, XQuery 3.0, JSONiq or pattern matching. It can also create new or transformed XML/HTML/JSON documents.
★ 844+1Star change over the last 7 days - #76★ 842+5Star change over the last 7 days
- #77
A general-purpose doujinshi and image gallery downloader
★ 841+0Star change over the last 7 days - #78
Converts some webnovels to epub format
★ 840+2Star change over the last 7 days - #79
A versatile Ruby web spidering library that can spider a site, multiple domains, certain links or infinitely. Spidr is designed to be fast and easy to use.
★ 836+0Star change over the last 7 days - #80★ 805+0Star change over the last 7 days
- #81
An open database of international sanctions data, persons of interest and politically exposed persons
★ 797+6Star change over the last 7 days - #82
Self-hosted scraping engine — bypasses any JS challenge & captcha: Cloudflare, Turnstile, reCAPTCHA, hCaptcha, GeeTest. FlareSolverr & Byparr alternative and drop-in replacement for your *arr stack.
★ 786+46Star change over the last 7 days - #83
🕵️♂️ LinkedIn profile scraper returning structured profile data in JSON.
★ 776+0Star change over the last 7 days - #84
operative framework is a rust investigation OSINT framework, you can interact with multiple targets, execute multiple modules, create links with target, export rapport to PDF file, add note to target or results, interact with RESTFul API, write your own modules.
★ 749-1Star change over the last 7 days - #85
A Multi-User Selfhosted price tracker for Amazon, Aliexpress, ebay and many more along with custom stores. Get notified when price matches one or more criterias set by the user
★ 740+3Star change over the last 7 days - #86
Python package for scraping real estate property data
★ 739+3Star change over the last 7 days - #87★ 738+1Star change over the last 7 days
- #88
A Scala library for scraping content from HTML pages
★ 732+0Star change over the last 7 days - #89★ 730+0Star change over the last 7 days
- #90
Twitter Internal API Document
★ 711+2Star change over the last 7 days