Skip to main content
buildradar
Sign in
Topic · scraper

scraper

Tracked open-source repos tagged scraper, sorted by stars.

114 repos
  • newspaper4k@AndyTheFactory

    📰 Newspaper4k a fork of the beloved Newspaper3k. Extraction of articles, titles, and metadata from news websites.

    1,141+1Star change over the last 7 days
  • crawly@elixir-crawly

    Crawly, a high-level web crawling & scraping framework for Elixir.

    1,116+2Star change over the last 7 days
  • kimuraframework@vifreefly

    Write web scrapers in Ruby using a clean, AI-assisted DSL. Kimurai uses AI to figure out where the data lives, then caches the selectors and scrapes with pure Ruby. Get the intelligence of an LLM without the per-request latency or token costs.

    1,102+0Star change over the last 7 days
  • scrapfly-scrapers@scrapfly

    Scalable Python web scraping scripts for +40 popular domains

    1,076+5Star change over the last 7 days
  • jikan@jikan-me

    Unofficial MyAnimeList PHP+REST API which provides functions other than the official API

    1,072+6Star change over the last 7 days
  • URS@JosephLai241

    Universal Reddit Scraper - A comprehensive Reddit scraping/archival command-line tool.

    1,022+1Star change over the last 7 days
  • GreenResourcesManager@klsdf

    一个管理电脑内资源,并提供游戏化的综合管理器

    1,020+9Star change over the last 7 days
  • wreq@0x676e67

    An ergonomic, privacy-aware Rust HTTP Client

    1,018+13Star change over the last 7 days
  • google-play-scraper@JoMingyu

    Google play scraper for Python inspired by <facundoolano/google-play-scraper>

    1,010+3Star change over the last 7 days
  • First ever tool to view "Instagram private posts" anonymously

    995+14Star change over the last 7 days
  • crawler@fredwu

    A high performance web crawler / scraper in Elixir.

    956+0Star change over the last 7 days
  • x-tweet-fetcher@ythx-101

    Fetch X/Twitter tweets, replies, timelines, and articles without login or API keys — field tool for AI agents.

    955+7Star change over the last 7 days
  • scrapyrt@scrapinghub

    HTTP API for Scrapy spiders

    884+2Star change over the last 7 days
  • skrape.it@skrapeit

    A Kotlin-based testing/scraping/parsing library providing the ability to analyze and extract data from HTML (server & client-side rendered). It places particular emphasis on ease of use and a high level of readability by providing an intuitive DSL. It aims to be a testing lib, but can also be used to scrape websites in a convenient fashion.

    875+0Star change over the last 7 days
  • xidel@benibela

    Command line tool to download and extract data from HTML/XML pages or JSON-APIs, using CSS, XPath 3.0, XQuery 3.0, JSONiq or pattern matching. It can also create new or transformed XML/HTML/JSON documents.

    844+1Star change over the last 7 days
  • zimit@openzim

    Make a ZIM file from any Web site and surf offline!

    842+5Star change over the last 7 days
  • HDoujinDownloader@HDoujinDownloader

    A general-purpose doujinshi and image gallery downloader

    841+0Star change over the last 7 days
  • epublifier@maoserr

    Converts some webnovels to epub format

    840+2Star change over the last 7 days
  • spidr@postmodern

    A versatile Ruby web spidering library that can spider a site, multiple domains, certain links or infinitely. Spidr is designed to be fast and easy to use.

    836+0Star change over the last 7 days
  • jvppeteer@fanyong920

    Java API For Chrome and Firefox

    805+0Star change over the last 7 days
  • opensanctions@opensanctions

    An open database of international sanctions data, persons of interest and politically exposed persons

    797+6Star change over the last 7 days
  • trawl@germondai

    Self-hosted scraping engine — bypasses any JS challenge & captcha: Cloudflare, Turnstile, reCAPTCHA, hCaptcha, GeeTest. FlareSolverr & Byparr alternative and drop-in replacement for your *arr stack.

    786+46Star change over the last 7 days
  • linkedin-profile-scraper-api@josephlimtech

    🕵️‍♂️ LinkedIn profile scraper returning structured profile data in JSON.

    776+0Star change over the last 7 days
  • operative framework is a rust investigation OSINT framework, you can interact with multiple targets, execute multiple modules, create links with target, export rapport to PDF file, add note to target or results, interact with RESTFul API, write your own modules.

    749-1Star change over the last 7 days
  • Discount-Bandit@Cybrarist

    A Multi-User Selfhosted price tracker for Amazon, Aliexpress, ebay and many more along with custom stores. Get notified when price matches one or more criterias set by the user

    740+3Star change over the last 7 days
  • HomeHarvest@ZacharyHampton

    Python package for scraping real estate property data

    739+3Star change over the last 7 days
  • slash@theahmadov

    The Slash OSINT Tool

    738+1Star change over the last 7 days
  • scala-scraper@ruippeixotog

    A Scala library for scraping content from HTML pages

    732+0Star change over the last 7 days
  • scrapers@cassidoo

    A list of scrapers from around the web.

    730+0Star change over the last 7 days
  • Twitter Internal API Document

    711+2Star change over the last 7 days
← Back to topics