crawler
Tracked open-source repos tagged crawler, sorted by stars.
- #61
Leaked GPTs Prompts Bypass the 25 message limit or to try out GPTs without a Plus subscription.
★ 2,469+1Star change over the last 7 days - #62★ 2,463+0Star change over the last 7 days
- #63
🕷 CrawlerDetect is a PHP class for detecting bots/crawlers/spiders via the user agent
★ 2,403+2Star change over the last 7 days - #64
Cross Platform C# web crawler framework built for speed and flexibility. Please star this project! +1.
★ 2,310+0Star change over the last 7 days - #65
Website Cloner - Utilizes powerful Go routines to clone websites to your computer within seconds.
★ 2,239+5Star change over the last 7 days - #66
🏳️🌈 Media downloader from any sites, including Twitter, Reddit, Instagram, BlueSky, TikTok, Threads, Facebook, OnlyFans, YouTube, Pinterest, PornHub, XHamster, XVIDEOS, ThisVid etc.
★ 2,161+8Star change over the last 7 days - #67
Transform Web Content into LLM-Ready Data
★ 2,158+32Star change over the last 7 days - #68
蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统
★ 2,088-1Star change over the last 7 days - #69★ 2,019+0Star change over the last 7 days
- #70
爬虫案例合集。包括但不限于《淘宝、京东、天猫、豆瓣、抖音、快手、微博、微信、阿里、头条、pdd、优酷、爱奇艺、携程、12306、58、搜狐、各种指数、维普万方、Zlibraty、Oalib、小说、招标网、采购网、小红书、大众点评、推特、脉脉、知乎》
★ 1,967+1Star change over the last 7 days - #71
Free antidetect browser stealth for Playwright: undetected headless Firefox fingerprint. Python scraping, recaptcha and bot detection bypass. Open source
★ 1,962+13Star change over the last 7 days - #72
NewPipe's core library for extracting data from streaming sites
★ 1,960+3Star change over the last 7 days - #73
浏览器内存漫游解决方案(探索中...)
★ 1,912+0Star change over the last 7 days - #74★ 1,891+3Star change over the last 7 days
- #75★ 1,879+2Star change over the last 7 days
- #76
多平台内容监控·采集·搬运 一个 Web 面板管起抖音 / 小红书 / 快手
★ 1,833+65Star change over the last 7 days - #77
Diskover Community Edition - Open source file indexer, file search engine and data management and analytics powered by Elasticsearch
★ 1,816-2Star change over the last 7 days - #78
Google, Naver multiprocess image web crawler (Selenium)
★ 1,692+1Star change over the last 7 days - #79
Python binding to Modest and Lexbor engines. Fast HTML5 parser with CSS selectors for Python.
★ 1,671+4Star change over the last 7 days - #80
python爬虫,目前库存:网易云音乐歌曲爬取,B站视频爬取,知乎问答爬取,壁纸爬取,xvideos视频爬取,有声书爬取,微博爬虫,安居客信息爬取+数据可视化,哔哩哔哩视频封面提取器,ip代理池封装,知乎百万级用户爬虫+数据分析,github用户爬虫
★ 1,649+4Star change over the last 7 days - #81
The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns
★ 1,609+2Star change over the last 7 days - #82★ 1,605+1Star change over the last 7 days
- #83
ScopeSentry-Cyberspace mapping, subdomain enumeration, port scanning, sensitive information discovery, vulnerability scanning, distributed nodes
★ 1,601+9Star change over the last 7 days - #84
CLI tool for saving a faithful copy of a complete web page in a single HTML file (based on SingleFile)
★ 1,579+9Star change over the last 7 days - #85
❤️ Fredy - [F]ind [R]eal [E]state [D]amn Eas[y] - Fredy keeps searching for new apartments, houses, and flats in Europe on platforms like ImmoScout24, Immowelt, eBay Kleinanzeigen and instantly delivers the results to you via Slack, Telegram, Email, Discord or ntfy, so you can focus on the more important things in life ;)
★ 1,462+11Star change over the last 7 days - #86★ 1,451+0Star change over the last 7 days
- #87
An ergonomic, privacy-aware Python HTTP Client
★ 1,438+1Star change over the last 7 days - #88
小红书数据采集、网站图片、视频资源批量下载工具,颜值超高的数据采集工具(批量下载,视频提取,图片)Telegram:https://t.me/+ZtLSwuIKTo44MDY1
★ 1,436+2Star change over the last 7 days - #89
Collection of patches for puppeteer and playwright to avoid automation detection and leaks. Helps to avoid Cloudflare and DataDome CAPTCHA pages. Easy to patch/unpatch, can be enabled/disabled on demand.
★ 1,422-1Star change over the last 7 days - #90★ 1,417-1Star change over the last 7 days