crawler
Tracked open-source repos tagged crawler, sorted by stars.
- #31
实战🐍多种网站、电商数据爬虫🕷。包含🕸:淘宝商品、微信公众号、大众点评、企查查、招聘网站、闲鱼、阿里任务、博客园、微博、百度贴吧、豆瓣电影、包图网、全景网、豆瓣音乐、某省药监局、搜狐新闻、机器学习文本采集、fofa资产采集、汽车之家、国家统计局、百度关键词收录数、蜘蛛泛目录、今日头条、豆瓣影评、携程、小米应用商店、安居客、途家民宿❤️❤️❤️。微信爬虫展示项目:
★ 5,670+4Star change over the last 7 days - #32
Redis-based components for Scrapy.
★ 5,643+1Star change over the last 7 days - #33
Analysis of Bot Protection systems with available countermeasures 🚿. How to defeat anti-bot system 👻 and get around browser fingerprinting scripts 🕵️♂️ when scraping the web?
★ 5,131+2Star change over the last 7 days - #34★ 4,767+29Star change over the last 7 days
- #35
Collection of China illegal cases about web crawler 本项目用来整理所有中国大陆爬虫开发者涉诉与违规相关的新闻、资料与法律法规。致力于帮助在中国大陆工作的爬虫行业从业者了解我国相关法律,避免触碰数据合规红线。
★ 4,724+6Star change over the last 7 days - #36
新浪微博爬虫,用python爬取新浪微博数据,并下载微博图片和微博视频
★ 4,632+3Star change over the last 7 days - #37
A community-driven way to read and chat with AI bots - powered by chatGPT.
★ 4,417-1Star change over the last 7 days - #38
Download comics novels 小说漫画下载工具 小説漫画のダウンローダ 小說漫畫下載:腾讯漫画 大角虫漫画 有妖气 咪咕 SF漫画 哦漫画 看漫画 漫画柜 汗汗酷漫 動漫伊甸園 快看漫画 微博动漫 733动漫网 大古漫画网 漫画DB 無限動漫 動漫狂 卡推漫画 动漫之家 动漫屋 古风漫画网 36漫画网 亲亲漫画网 乙女漫画 webtoons 咚漫 ニコニコ静画 ComicWalker ヤングエースUP モアイ pixivコミック サイコミ;アルファポリス カクヨム ハーメルン 小説家になろう 起点中文网 八一中文网 顶点小说 落霞小说网 努努书坊 笔趣阁→epub.
★ 4,203+5Star change over the last 7 days - #39
Proxy [Finder | Checker | Server]. HTTP(S) & SOCKS :performing_arts:
★ 4,156-2Star change over the last 7 days - #40
DotnetSpider, a .NET standard web crawling library. It is lightweight, efficient and fast high-level web crawling & scraping framework
★ 4,138+0Star change over the last 7 days - #41
Intelligent proxy pool for Humans™ to extract content from the internet and build your own Large Language Models in this new AI era
★ 4,020+1Star change over the last 7 days - #42
Headless Chrome .NET API
★ 3,917+1Star change over the last 7 days - #43
Take a list of domains, crawl urls and scan for endpoints, secrets, api keys, file extensions, tokens and more
★ 3,758+2Star change over the last 7 days - #44
All in one tool for Information Gathering, Vulnerability Scanning and Crawling. A must have tool for all penetration testers
★ 3,749-1Star change over the last 7 days - #45
🚀🚀🚀feapder is an easy to use, powerful crawler framework | feapder是一款上手简单,功能强大的Python爬虫框架。内置AirSpider、Spider、TaskSpider、BatchSpider四种爬虫解决不同场景的需求。且支持断点续爬、监控报警、浏览器渲染、海量数据去重等功能。更有功能强大的爬虫管理系统feaplat为其提供方便的部署及调度
★ 3,732-3Star change over the last 7 days - #46★ 3,724-1Star change over the last 7 days
- #47★ 3,555+0Star change over the last 7 days
- #48★ 3,036+1Star change over the last 7 days
- #49★ 2,995+1Star change over the last 7 days
- #50
All In One Web Recon
★ 2,957+3Star change over the last 7 days - #51
Node.js scraper to get data from Google Play
★ 2,956+4Star change over the last 7 days - #52★ 2,875+1Star change over the last 7 days
- #53
DecryptLogin: APIs for loginning some websites by using requests.
★ 2,852+0Star change over the last 7 days - #54★ 2,829+0Star change over the last 7 days
- #55
Geziyor, blazing fast web crawling & scraping framework for Go. Supports JS rendering.
★ 2,776+0Star change over the last 7 days - #56
Videodl: A lightweight video downloader written in pure python. (轻量级视频下载器,优先高清无水印,支持抖音,快手,小红书,B站,TikTok,YouTube,FIFA+,优酷,腾讯,爱奇艺,1905电影网,乐视,芒果,咪咕,PPTV,搜狐,Facebook,Twitter,新浪微博,今日头条,网易公开课,全民K歌,CCTV央视频,酷狗音乐MV,新片场,知乎,百度贴吧,TED等海量流媒体平台)
★ 2,713+20Star change over the last 7 days - #57★ 2,690-1Star change over the last 7 days
- #58★ 2,687+6Star change over the last 7 days
- #59★ 2,510+0Star change over the last 7 days
- #60
news-please - an integrated web crawler and information extractor for news that just works
★ 2,485+0Star change over the last 7 days