spider
Tracked open-source repos tagged spider, sorted by stars.
- #31★ 2,463+0Star change over the last 7 days
- #32
🕷 CrawlerDetect is a PHP class for detecting bots/crawlers/spiders via the user agent
★ 2,403+2Star change over the last 7 days - #33
Download images from Google, Bing, Baidu. 谷歌、百度、必应图片下载.
★ 2,360-2Star change over the last 7 days - #34
Cross Platform C# web crawler framework built for speed and flexibility. Please star this project! +1.
★ 2,310+0Star change over the last 7 days - #35
蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统
★ 2,088-1Star change over the last 7 days - #36★ 2,019+0Star change over the last 7 days
- #37★ 1,891+3Star change over the last 7 days
- #38★ 1,879+2Star change over the last 7 days
- #39
A Python based web application scanner to gather OSINT and fuzz for OWASP vulnerabilities on a target website.
★ 1,820-1Star change over the last 7 days - #40
QQ空间导出助手,用于备份QQ空间的说说、日志、私密日记、相册、视频、留言板、QQ好友、收藏夹、分享、最近访客为文件,便于迁移与保存
★ 1,815+13Star change over the last 7 days - #41
python爬虫,目前库存:网易云音乐歌曲爬取,B站视频爬取,知乎问答爬取,壁纸爬取,xvideos视频爬取,有声书爬取,微博爬虫,安居客信息爬取+数据可视化,哔哩哔哩视频封面提取器,ip代理池封装,知乎百万级用户爬虫+数据分析,github用户爬虫
★ 1,649+4Star change over the last 7 days - #42
The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns
★ 1,609+2Star change over the last 7 days - #43★ 1,605+1Star change over the last 7 days
- #44★ 1,410+0Star change over the last 7 days
- #45★ 1,379+0Star change over the last 7 days
- #46
📰 Diários oficiais brasileiros acessíveis a todos | 📰 Brazilian government gazettes, accessible to everyone.
★ 1,375+1Star change over the last 7 days - #47★ 1,303+1Star change over the last 7 days
- #48
A download tools for clawing the ebooks from internets.
★ 1,294+0Star change over the last 7 days - #49
Boss直聘爬虫 / BOSS直聘职位数据抓取工具,基于 Chrome CDP 协议复用真实登录态,绕过字体反爬,输出明文薪资 JSON/CSV + 薪资技能分析。A Chrome-CDP-based BOSS Zhipin job scraper/crawler.
★ 1,282+33Star change over the last 7 days - #50★ 1,258-1Star change over the last 7 days
- #51★ 1,117+3Star change over the last 7 days
- #52★ 1,116+2Star change over the last 7 days
- #53★ 1,104+2Star change over the last 7 days
- #54
爬取微信公众号文章
★ 1,079+4Star change over the last 7 days - #55
Scalable Python web scraping scripts for +40 popular domains
★ 1,076+4Star change over the last 7 days - #56
JS破解逆向,破解JS反爬虫加密参数,已破解极验滑块w(2022.2.19),QQ音乐sign(2022.2.13),拼多多anti_content,boss直聘zp_token,知乎x-zse-96,酷狗kg_mid/dfid,唯品会mars_cid,中国裁判文书网(2020-06-30更新),淘宝密码,天安保险登录,b站登录,房天下登录,WPS登录,微博登录,有道翻译,网易登录,微信公众号登录,空中网登录,今目标登录,学生信息管理系统登录,共赢金融登录,重庆科技资源共享平台登录,网易云音乐下载,一键解析视频链接,财联社登录。
★ 1,035+0Star change over the last 7 days - #57
Golang短视频去水印:抖音,皮皮虾,火山,微视,最右,快手,全民小视频,皮皮搞笑,西瓜视频,虎牙,梨视频,acfun,好看视频...
★ 1,032+10Star change over the last 7 days - #58
小说下载|小说爬取|起点|笔趣阁|导出Markdown|导出txt|转换epub|广告过滤|自动校对
★ 1,017+8Star change over the last 7 days - #59★ 998+0Star change over the last 7 days
- #60
SpiderSuite (web security crawler) releases, wiki and roadmap
★ 974+1Star change over the last 7 days