crawler
Tracked open-source repos tagged crawler, sorted by stars.
- #91★ 1,410+0Star change over the last 7 days
- #92★ 1,386+1Star change over the last 7 days
- #93★ 1,379+0Star change over the last 7 days
- #94
Lightweight Ruby web crawler/scraper with an elegant DSL which extracts structured data from pages.
★ 1,360+0Star change over the last 7 days - #95★ 1,303+1Star change over the last 7 days
- #96
📝 quickly crawl the information (e.g. followers, tags etc...) of an instagram profile.
★ 1,302+1Star change over the last 7 days - #97★ 1,302+0Star change over the last 7 days
- #98
Boss直聘爬虫 / BOSS直聘职位数据抓取工具,基于 Chrome CDP 协议复用真实登录态,绕过字体反爬,输出明文薪资 JSON/CSV + 薪资技能分析。A Chrome-CDP-based BOSS Zhipin job scraper/crawler.
★ 1,282+33Star change over the last 7 days - #99★ 1,258-1Star change over the last 7 days
- #100
一个好用的哔哩哔哩漫画下载器,拥有图形界面,支持关键词搜索漫画和二维码登入,黑科技下载未解锁章节,多线程下载,多种保存格式,本地漫画管理,一键检查更新!
★ 1,243+1Star change over the last 7 days - #101
基于appium的app自动遍历工具
★ 1,237+0Star change over the last 7 days - #102
Easily download all the photos/videos from tumblr blogs. 下载指定的 Tumblr 博客中的图片,视频
★ 1,157+0Star change over the last 7 days - #103
📰 Newspaper4k a fork of the beloved Newspaper3k. Extraction of articles, titles, and metadata from news websites.
★ 1,141+1Star change over the last 7 days - #104
Open-source, self-hosted enterprise & site search server built on OpenSearch. Crawls web / file / DB / cloud sources, 20+ languages, REST API, and AI/RAG & semantic search. Apache-2.0.
★ 1,128+1Star change over the last 7 days - #105
Run a high-fidelity browser-based web archiving crawler in a single Docker container
★ 1,127+5Star change over the last 7 days - #106★ 1,116+2Star change over the last 7 days
- #107
Write web scrapers in Ruby using a clean, AI-assisted DSL. Kimurai uses AI to figure out where the data lives, then caches the selectors and scrapes with pure Ruby. Get the intelligence of an LLM without the per-request latency or token costs.
★ 1,102+0Star change over the last 7 days - #108
Scalable Python web scraping scripts for +40 popular domains
★ 1,075+4Star change over the last 7 days - #109
Golang短视频去水印:抖音,皮皮虾,火山,微视,最右,快手,全民小视频,皮皮搞笑,西瓜视频,虎牙,梨视频,acfun,好看视频...
★ 1,031+10Star change over the last 7 days - #110★ 1,016+12Star change over the last 7 days
- #111
小说下载|小说爬取|起点|笔趣阁|导出Markdown|导出txt|转换epub|广告过滤|自动校对
★ 1,016+8Star change over the last 7 days - #112★ 1,009-1Star change over the last 7 days
- #113
Google play scraper for Python inspired by <facundoolano/google-play-scraper>
★ 1,009+2Star change over the last 7 days - #114
A scalable, mature and versatile web crawler based on Apache Storm
★ 994+1Star change over the last 7 days - #115
SpiderSuite (web security crawler) releases, wiki and roadmap
★ 974+1Star change over the last 7 days - #116★ 956+0Star change over the last 7 days
- #117
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
★ 952+87Star change over the last 7 days - #118
Pure-Python protocol tools and research notes for X.com, with VibeLoft Twitter account datasets
★ 940-1Star change over the last 7 days - #119★ 933+1Star change over the last 7 days
- #120
ChatWeb can crawl web pages, read PDF, DOCX, TXT, and extract the main content, then answer your questions based on the content, or summarize the key points.
★ 916+0Star change over the last 7 days