web-crawling
Tracked open-source repos tagged web-crawling, sorted by stars.
Related topics
Topics that frequently appear alongside web-crawling on the same repo.
Recent risers
Repos created in the last 90 days, tagged web-crawling.
- #1
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
★ 25,556+98Star change over the last 7 days - #2
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
★ 9,477+19Star change over the last 7 days - #3
Declarative data automation language and Go runtime for structured extraction workflows.
★ 6,008-1Star change over the last 7 days - #4
The All in One Framework to Build Undefeatable Scrapers
★ 5,692+12Star change over the last 7 days - #5
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
★ 2,614+10Star change over the last 7 days - #6★ 1,512+556Star change over the last 7 days
- #7★ 663+1Star change over the last 7 days