D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Explore 3 GitHub repositories focused on crawling. Discover top-starred projects and those trending this week.
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
Take a list of domains, crawl urls and scan for endpoints, secrets, api keys, file extensions, tokens and more
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
Take a list of domains, crawl urls and scan for endpoints, secrets, api keys, file extensions, tokens and more