🔥 10 GitHub Repositories to Scrape Almost Any Website
1. Firecrawl
Turns entire websites into clean, AI-ready Markdown or structured data with just a few API calls. Perfect for feeding LLMs.
2. Crawl4AI
An open source python crawler built specifically for AI. Extracts clean, structured content optimized for LLMs.
3. Browser Use
AI Agent that control browsers like a human. It allows an AI agent to dynamically visually navigate, click elements, bypass popups, and extract data.
4. Crawlee
A powerful scraping framework for building fast, reliable crawlers with support for Playwright, Puppeteer, and Cheerio.
5. Scrapy
One of the most popular Python frameworks for large-scale web scraping and crawling projects.
6. MarkItDown
Converts PDFs, Office documents, HTML, and many other file types into clean Markdown for AI workflows.
7. Scrapling
A modern Python scraping library that combines speed, browser automation, and smart parsing with a simple API.
8. Skyvern
An AI-powered scraping tool that dynamically solve CAPTCHAs, log into complex portals, and extract data without requiring any pre-defined HTML selectors or XPaths.
9. AutoScraper
Automatically learns how to extract similar data from web pages by showing it just a few examples.
10. curl-impersonate
Makes cURL mimic real browsers like Chrome and Safari to bypass bot detection and access protected websites more reliably.
💡 Save this list for your next web scraping or AI automation project.
Post #897
4.47K
- ❤ 8
- 🔥 2