crawl4ai

Web crawler

A web crawling tool designed to extract structured data from the web for use in AI applications

🚀🤖 Crawl4AI: Crawl Smarter, Faster, Freely. For AI.

GitHub

19k stars
109 watching
1k forks
Language: HTML
last commit: almost 2 years ago

Related projects:

RepositoryDescriptionStars
code4craft/webmagicA framework for building scalable web crawlers in Java.11,456
apify/crawleeA tool for building reliable web scraping and browser automation pipelines in Node.js.16,081
spatie/crawlerA powerful web crawler written in PHP that can execute JavaScript and crawl multiple URLs concurrently.2,552
gocolly/collyA framework for extracting structured data from websites in a fast and elegant way23,444
yasserg/crawler4jA Java-based web crawler for extracting and processing web page content4,563
yujiosaka/headless-chrome-crawlerA distributed crawling framework that leverages Headless Chrome to scrape dynamic websites5,534
s0md3v/photonA fast and flexible web crawler designed to gather information from the internet11,122
bda-research/node-crawlerA NodeJS-based web crawler and spider that extracts data from websites.6,718
s0rg/crawleyA utility for systematically extracting URLs from web pages and printing them to the console.268
ionicabizau/scrape-itA Node.js library and CLI tool for automating web page scraping and parsing4,024
howie6879/ruiaAn async web scraping micro-framework built with asyncio and aiohttp to simplify URL crawling1,753
internetarchive/heritrix3A web crawler designed to collect and preserve digital artifacts while respecting site policies and load constraints.2,857
elliotgao2/gainA Python web crawling framework utilizing asyncio and aiohttp for efficient data extraction from websites.2,037
stewartmckee/cobwebA flexible web crawler that can be used to extract data from websites in a scalable and efficient manner226
archiveteam/grab-siteA web crawler designed to backup websites by recursively crawling and writing WARC files.1,406