haul

Image crawler

A tool to extract images from web pages and URLs

An Extensible Image Crawler

GitHub

158 stars
11 watching
38 forks
Language: Python
last commit: over 9 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
archiveteam/grab-siteA web crawler designed to backup websites by recursively crawling and writing WARC files.1,406
sananth12/imagescraperDownloads images from a webpage in parallel using multiple threads and saves them to a specified directory763
webrecorder/browsertrix-crawlerA containerized browser-based crawler system for capturing web content in a high-fidelity and customizable manner.677
vida-nyu/acheA web crawler designed to efficiently collect and prioritize relevant content from the web459
vicktornl/wagtail-stock-imagesA tool to search and add stock images to the Wagtail content management system.10
fredwu/crawlerA high-performance web crawling and scraping solution with customizable settings and worker pooling.945
puerkitobio/fetchbotA flexible web crawler that follows robots.txt policies and crawl delays.787
jmg/crawleyA Pythonic framework for building high-speed web crawlers with flexible data extraction and storage options.188
evyatarmeged/stegextractA tool to extract hidden data from images by detecting embedded files and strings.116
mapbox/robosatAn end-to-end pipeline for extracting features from aerial and satellite imagery using convolutional neural networks2,027
internetarchive/brozzlerA distributed web crawler that fetches and extracts links from websites using a real browser.678
archiveteam/wpullDownloads and crawls web pages, allowing for the archiving of websites.556
azubieta/appimages.scraperA tool to extract AppImage release data from web pages11
elliotgao2/gainA Python web crawling framework utilizing asyncio and aiohttp for efficient data extraction from websites.2,037
stewartmckee/cobwebA flexible web crawler that can be used to extract data from websites in a scalable and efficient manner226