xidel

Web scraper

A tool to extract data from web pages using various query languages and selectors.

Command line tool to download and extract data from HTML/XML pages or JSON-APIs, using CSS, XPath 3.0, XQuery 3.0, JSONiq or pattern matching. It can also create new or transformed XML/HTML/JSON documents.

GitHub

690 stars
27 watching
42 forks
Language: Pascal
last commit: over 2 years ago
Linked from 1 awesome list

clicommand-linecss-selectorcurldata-processingdatascrapinghtmlhttphttpiejsonrestscraperwebwebscraperwebscrapingwgetxmlxmlstarletxpathxquery

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
felipecsl/wombatA Ruby-based web crawler and data extraction tool with an elegant DSL.1,315
the-markup/blacklight-collectorA tool for scraping website content and analyzing browser behavior205
miyagawa/web-scraperA Perl toolkit for extracting structured data from HTML documents using a DSL-like interface.104
joseconstela/webparsyA Node.js library and CLI for scraping websites using Puppeteer and YAML definitions44
spekulatius/phpscraperA web scraping utility for PHP that simplifies the process of extracting information from websites.544
slotix/dataflowkitA framework for extracting structured data from web pages using CSS selectors.667
medialab/minetA command line tool and Python library for extracting data from various web sources.293
jaimeiniesta/metainspectorA Ruby gem for web scraping and extracting metadata from web pages.1,038
bplawler/crawlerA Scala-based DSL for programmatically accessing and interacting with web pages149
oscarotero/embedA PHP library to retrieve metadata and embed code from any web page2,100
zhuyingda/websterA framework for automating web scraping and crawling tasks using Node.js518
spider-rs/spiderA tool for web data extraction and processing using Rust1,234
gushonorato/mechanizeA web scraping and automation tool for Elixir.30
meilisearch/docs-scraperAutomates scraping and indexing of documentation content into a search engine297
jakopako/goskyrA tool to simplify web scraping of list-like structured data from web pages36