galer

URL extractor

A tool to extract URLs from HTML attributes by crawling in and evaluating JavaScript

A fast tool to fetch URLs from HTML attributes by crawl-in.

GitHub

255 stars
6 watching
38 forks
Language: Go
last commit: almost 2 years ago
crawlerdevtoolextractorgalergogolangspiderurl-extractorurl-parserwaybackurls

Related projects:

RepositoryDescriptionStars
s0rg/crawleyA utility for systematically extracting URLs from web pages and printing them to the console.268
mvdan/xurlsA tool to extract URLs from text using regular expressions in the Go programming language.1,193
003random/getjsA tool to extract JavaScript sources from URLs and web pages efficiently732
eloopwoo/chrome-url-dumperA tool to extract and dump URLs from Chrome's stored databases.34
karust/gogetcrawlA tool and package for extracting web archive data from popular sources like Wayback Machine and Common Crawl using the Go programming language.148
coleifer/micawberA library for extracting metadata and content from URLs635
davemolk/gogetjsTools for extracting and analyzing JavaScript files from web pages41
foolin/pagserA tool for automatically extracting structured data from HTML pages105
iamstoxe/urlgrabA tool to crawl websites by exploring links recursively with support for JavaScript rendering.331
offensivedev/urldozerA tool for analyzing URLs to extract various information such as paths, domains, and parameters.29
archiveteam/grab-siteA web crawler designed to backup websites by recursively crawling and writing WARC files.1,406
gamallo/galextraA multi-language term extractor that uses morphosyntax tagging and filtering to identify multi-word terms from plain text input.2
limiu82214/gojmaprA library to extract specific properties from complex JSON structures into Go structs with minimal code changes.22
plainas/tqTool that extracts content from HTML documents based on CSS selectors236
patternhelloworld/url-knifeA JavaScript library to extract and decompose URLs in texts with robust patterns, including email addresses.341