cascalog

Data processor

A library for data processing and querying on large datasets without the need for Hadoop expertise

Data processing on Hadoop without the hassle.

GitHub

1k stars
80 watching
178 forks
Language: Clojure
last commit: over 3 years ago

Related projects:

RepositoryDescriptionStars
scicloj/tableclothA dataset manipulation library built on top of tech.ml.dataset, providing a simplified API for data processing and analysis.308
dkogan/vnlogA toolkit for manipulating tabular ASCII data with normal UNIX tools.161
ndmitchell/cmdargsA Haskell library for building command line applications with minimal code91
techascent/tech.ml.datasetA Clojure library for efficient tabular data processing and analysis687
nysol/mcmdA set of commands for high-speed processing of large-scale CSV data33
tweag/haskellrAn environment for efficient data processing using Haskell or R code.587
zepgram/module-multi-threadingA module that enables parallel processing of large data sets in Magento 2 using multiple child processes.80
snoyberg/conduitA framework for handling and transforming streaming data in a consistent and efficient way903
kapolos/pramdaA PHP implementation of functional programming concepts to simplify data processing and analysis.245
netflix/pigpenA map-reduce framework for Clojure that compiles to Apache Pig or Cascading without requiring prior knowledge of those systems.567
travitch/datalogA Haskell implementation of Datalog, allowing recursive queries in a logic language.104
apache/samzaA distributed stream processing framework for handling high-volume data streams with fault tolerance and durability guarantees817
hashrock/deno-fnparseA parser combinator library for Deno that provides a simple way to parse CSV data.11
sodiumjoe/lobarA command-line wrapper around lodash's chain method for functional data processing28
reubano/mezaA lightweight toolkit for processing tabular data with a focus on functional programming and PyPy compatibility.417