datafu

Hadoop data processing library

A collection of libraries for working with large-scale data in Hadoop, providing incremental processing capabilities and user-defined functions.

Hadoop library for large-scale data processing, now an Apache Incubator project

GitHub

583 stars
75 watching
133 forks
Language: Java
last commit: about 12 years ago

Related projects:

RepositoryDescriptionStars
apache/datafuA collection of libraries for data mining and statistics in large-scale Hadoop environments119
datasalt/pangoolA Java framework that simplifies Hadoop's MapReduce API to build efficient data processing pipelines57
linkedinattic/cleoA flexible library for enabling rapid development of typeahead search functionality565
linkedinattic/kamikazeA utility package wrapping set implementations on document lists with compression and set operation support.22
lacuna/bifurcanA Java library providing efficient, functional data structures with customizable equality semantics and high performance.968
apache/tezA system that enables flexible data processing pipelines using a low-level engine for higher-level frameworks482
linkeddata/rdflib.jsA JavaScript library for working with RDF data in various formats and querying RDF stores567
dfianthdl/dfhdlA programming language and library for describing dataflow-based digital hardware in a high-level, object-oriented way82
frappe/datatableA modern javascript library for creating interactive and editable tables on the web1,050
twitter/scaldingA Scala library for specifying and executing MapReduce jobs in Hadoop3,506
rbrahul/gofpA utility library providing common functions for working with data structures like slices and maps in Go.148
linkedin/ambryA distributed object store designed to efficiently store and serve large media objects in web applications.1,749
mhausenblas/mrlinMaps RDF data into HBase for scalable storage and processing of Linked Data17
alangrafu/lodspeakrA framework for building Linked Data applications using PHP32
intentmedia/marioA library that enables the definition of complex data pipelines in a functional, typesafe, and efficient way using a declarative syntax139