datatrove

Data pipeline framework

A platform-agnostic data processing framework for large-scale text data pipelines

Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.

GitHub

2k stars
47 watching
155 forks
Language: Python
last commit: almost 2 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
giacbrd/smartpipelineA framework for designing and executing concurrent data pipelines with a focus on simplicity and efficiency25
vectaport/flowgraphA software framework for building scalable, asynchronous data pipelines with explicit back-pressure management and logging capabilities.60
pdpipe/pdpipeProvides a set of pre-defined data processing pipelines for pandas DataFrames.718
ypares/porcupineA tool that enables data manipulation and analysis pipelines to be flexible, reusable, and reproducible in different environments89
databiosphere/toilA workflow management system designed to efficiently run pipelines in various environments.901
log2timeline/dftimewolfA framework for orchestrating data collection, processing, and export299
dataform-co/dataformA framework for managing data operations in BigQuery using SQL and software engineering best practices860
galaxyproject/galaxyA platform for data-intensive scientific analysis and workflow management1,431
mara/mara-pipelinesA lightweight ETL framework providing a simple way to define and execute data transformation pipelines using declarative Python code.2,082
olirice/flupyA library that provides a fluent interface for processing data pipelines in Python without holding large amounts of memory193
johnsonc/lambdoA workflow engine for unifying feature engineering and machine learning operations in data analysis pipelines1
valeriobasile/learningbyreadingA software framework for building NLP and entity linking pipelines with semantic parsing, word sense disambiguation, and entity linking capabilities.82
intentmedia/marioA library that enables the definition of complex data pipelines in a functional, typesafe, and efficient way using a declarative syntax139
druths/xpA tool for creating flexible and self-documenting data science pipelines56
datasalt/pangoolA Java framework that simplifies Hadoop's MapReduce API to build efficient data processing pipelines57