dud

Data pipeline manager

A lightweight tool for managing and versioning large data alongside source code in data pipelines

A lightweight CLI tool for versioning data alongside source code and building data pipelines.

GitHub

184 stars
8 watching
8 forks
Language: Go
last commit: almost 2 years ago
Linked from 1 awesome list

data-engineeringdata-pipelinesdata-sciencedatasetdvcsmachine-learningmlops

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
prodmodel/prodmodelA tool for managing data science pipelines by automating build, testing, and deployment processes while ensuring correctness and performance.58
linkedin/brooklinA distributed system for streaming data between heterogeneous systems with high reliability and throughput at scale931
johnsonc/lambdoA workflow engine for unifying feature engineering and machine learning operations in data analysis pipelines1
hyfather/pipelineA package implementing pipelines using goroutines to manage concurrency in Go applications.56
fluidattacks/makesA framework for building and managing CI/CD pipelines and application environments with cryptographic signed dependencies.461
galaxyproject/galaxyA platform for data-intensive scientific analysis and workflow management1,431
ypares/porcupineA tool that enables data manipulation and analysis pipelines to be flexible, reusable, and reproducible in different environments89
samapriya/planet-gee-pipeline-cliA command-line tool for automating data processing and uploads from Planet's API to Google Earth Engine.42
pdpipe/pdpipeProvides a set of pre-defined data processing pipelines for pandas DataFrames.718
apache/streampipesA toolbox for industrial data analytics and stream processing614
druths/xpA tool for creating flexible and self-documenting data science pipelines56
montilab/pipelinerA framework for defining and automating bioinformatics pipelines using Nextflow.44
calebwin/pipelinesA language and runtime for crafting massively parallel data pipelines375
ssadedin/bpipeA tool for running and managing bioinformatics pipelines by abstracting away low-level details and providing features such as dependency tracking, transactional management, and parallelism.233
giacbrd/smartpipelineA framework for designing and executing concurrent data pipelines with a focus on simplicity and efficiency25