kedro

Data pipeline toolkit

A toolbox for production-ready data science pipelines with software engineering best practices for reproducibility and modularity

Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, and modular.

GitHub

10k stars
110 watching
909 forks
Language: Python
last commit: almost 2 years ago
Linked from 5 awesome lists

experiment-trackinghacktoberfestkedromachine-learningmachine-learning-engineeringmlopspipelinepython

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
harisekhon/devops-python-toolsTools for managing and automating DevOps tasks, data processing, and cloud infrastructure using Python.783
gradio-app/gradioEnables rapid creation and deployment of web applications for machine learning models and functions using Python34,557
jakevdp/pythondatasciencehandbookAn online guide and set of executable Jupyter notebooks providing an introduction to core libraries for data science in Python.43,422
donnemartin/data-science-ipython-notebooksA comprehensive collection of data science and machine learning notebooks using Python and various deep learning frameworks.27,601
openmined/pysyftEnables data scientists to perform analysis on private data without accessing the underlying data, using a secure and decentralized server architecture.9,557
pypi/warehouseThe software behind the Python Package Index.3,617
sdv-dev/sdvA library for generating synthetic tabular data based on real-world patterns2,416
pachyderm/pachydermAutomates data transformations with versioning and lineage tracking for scalable data pipelines6,191
mito-ds/mitoA Jupyter Notebook add-on for spreadsheet-like editing and automation of Pandas dataframes2,318
unstructured-io/unstructuredA toolkit for building custom machine learning pipelines from unstructured data9,452
jupyterlab/jupyterlabAn extensible environment for interactive and reproducible computing using the Jupyter Notebook architecture14,263
ahkarami/deep-learning-in-productionA collection of notes and references on deploying deep learning models in production environments4,313
ploomber/ploomberA platform for building and deploying data pipelines using Python, with features for caching, automation, and modularization.3,530
pandas-dev/pandasA powerful data analysis toolkit for Python that provides flexible and expressive data structures for efficient data manipulation and analysis.44,052
rdkit/rdkitA comprehensive software suite for cheminformatics and machine learning tasks2,712