butterfree
Feature pipeline builder
A Python library for building data pipelines to create and load features into a feature store using Apache Spark.
A tool for building feature stores.
288 stars
195 watching
36 forks
Language: Python
last commit: almost 2 years agoLinked from 2 awesome lists
data-engineeringdata-scienceetletl-frameworkfeature-storepackagepysparkpython
Related projects:
| Repository | Description | Stars |
|---|---|---|
| A tool that enables data analysts to create and manage data pipelines with an intuitive interface, generating Python code for deployment anywhere. | 933 | |
| An asset packaging library for Django that simplifies CSS and JavaScript concatenation and compression. | 1,520 | |
| Simplifies the deployment of Kubeflow Pipelines workflows by providing a graphical interface for Data Scientists to define and deploy pipelines directly from JupyterLab. | 632 | |
| A tool for creating flexible and self-documenting data science pipelines | 56 | |
| A tool that speeds up the development of Django Rest APIs by automating repetitive tasks. | 118 | |
| A framework for building pluggable business logic pipelines with a focus on modular and composable components. | 362 | |
| A workflow engine for unifying feature engineering and machine learning operations in data analysis pipelines | 1 | |
| A Python package to build and experiment with machine learning pipelines using Kedro, MLflow, and other tools | 226 | |
| A library for composing and chaining functions on Observables in RxJava to simplify complex data processing pipelines. | 49 | |
| A Python project implementing a novel approach to high-performance feature learning and dimensionality reduction in deep neural networks | 7 | |
| A tool that enables data manipulation and analysis pipelines to be flexible, reusable, and reproducible in different environments | 89 | |
| A Python-based feature store library with a simple, scalable, and flexible architecture for storing and managing data for machine learning applications. | 58 | |
| A Java framework that simplifies Hadoop's MapReduce API to build efficient data processing pipelines | 57 | |
| A Python framework for real-time data processing on Apache Kafka streams | 1,246 | |
| A framework for designing and executing concurrent data pipelines with a focus on simplicity and efficiency | 25 |