awesome-data-engineering

Data engineering toolkit

A curated list of tools and technologies for data engineering

A curated list of data engineering tools for software developers

GitHub

7k stars
261 watching
1k forks
last commit: almost 2 years ago
Linked from 5 awesome lists

awesomeawesome-list

Awesome Data Engineering / Databases / Relational

RQLite15,906almost 2 years agoReplicated SQLite using the Raft consensus protocol
MySQLThe world's most popular open source database

Awesome Data Engineering / Databases / Relational / MySQL

TiDB37,447almost 2 years agoTiDB is a distributed NewSQL database compatible with MySQL protocol
Percona XtraBackupPercona XtraBackup is a free, open source, complete online backup solution for all versions of Percona Server, MySQL® and MariaDB®
mysql_utils883about 7 years agoPinterest MySQL Management Tools

Awesome Data Engineering / Databases / Relational

MariaDBAn enhanced, drop-in replacement for MySQL
PostgreSQLThe world's most advanced open source database
Amazon RDSAmazon RDS makes it easy to set up, operate, and scale a relational database in the cloud
Crate.IOScalable SQL database with the NOSQL goodies

Awesome Data Engineering / Databases / Key-Value

RedisAn open source, BSD licensed, advanced key-value cache and store
RiakA distributed database designed to deliver maximum data availability by distributing data across multiple servers
AWS DynamoDBA fast and flexible NoSQL database service for all applications that need consistent, single-digit millisecond latency at any scale
HyperDex1,393over 2 years agoHyperDex is a scalable, searchable key-value store. Deprecated
SSDBA high performance NoSQL database supporting many data structures, an alternative to Redis
Kyoto Tycoon277almost 3 years agoKyoto Tycoon is a lightweight network server on top of the Kyoto Cabinet key-value database, built for high-performance and concurrency
IonDB592over 2 years agoA key-value store for microcontroller and IoT applications

Awesome Data Engineering / Databases / Column

CassandraThe right choice when you need scalability and high availability without compromising performance

Awesome Data Engineering / Databases / Column / Cassandra

Cassandra CalculatorThis simple form allows you to try out different values for your Apache Cassandra cluster and see what the impact is for your application
CCM1,219almost 2 years agoA script to easily create and destroy an Apache Cassandra cluster on localhost
ScyllaDB13,725almost 2 years agoNoSQL data store using the seastar framework, compatible with Apache Cassandra

Awesome Data Engineering / Databases / Column

HBaseThe Hadoop database, a distributed, scalable, big data store
AWS RedshiftA fast, fully managed, petabyte-scale data warehouse that makes it simple and cost-effective to analyze all your data using your existing business intelligence tools
FiloDB1,429almost 2 years agoDistributed. Columnar. Versioned. Streaming. SQL
VerticaDistributed, MPP columnar database with extensive analytics SQL
ClickHouseDistributed columnar DBMS for OLAP. SQL

Awesome Data Engineering / Databases / Document

MongoDBAn open-source, document database designed for ease of development and scaling

Awesome Data Engineering / Databases / Document / MongoDB

Percona Server for MongoDBPercona Server for MongoDB® is a free, enhanced, fully compatible, open source, drop-in replacement for the MongoDB® Community Edition that includes enterprise-grade features and functionality
MemDB596over 8 years agoDistributed Transactional In-Memory Database (based on MongoDB)

Awesome Data Engineering / Databases / Document

ElasticsearchSearch & Analyze Data in Real Time
CouchbaseThe highest performing NoSQL distributed database
RethinkDBThe open-source database for the realtime web
RavenDBFully Transactional NoSQL Document Database

Awesome Data Engineering / Databases / Graph

Neo4jThe world's leading graph database
OrientDB2nd Generation Distributed Graph Database with the flexibility of Documents in one product with an Open Source commercial friendly license
ArangoDBA distributed free and open-source database with a flexible data model for documents, graphs, and key-values
TitanA scalable graph database optimized for storing and querying graphs containing hundreds of billions of vertices and edges distributed across a multi-machine cluster
FlockDB3,337over 9 years agoA distributed, fault-tolerant graph database by Twitter. Deprecated

Awesome Data Engineering / Databases / Distributed

DAtomicThe fully transactional, cloud-ready, distributed database
Apache GeodeAn open source, distributed, in-memory database for scale-out applications
Gaffer1,774almost 2 years agoA large-scale graph database

Awesome Data Engineering / Databases / Timeseries

InfluxDB29,126almost 2 years agoScalable datastore for metrics, events, and real-time analytics
OpenTSDB5,009almost 2 years agoA scalable, distributed Time Series Database
QuestDBA relational column-oriented database designed for real-time analytics on time series and event data
kairosdb1,740almost 2 years agoFast scalable time series database
Heroic848over 5 years agoA scalable time series database based on Cassandra and Elasticsearch, by Spotify
Druid13,548almost 2 years agoColumn oriented distributed data store ideal for powering interactive applications
Riak-TSRiak TS is the only enterprise-grade NoSQL time series database optimized specifically for IoT and Time Series data
Akumuli835about 4 years agoAkumuli is a numeric time-series database. It can be used to capture, store and process time-series data in real-time. The word "akumuli" can be translated from esperanto as "accumulate"
RhombusA time-series object store for Cassandra that handles all the complexity of building wide row indexes
Dalmatiner DB694over 7 years agoFast distributed metrics database
Blueflood595about 2 years agoA distributed system designed to ingest and process time series data
Timely379about 2 years agoTimely is a time series database application that provides secure access to time series data based on Accumulo and Grafana

Awesome Data Engineering / Databases / Other

Tarantool3,437almost 2 years agoTarantool is an in-memory database and application server
GreenPlumThe Greenplum Database (GPDB) - An advanced, fully featured, open source data warehouse. It provides powerful and rapid analytics on petabyte scale data volumes
cayley14,868almost 2 years agoAn open-source graph database. Google
Snappydata1,041almost 4 years agoSnappyData: OLTP + OLAP Database built on Apache Spark
TimescaleDBBuilt as an extension on top of PostgreSQL, TimescaleDB is a time-series SQL database providing fast analytics, scalability, with automated data management on a proven storage engine
DuckDBDuckDB is a fast in-process analytical database that has zero external dependencies, runs on Linux/macOS/Windows, offers a rich SQL dialect, and is free and extensible

Awesome Data Engineering / Data Comparison

datacompy487almost 2 years agoDataComPy is a Python library that facilitates the comparison of two DataFrames in pandas, Polars, Spark and more. The library goes beyond basic equality checks by providing detailed insights into discrepancies at both row and column levels

Awesome Data Engineering / Data Ingestion

KafkaPublish-subscribe messaging rethought as a distributed commit log

Awesome Data Engineering / Data Ingestion / Kafka

BottledWater4over 3 years agoChange data capture from PostgreSQL into Kafka. Deprecated
kafkat504over 7 years agoSimplified command-line administration for Kafka brokers
kafkacat5,468about 2 years agoGeneric command line non-JVM Apache Kafka producer and consumer
pg-kafka111over 11 years agoA PostgreSQL extension to produce messages to Apache Kafka
librdkafka332almost 2 years agoThe Apache Kafka C/C++ library
kafka-docker6,943over 2 years agoKafka in Docker
kafka-manager11,853about 3 years agoA tool for managing Apache Kafka
kafka-node2,664about 3 years agoNode.js client for Apache Kafka 0.8
Secor1,846almost 2 years agoPinterest's Kafka to S3 distributed consumer
Kafka-logger45almost 8 years agoKafka-winston logger for Node.js from Uber

Awesome Data Engineering / Data Ingestion

AWS KinesisA fully managed, cloud-based service for real-time data processing over large, distributed data streams
RabbitMQRobust messaging for applications
dltA fast&simple pipeline building library for python data devs, runs in notebooks, cloud functions, airflow, etc
FluentDAn open source data collector for unified logging layer
EmbulkAn open source bulk data loader that helps data transfer between various databases, storages, file formats, and cloud services
Apache SqoopA tool designed for efficiently transferring bulk data between Apache Hadoop and structured datastores such as relational databases
Heka3,389over 2 years agoData Acquisition and Processing Made Easy. Deprecated
Gobblin2,232almost 2 years agoUniversal data ingestion framework for Hadoop from LinkedIn
NakadiNakadi is an open source event messaging platform that provides a REST API on top of Kafka-like queues
PravegaPravega provides a new storage abstraction - a stream - for continuous and unbounded data
Apache PulsarApache Pulsar is an open-source distributed pub-sub messaging system
AWS Data Wrangler3,951almost 2 years agoUtility belt to handle data on AWS
AirbyteOpen-source data integration for modern data teams
ArtieReal-time data ingestion tool leveraging change data capture
SlingSling is CLI data integration tool specialized in moving data between databases, as well as storage systems
MeltanoCLI & code-first ELT

Awesome Data Engineering / Data Ingestion / Meltano

Singer SDKThe fastest way to build custom data extractors and loaders compliant with the Singer Spec

Awesome Data Engineering / Data Ingestion

Google Sheets ETL18almost 2 years agoLive import all your Google Sheets to your data warehouse

Awesome Data Engineering / File System

HDFSA distributed file system designed to run on commodity hardware

Awesome Data Engineering / File System / HDFS

Snakebite854over 4 years agoA pure python HDFS client

Awesome Data Engineering / File System

AWS S3Object storage built to retrieve any amount of data from anywhere

Awesome Data Engineering / File System / AWS S3

smart_open3,233almost 2 years agoUtils for streaming large files (S3, HDFS, gzip, bz2)

Awesome Data Engineering / File System

AlluxioAlluxio is a memory-centric distributed storage system enabling reliable data sharing at memory-speed across cluster frameworks, such as Spark and MapReduce
CEPHCeph is a unified, distributed storage system designed for excellent performance, reliability, and scalability
JuiceFS11,030almost 2 years agoJuiceFS is a high-performance Cloud-Native file system driven by object storage for large-scale data storage
OrangeFSOrange File System is a branch of the Parallel Virtual File System
SnackFS14about 11 years agoSnackFS is our bite-sized, lightweight HDFS compatible file system built over Cassandra
GlusterFSGluster Filesystem
XtreemFSFault-tolerant distributed file system for all storage needs
SeaweedFS23,207almost 2 years agoSeaweed-FS is a simple and highly scalable distributed file system. There are two objectives: to store billions of files! to serve the files fast! Instead of supporting full POSIX file system semantics, Seaweed-FS choose to implement only a key~file mapping. Similar to the word "NoSQL", you can call it as "NoFS"
S3QL1,126almost 2 years agoS3QL is a file system that stores all its data online using storage services like Google Storage, Amazon S3, or OpenStack
LizardFSLizardFS Software Defined Storage is a distributed, parallel, scalable, fault-tolerant, Geo-Redundant and highly available file system

Awesome Data Engineering / Serialization format

Apache AvroApache Avro™ is a data serialization system
Apache ParquetApache Parquet is a columnar storage format available to any project in the Hadoop ecosystem, regardless of the choice of data processing framework, data model or programming language

Awesome Data Engineering / Serialization format / Apache Parquet

Snappy6,217about 2 years agoA fast compressor/decompressor. Used with Parquet
PigZA parallel implementation of gzip for modern multi-processor, multi-core machines

Awesome Data Engineering / Serialization format

Apache ORCThe smallest, fastest columnar storage for Hadoop workloads
Apache ThriftThe Apache Thrift software framework, for scalable cross-language services development
ProtoBuf65,999almost 2 years agoProtocol Buffers - Google's data interchange format
SequenceFileSequenceFile is a flat file consisting of binary key/value pairs. It is extensively used in MapReduce as input/output formats
Kryo6,217almost 2 years agoKryo is a fast and efficient object graph serialization framework for Java

Awesome Data Engineering / Stream Processing

Apache BeamApache Beam is a unified programming model that implements both batch and streaming data processing jobs that run on many execution engines
Spark StreamingSpark Streaming makes it easy to build scalable fault-tolerant streaming applications
Apache FlinkApache Flink is a streaming dataflow engine that provides data distribution, communication, and fault tolerance for distributed computations over data streams
Apache StormApache Storm is a free and open source distributed realtime computation system
Apache SamzaApache Samza is a distributed stream processing framework
Apache NiFiAn easy to use, powerful, and reliable system to process and distribute data
Apache HudiAn open source framework for managing storage for real time processing, one of the most interesting feature is the Upsert
VoltDBVoltDb is an ACID-compliant RDBMS which uses a
PipelineDB2,639over 4 years agoThe Streaming SQL Database
Spring Cloud DataflowStreaming and tasks execution between Spring Boot apps
BonoboBonobo is a data-processing toolkit for python 3.5+
Robinhood's Faust1,675almost 2 years agoForever scalable event processing & in-memory durable K/V store as a library with asyncio & static typing
HStreamDB713almost 2 years agoThe streaming database built for IoT data storage and real-time processing
Kuiper1,505almost 2 years agoAn edge lightweight IoT data analytics/streaming software implemented by Golang, and it can be run at all kinds of resource-constrained edge devices
Zilla553almost 2 years ago- An API gateway built for event-driven architectures and streaming that supports standard protocols such as HTTP, SSE, gRPC, MQTT, and the native Kafka protocol
SwimOS321almost 2 years agoA framework for building real-time streaming data processing applications that supports a wide range of ingestion sources

Awesome Data Engineering / Batch Processing

Hadoop MapReduceHadoop MapReduce is a software framework for easily writing applications which process vast amounts of data (multi-terabyte data-sets) - in-parallel on large clusters (thousands of nodes) - of commodity hardware in a reliable, fault-tolerant manner
SparkA multi-language engine for executing data engineering, data science, and machine learning on single-node machines or clusters

Awesome Data Engineering / Batch Processing / Spark

Spark PackagesA community index of packages for Apache Spark
Deep Spark197about 10 years agoConnecting Apache Spark with different data stores. Deprecated
Spark RDD API ExamplesExamples by Zhen He
LivyThe REST Spark Server
Delight344over 2 years agoA free & cross platform monitoring tool (Spark UI / Spark History Server alternative)

Awesome Data Engineering / Batch Processing

AWS EMRA web service that makes it easy to quickly and cost-effectively process vast amounts of data
Data MechanicsA cloud-based platform deployed on Kubernetes making Apache Spark more developer-friendly and cost-effective
TezAn application framework which allows for a complex directed-acyclic-graph of tasks for processing data
Bistro7about 8 years agoA light-weight engine for general-purpose data processing including both batch and stream analytics. It is based on a novel unique data model, which represents data via and processes data via as opposed to having only set operations in conventional approaches like MapReduce or SQL

Awesome Data Engineering / Batch Processing / Batch ML

H2OFast scalable machine learning API for smarter applications
MahoutAn environment for quickly creating scalable performant machine learning applications
Spark MLlibSpark's scalable machine learning library consisting of common learning algorithms and utilities, including classification, regression, clustering, collaborative filtering, dimensionality reduction, as well as underlying optimization primitives

Awesome Data Engineering / Batch Processing / Batch Graph

GraphLab CreateA machine learning platform that enables data scientists and app developers to easily create intelligent apps at scale
GiraphAn iterative graph processing system built for high scalability
Spark GraphXApache Spark's API for graphs and graph-parallel computation

Awesome Data Engineering / Batch Processing / Batch SQL

PrestoA distributed SQL query engine designed to query large data sets distributed over one or more heterogeneous data sources
HiveData warehouse software facilitates querying and managing large datasets residing in distributed storage

Awesome Data Engineering / Batch Processing / Batch SQL / Hive

Hivemall311about 4 years agoScalable machine learning library for Hive/Hadoop
PyHive1,676about 2 years agoPython interface to Hive and Presto

Awesome Data Engineering / Batch Processing / Batch SQL

DrillSchema-free SQL Query Engine for Hadoop, NoSQL and Cloud Storage

Awesome Data Engineering / Charts and Dashboards

HighchartsA charting library written in pure JavaScript, offering an easy way of adding interactive charts to your web site or web application
ZingChartFast JavaScript charts for any data set
C3.jsD3-based reusable chart library
D3.jsA JavaScript library for manipulating documents based on data

Awesome Data Engineering / Charts and Dashboards / D3.js

D3PlusD3's simpler, easier to use cousin. Mostly predefined templates that you can just plug data in

Awesome Data Engineering / Charts and Dashboards

SmoothieChartsA JavaScript Charting Library for Streaming Data
PyXley2,272over 8 years agoPython helpers for building dashboards using Flask and React
Plotly21,641almost 2 years agoFlask, JS, and CSS boilerplate for interactive, web-based visualization apps in Python
Apache Superset63,320almost 2 years agoApache Superset (incubating) - A modern, enterprise-ready business intelligence web application
RedashMake Your Company Data Driven. Connect to any data source, easily visualize and share your data
Metabase39,103almost 2 years agoMetabase is the easy, open source way for everyone in your company to ask questions and learn from data
PyQtGraphPyQtGraph is a pure-python graphics and GUI library built on PyQt4 / PySide and numpy. It is intended for use in mathematics / scientific / engineering applications

Awesome Data Engineering / Workflow

Luigi17,950almost 2 years agoLuigi is a Python module that helps you build complex pipelines of batch jobs
CronQAn application cron-like system. w/Luige. Deprecated
CascadingJava based application development platform
Airflow37,580almost 2 years agoAirflow is a system to programmatically author, schedule, and monitor data pipelines
AzkabanAzkaban is a batch workflow job scheduler created at LinkedIn to run Hadoop jobs. Azkaban resolves the ordering through job dependencies and provides an easy-to-use web user interface to maintain and track your workflows
OozieOozie is a workflow scheduler system to manage Apache Hadoop jobs
Pinball1,046almost 7 years agoDAG based workflow manager. Job flows are defined programmatically in Python. Support output passing between jobs
Dagster12,055almost 2 years agoDagster is an open-source Python library for building data applications
Hamilton1,900almost 2 years agoHamilton is a lightweight library to define data transformations as a directed-acyclic graph (DAG). If you like dbt for SQL transforms, you will like Hamilton for Python processing
KedroKedro is a framework that makes it easy to build robust and scalable data pipelines by providing uniform project templates, data abstraction, configuration and pipeline assembly
DataformAn open-source framework and web based IDE to manage datasets and their dependencies. SQLX extends your existing SQL warehouse dialect to add features that support dependency management, testing, documentation and more
CensusA reverse-ETL tool that let you sync data from your cloud data warehouse to SaaS applications like Salesforce, Marketo, HubSpot, Zendesk, etc. No engineering favors required—just SQL
dbtA command line tool that enables data analysts and engineers to transform data in their warehouses more effectively
KestraScalable, event-driven, language-agnostic orchestration and scheduling platform to manage millions of workflows declaratively in code
RudderStack4,109almost 2 years agoA warehouse-first Customer Data Platform that enables you to collect data from every application, website and SaaS platform, and then activate it in your warehouse and business tools
PACE34almost 2 years agoAn open source framework that allows you to enforce agreements on how data should be accessed, used, and transformed, regardless of the data platform (Snowflake, BigQuery, DataBricks, etc.)
PrefectPrefect is an orchestration and observability platform. With it, developers can rapidly build and scale resilient code, and triage disruptions effortlessly
Multiwoven1,556almost 2 years agoThe open-source reverse ETL, data activation platform for modern data teams
SuprSendCreate automated workflows and logic using API's for your notification service. Add templates, batching, preferences, inapp inbox with workflows to trigger notifications directly from your data warehouse
Kestra14,708almost 2 years agoA versatile open source orchestrator and scheduler built on Java, designed to handle a broad range of workflows with a language-agnostic, API-first architecture
MageOpen-source data pipeline tool for transforming and integrating data

Awesome Data Engineering / Data Lake Management

lakeFS4,496almost 2 years agolakeFS is an open source platform that delivers resilience and manageability to object-storage based data lakes
Project Nessie1,064almost 2 years agoProject Nessie is a Transactional Catalog for Data Lakes with Git-like semantics. Works with Apache Iceberg tables

Awesome Data Engineering / ELK Elastic Logstash Kibana

docker-logstash236almost 11 years agoA highly configurable Logstash (1.4.4) - Docker image running Elasticsearch (1.7.0) - and Kibana (3.1.2)
elasticsearch-jdbc2,838almost 5 years agoJDBC importer for Elasticsearch
ZomboDB4,687almost 2 years agoPostgres Extension that allows creating an index backed by Elasticsearch

Awesome Data Engineering / Docker

Gockerize666over 8 years agoPackage golang service into minimal Docker containers
Flocker3,390over 9 years agoEasily manage Docker containers & their data
RancherRancherOS is a 20mb Linux distro that runs the entire OS as Docker containers
KontenaApplication Containers for Masses
Weave6,621about 2 years agoWeaving Docker containers into applications
Zodiac198over 6 years agoA lightweight tool for easy deployment and rollback of dockerized applications
cAdvisor17,304almost 2 years agoAnalyzes resource usage and performance characteristics of running containers
Micro S3 persistence14almost 7 years agoDocker microservice for saving/restoring volume data to S3
Rocker-compose406over 3 years agoDocker composition tool with idempotency features for deploying apps composed of multiple containers. Deprecated
Nomad15,029almost 2 years agoNomad is a cluster manager, designed for both long-lived services and short-lived batch processing workloads
ImageLayersVisualize Docker images and the layers that compose them

Awesome Data Engineering / Datasets / Realtime

Twitter RealtimeThe Streaming APIs give developers low latency access to Twitter's global stream of Tweet data
Eventsim508over 4 years agoEvent data simulator. Generates a stream of pseudo-random events from a set of users, designed to simulate web traffic
RedditReal-time data is available including comments, submissions and links posted to reddit

Awesome Data Engineering / Datasets / Data Dumps

GitHub ArchiveGitHub's public timeline since 2011, updated every hour
Common CrawlOpen source repository of web crawl data
WikipediaWikipedia's complete copy of all wikis, in the form of Wikitext source and metadata embedded in XML. A number of raw database tables in SQL form are also available

Awesome Data Engineering / Monitoring / Prometheus

Prometheus.io56,244almost 2 years agoAn open-source service monitoring system and time series database
HAProxy Exporter619over 3 years agoSimple server that scrapes HAProxy stats and exports them via HTTP for Prometheus consumption

Awesome Data Engineering / Profiling / Data Profiler

Data Profiler1,442almost 2 years agoThe DataProfiler is a Python library designed to make data analysis, monitoring, and sensitive data detection easy

Awesome Data Engineering / Testing

Grai301almost 2 years agoA data catalog tool that integrates into your CI system exposing downstream impact testing of data changes. These tests prevent data changes which might break data pipelines or BI dashboards from making it to production
DQOps118almost 2 years agoAn open-source data quality platform for the whole data platform lifecycle from profiling new data sources to applying full automation of data quality monitoring
DataKitchenOpen Source Data Observability for end-to-end Data Journey Observability, data profiling, anomaly detection, and auto-created data quality validation tests

Awesome Data Engineering / Community / Forums

/r/dataengineeringNews, tips, and background on Data Engineering
/r/etlSubreddit focused on ETL

Awesome Data Engineering / Community / Conferences

Data CouncilData Council is the first technical conference that bridges the gap between data scientists, data engineers and data analysts

Awesome Data Engineering / Community / Podcasts

Data Engineering PodcastThe show about modern data infrastructure
The Data Stack ShowA show where they talk to data engineers, analysts, and data scientists about their experience around building and maintaining data infrastructure, delivering data and data products, and driving better outcomes across their businesses with data

Backlinks from these awesome lists:

More related projects: