awesome-dataops

DataOps toolkit

A curated list of tools and technologies for DataOps, covering data cataloging, exploration, ingestion, processing, and more.

sunglasses A curated list of awesome DataOps tools

GitHub

163 stars
9 watching
20 forks
Language: Python
last commit: almost 2 years ago
Linked from 1 awesome list

awesomeawesome-listdata-engineerdata-engineeringdataops

Awesome DataOps / Data Catalog

AmundsenData discovery and metadata engine for improving the productivity when interacting with data
Apache AtlasProvides open metadata management and governance capabilities to build a data catalog
CKAN4,509over 1 year agoOpen-source DMS (data management system) for powering data hubs and data portals
DataHub10,046over 1 year agoLinkedIn's generalized metadata search & discovery tool
Magda518over 1 year agoA federated, open-source data catalog for all your big data and small data
Marquez1,800almost 2 years agoService for the collection, aggregation, and visualization of a data ecosystem's metadata
Metacat1,616over 1 year agoUnified metadata exploration API service for Hive, RDS, Teradata, Redshift, S3 and Cassandra
OpenLineage1,802over 1 year agoOpen standard for metadata and lineage collection
OpenMetadataA Single place to discover, collaborate and get your data right
Unity CatalogIndustry’s only universal catalog for data and AI

Awesome DataOps / Data Exploration

Apache ZeppelinEnables data-driven, interactive data analytics and collaborative documents
Jupyter NotebookWeb-based notebook environment for interactive computing
JupyterLabThe next-generation user interface for Project Jupyter
Jupytext6,673almost 2 years agoJupyter Notebooks as Markdown Documents, Julia, Python or R scripts
PolynoteThe polyglot notebook with first-class Scala support

Awesome DataOps / Data Ingestion

Amazon KinesisEasily collect, process, and analyze video and data streams in real time
Apache Gobblin2,232almost 2 years agoA framework that simplifies common aspects of big data such as data ingestion
Apache Kafka29,060almost 2 years agoOpen-source distributed event streaming platform used by thousands of companies
Apache Pulsar14,315almost 2 years agoDistributed pub-sub messaging platform with a flexible messaging model and intuitive API
Embulk1,758almost 2 years agoA parallel bulk data loader that helps data transfer between various storages
Fluentd12,963over 1 year agoCollects events from various data sources and writes them to files
Google PubSubIngest events for streaming into BigQuery, data lakes or operational databases
Nakadi958over 2 years agoA distributed event bus that implements a RESTful API abstraction on top of Kafka-like queues
Pravega1,983about 2 years agoAn open source distributed storage service implementing Streams
RabbitMQOne of the most popular open source message brokers

Awesome DataOps / Data Workflow

Apache Airflow37,580almost 2 years agoA platform to programmatically author, schedule, and monitor workflows
Apache Oozie717about 2 years agoAn extensible, scalable and reliable system to manage complex Hadoop workloads
Azkaban4,481about 2 years agoBatch workflow job scheduler created at LinkedIn to run Hadoop jobs
Dagster12,055almost 2 years agoAn orchestration platform for the development, production, and observation of data assets
Luigi17,950almost 2 years agoPython module that helps you build complex pipelines of batch jobs
PrefectA workflow management system, designed for modern infrastructure

Awesome DataOps / Data Processing

Apache Beam7,911almost 2 years agoA unified model for defining both batch and streaming data-parallel processing pipelines
Apache Flink24,261almost 2 years agoAn open source stream processing framework with powerful capabilities
Apache Hadoop MapReduceA framework for writing applications which process vast amounts of data
Apache Nifi4,955almost 2 years agoAn easy to use, powerful, and reliable system to process and distribute data
Apache Samza817almost 2 years agoA distributed stream processing framework which uses Apache Kafka and Hadoop YARN
Apache Spark40,170almost 2 years agoA unified analytics engine for large-scale data processing
Apache Storm6,603almost 2 years agoAn open source distributed realtime computation system
Apache Tez482almost 2 years agoA generic data-processing pipeline engine envisioned as a low-level engine
Faust6,751about 2 years agoA stream processing library, porting the ideas from Kafka Streams to Python

Awesome DataOps / Data Quality

Cerberus3,179about 2 years agoLightweight, extensible data validation library for Python
Cleanlab9,820almost 2 years agoData-centric AI tool to detect (non-predefined) issues in ML data like label errors or outliers
DataProfiler1,442almost 2 years agoA Python library designed to make data analysis, monitoring, and sensitive data detection easy
Deequ3,324almost 2 years agoA library built on top of Apache Spark for measuring data quality in large datasets
Great ExpectationsA Python data validation framework that allows to test your data against datasets
JSON SchemaA vocabulary that allows you to annotate and validate JSON documents
SodaSQL61almost 4 years agoData profiling, testing, and monitoring for SQL accessible data

Awesome DataOps / Data Serialization

Apache Avro2,973almost 2 years agoA data serialization system which is compact, fast and provides rich data structures
Apache ORC698almost 2 years agoA self-describing type-aware columnar file format designed for Hadoop workloads
Apache Parquet2,665almost 2 years agoA columnar storage format which provides efficient storage and encoding of data
Kryo6,217almost 2 years agoA fast and efficient binary object graph serialization framework for Java
ProtoBuf65,999almost 2 years agoLanguage-neutral, platform-neutral, extensible mechanism for serializing structured data

Awesome DataOps / Data Serialization / Data Compression

Pigz2,669almost 2 years agoA parallel implementation of gzip for modern multi-processor, multi-core machines
Snappy6,217about 2 years agoOpen source compression library that is fast, stable and robuts

Awesome DataOps / Data Serialization / Data Table Format

Apache Hudi5,498almost 2 years agoManages the storage of large analytical datasets on DFS
Apache Iceberg6,621almost 2 years agoOpen table format for huge analytic datasets
Delta Lake7,677almost 2 years agoAn open source project that enables building a Lakehouse architecture on top of data lakes

Awesome DataOps / Data Visualization

Apache Superset63,320almost 2 years agoA modern data exploration and data visualization platform
CountSQL/drag-and-drop querying and visualisation tool based on notebooks
Dash21,641almost 2 years agoAnalytical Web Apps for Python, R, Julia, and Jupyter
Data StudioReporting solution for power users who want to go beyond the data and dashboards of GA
HUE1,188over 1 year agoA mature SQL Assistant for querying Databases & Data Warehouses
Lux5,226over 2 years agoFast and easy data exploration by automating the visualization and data analysis process
MetabaseThe simplest, fastest way to get business intelligence and analytics to everyone
RedashConnect to any data source, easily visualize, dashboard and share your data
TableauPowerful and fastest growing data visualization tool used in the business intelligence industry

Awesome DataOps / Data Warehouse

Amazon RedshiftAccelerate your time to insights with fast, easy, and secure cloud data warehousing
Apache Hive5,577almost 2 years agoFacilitates reading, writing, and managing large datasets residing in distributed storage
Apache Kylin3,661almost 2 years agoAn open source, distributed analytical data warehouse for big data
Google BigQueryServerless, highly scalable, and cost-effective multicloud data warehouse

Awesome DataOps / Database / Columnar Database

Apache Cassandra8,906almost 2 years agoOpen source column based DBMS designed to handle large amounts of data
Apache Druid13,548almost 2 years agoDesigned to quickly ingest massive quantities of event data, and provide low-latency queries
Apache HBase5,246almost 2 years agoAn open-source, distributed, versioned, column-oriented store
Scylla13,725almost 2 years agoDesigned to be compatible with Cassandra while achieving higher throughputs and lower latencies

Awesome DataOps / Database / Document-Oriented Database

Apache CouchDB6,298almost 2 years agoAn open-source document-oriented NoSQL database, implemented in Erlang
Elasticsearch71,007almost 2 years agoA distributed document oriented database with a RESTful search engine
MongoDB26,503almost 2 years agoA cross-platform document database that uses JSON-like documents with optional schemas
RethinkDB26,806almost 2 years agoThe first open-source scalable database built for realtime applications

Awesome DataOps / Database / Graph Database

Age3,191almost 2 years agoA multi-model database that supports both graph and relational data models
ArangoDB13,613almost 2 years agoA scalable open-source multi-model database natively supporting graph, document and search
JanusGraph5,351almost 2 years agoManage large graphs with billions of data distributed across a multi-machine cluster
Memgraph2,520over 1 year agoAn open source graph database, built for real-time streaming data, compatible with Neo4j
Neo4j13,537almost 2 years agoA high performance graph store with all the features expected of a mature and robust database
Titan5,243almost 4 years agoA highly scalable graph database optimized for storing and querying large graphs

Awesome DataOps / Database / Key-Value Database

Apache Accumulo1,075almost 2 years agoA sorted, distributed key-value store that provides robust and scalable data storage
Dragonfly26,326over 1 year agoA modern in-memory datastore, fully compatible with Redis and Memcached APIs
DynamoDBFast, flexible NoSQL database service for single-digit millisecond performance at any scale
etcd48,056over 1 year agoDistributed reliable key-value store for the most critical data of a distributed system
EVCache2,071almost 2 years agoA distributed in-memory data store for the cloud
Memcached13,601almost 2 years agoA high performance multithreaded event-based key/value cache store
Redis67,358almost 2 years agoAn in-memory key-value database that persists on disk

Awesome DataOps / Database / Relational Database

CockroachDB30,270almost 2 years agoA distributed database designed to build, scale, and manage data-intensive apps
Crate4,139almost 2 years agoA distributed SQL database that makes it simple to store and analyze massive amounts of data
MariaDB5,752over 1 year agoA replacement of MySQL with more features, new storage engines and better performance
MySQL10,964almost 2 years agoOne of the most popular open source transactional databases
PostgreSQL16,442almost 2 years agoAn advanced RDBMS that supports an extended subset of the SQL standard
RQLite15,906almost 2 years agoA lightweight, distributed relational database, which uses SQLite as its storage engine
SQLite6,902over 1 year agoA popular choice as embedded database software for local/client storage

Awesome DataOps / Database / Time Series Database

Akumuli835about 4 years agoCan be used to capture, store and process time-series data in real-time
Atlas3,459almost 2 years agoAn in-memory dimensional time series database
InfluxDB29,126almost 2 years agoScalable datastore for metrics, events, and real-time analytics
QuestDB14,699almost 2 years agoAn open source SQL database designed to process time series data, faster
TimescaleDB18,066over 1 year agoOpen-source time-series SQL database optimized for fast ingest and complex queries

Awesome DataOps / Database / Vector Database

Milvus31,283almost 2 years agoAn open source embedding vector similarity search engine powered by Faiss, NMSLIB and Annoy
PineconeManaged and distributed vector similarity search used with a lightweight SDK
Qdrant21,001over 1 year agoAn open source vector similarity search engine with extended filtering support

Awesome DataOps / File System

Alluxio6,880almost 2 years agoA virtual distributed storage system
Amazon Simple Storage Service (S3)Object storage built to retrieve any amount of data from anywhere
Apache Hadoop Distributed File System (HDFS)A distributed file system
GlusterFS4,774almost 2 years agoA software defined distributed storage that can scale to several petabytes
Google Cloud Storage (GCS)Object storage for companies of all sizes, to store any amount of data
LakeFS4,496over 1 year agoOpen source tool that transforms your object storage into a Git-like repository
LizardFS958about 2 years agoA highly reliable, scalable and efficient distributed file system
MinIO48,833almost 2 years agoHigh Performance, Kubernetes Native Object Storage compatible with Amazon S3 API
SeaweedFS23,207almost 2 years agoA fast distributed storage system for blobs, objects, files, and data lake
Swift2,639over 1 year agoA distributed object storage system designed to scale from a single machine to thousands of servers

Awesome DataOps / Logging and Monitoring

Grafana65,525over 1 year agoVisualize metrics, logs, and traces from multiple sources like Prometheus, Loki, InfluxDB and more
Loki24,172over 1 year agoA horizontally-scalable, highly-available, multi-tenant log aggregation system inspired by Prometheus
Prometheus56,244almost 2 years agoA monitoring system and time series database
Whylogs2,664almost 2 years agoA tool for creating data logs, enabling monitoring for data drift and data quality issues

Awesome DataOps / Metadata Service

Hive MetastoreService that stores metadata related to Apache Hive and other services
Metacat1,616over 1 year agoProvides you information about what data you have, where it resides and how to process it

Awesome DataOps / SQL Query Engine

Apache Drill1,949almost 2 years agoSchema-free SQL Query Engine for Hadoop, NoSQL and Cloud Storage
Apache Impala1,164almost 2 years agoLightning-fast, distributed SQL queries for petabytes of data
DremioPower high-performing BI dashboards and interactive analytics directly on data lake
Presto16,114over 1 year agoA distributed SQL query engine for big data
Trino10,601over 1 year agoA fast distributed SQL query engine for big data analytics

Resources / Books

Data Mesh: Delivering Data-Driven Value at Scale(O'Reilly)
Designing Data-Intensive Applications(O'Reilly)
Fundamentals of Data Engineering(O'Reilly)
Getting Started with Impala(O'Reilly)
Learning and Operating Presto(O'Reilly)
Learning Spark: Lightning-Fast Data Analytics(O'Reilly)
Spark in Action(O'Reilly)
Spark: The Definitive Guide(O'Reilly)

Resources / Other Lists

Awesome Data Engineering6,889almost 2 years ago
Awesome MLOps4,181almost 2 years ago
DataOps Resource24about 6 years ago

Resources / Slack

Delta Lake Workspace
Trino Workspace

Backlinks from these awesome lists:

More related projects: