Awesome Lists

awesome-dataops

by kelvins

awesome listPythonpushed almost 2 years ago

sunglasses A curated list of awesome DataOps tools

AI summary

DataOps toolkit

A curated list of tools and technologies for DataOps, covering data cataloging, exploration, ingestion, processing, and more.

stars
163
forks
20
watching
9
awesome list
1
entries
140
View on GitHub

Embed the badge

Show how many awesome lists link to your project. The count updates automatically.

Awesome Lists badge
Markdown
[![Awesome Lists Badge](https://awesome.facts.dev/shield/kelvins/awesome-dataops/links.svg)](https://awesome.facts.dev/awesome/kelvins/awesome-dataops)
HTML
<a href="https://awesome.facts.dev/awesome/kelvins/awesome-dataops"><img src="https://awesome.facts.dev/shield/kelvins/awesome-dataops/links.svg" alt="Awesome Lists Badge" /></a>
Image URL
https://awesome.facts.dev/shield/kelvins/awesome-dataops/links.svg

What's in the list

140 links in 25 sections, with live GitHub stats.activeno commit in 2y

Data Catalog

  • Amundsen

    Data discovery and metadata engine for improving the productivity when interacting with data

  • Apache Atlas

    Provides open metadata management and governance capabilities to build a data catalog

  • CKAN

    Open-source DMS (data management system) for powering data hubs and data portals

  • DataHub

    LinkedIn's generalized metadata search & discovery tool

  • Magda

    A federated, open-source data catalog for all your big data and small data

  • Marquez

    Service for the collection, aggregation, and visualization of a data ecosystem's metadata

  • Metacat

    Unified metadata exploration API service for Hive, RDS, Teradata, Redshift, S3 and Cassandra

  • OpenLineage

    Open standard for metadata and lineage collection

  • OpenMetadata

    A Single place to discover, collaborate and get your data right

  • Unity Catalog

    Industry’s only universal catalog for data and AI

Data Exploration

  • Apache Zeppelin

    Enables data-driven, interactive data analytics and collaborative documents

  • Jupyter Notebook

    Web-based notebook environment for interactive computing

  • JupyterLab

    The next-generation user interface for Project Jupyter

  • Jupytext

    Jupyter Notebooks as Markdown Documents, Julia, Python or R scripts

  • Polynote

    The polyglot notebook with first-class Scala support

Data Ingestion

  • Amazon Kinesis

    Easily collect, process, and analyze video and data streams in real time

  • Apache Gobblin

    A framework that simplifies common aspects of big data such as data ingestion

  • Apache Kafka

    Open-source distributed event streaming platform used by thousands of companies

  • Apache Pulsar

    Distributed pub-sub messaging platform with a flexible messaging model and intuitive API

  • Embulk

    A parallel bulk data loader that helps data transfer between various storages

  • Fluentd

    Collects events from various data sources and writes them to files

  • Google PubSub

    Ingest events for streaming into BigQuery, data lakes or operational databases

  • Nakadi

    A distributed event bus that implements a RESTful API abstraction on top of Kafka-like queues

  • Pravega

    An open source distributed storage service implementing Streams

  • RabbitMQ

    One of the most popular open source message brokers

Data Workflow

  • Apache Airflow

    A platform to programmatically author, schedule, and monitor workflows

  • Apache Oozie

    An extensible, scalable and reliable system to manage complex Hadoop workloads

  • Azkaban

    Batch workflow job scheduler created at LinkedIn to run Hadoop jobs

  • Dagster

    An orchestration platform for the development, production, and observation of data assets

  • Luigi

    Python module that helps you build complex pipelines of batch jobs

  • Prefect

    A workflow management system, designed for modern infrastructure

Data Processing

  • Apache Beam

    A unified model for defining both batch and streaming data-parallel processing pipelines

  • Apache Flink

    An open source stream processing framework with powerful capabilities

  • Apache Hadoop MapReduce

    A framework for writing applications which process vast amounts of data

  • Apache Nifi

    An easy to use, powerful, and reliable system to process and distribute data

  • Apache Samza

    A distributed stream processing framework which uses Apache Kafka and Hadoop YARN

  • Apache Spark

    A unified analytics engine for large-scale data processing

  • Apache Storm

    An open source distributed realtime computation system

  • Apache Tez

    A generic data-processing pipeline engine envisioned as a low-level engine

  • Faust

    A stream processing library, porting the ideas from Kafka Streams to Python

Data Quality

  • Cerberus

    Lightweight, extensible data validation library for Python

  • Cleanlab

    Data-centric AI tool to detect (non-predefined) issues in ML data like label errors or outliers

  • DataProfiler

    A Python library designed to make data analysis, monitoring, and sensitive data detection easy

  • Deequ

    A library built on top of Apache Spark for measuring data quality in large datasets

  • Great Expectations

    A Python data validation framework that allows to test your data against datasets

  • JSON Schema

    A vocabulary that allows you to annotate and validate JSON documents

  • SodaSQL

    Data profiling, testing, and monitoring for SQL accessible data

Data Serialization

  • Apache Avro

    A data serialization system which is compact, fast and provides rich data structures

  • Apache ORC

    A self-describing type-aware columnar file format designed for Hadoop workloads

  • Apache Parquet

    A columnar storage format which provides efficient storage and encoding of data

  • Kryo

    A fast and efficient binary object graph serialization framework for Java

  • ProtoBuf

    Language-neutral, platform-neutral, extensible mechanism for serializing structured data

Data Serialization / Data Compression

  • Pigz

    A parallel implementation of gzip for modern multi-processor, multi-core machines

  • Snappy

    Open source compression library that is fast, stable and robuts

Data Serialization / Data Table Format

  • Apache Hudi

    Manages the storage of large analytical datasets on DFS

  • Apache Iceberg

    Open table format for huge analytic datasets

  • Delta Lake

    An open source project that enables building a Lakehouse architecture on top of data lakes

Data Visualization

  • Apache Superset

    A modern data exploration and data visualization platform

  • Count

    SQL/drag-and-drop querying and visualisation tool based on notebooks

  • Dash

    Analytical Web Apps for Python, R, Julia, and Jupyter

  • Data Studio

    Reporting solution for power users who want to go beyond the data and dashboards of GA

  • HUE

    A mature SQL Assistant for querying Databases & Data Warehouses

  • Lux

    Fast and easy data exploration by automating the visualization and data analysis process

  • Metabase

    The simplest, fastest way to get business intelligence and analytics to everyone

  • Redash

    Connect to any data source, easily visualize, dashboard and share your data

  • Tableau

    Powerful and fastest growing data visualization tool used in the business intelligence industry

Data Warehouse

  • Amazon Redshift

    Accelerate your time to insights with fast, easy, and secure cloud data warehousing

  • Apache Hive

    Facilitates reading, writing, and managing large datasets residing in distributed storage

  • Apache Kylin

    An open source, distributed analytical data warehouse for big data

  • Google BigQuery

    Serverless, highly scalable, and cost-effective multicloud data warehouse

Database / Columnar Database

  • Apache Cassandra

    Open source column based DBMS designed to handle large amounts of data

  • Apache Druid

    Designed to quickly ingest massive quantities of event data, and provide low-latency queries

  • Apache HBase

    An open-source, distributed, versioned, column-oriented store

  • Scylla

    Designed to be compatible with Cassandra while achieving higher throughputs and lower latencies

Database / Document-Oriented Database

  • Apache CouchDB

    An open-source document-oriented NoSQL database, implemented in Erlang

  • Elasticsearch

    A distributed document oriented database with a RESTful search engine

  • MongoDB

    A cross-platform document database that uses JSON-like documents with optional schemas

  • RethinkDB

    The first open-source scalable database built for realtime applications

Database / Graph Database

  • Age

    A multi-model database that supports both graph and relational data models

  • ArangoDB

    A scalable open-source multi-model database natively supporting graph, document and search

  • JanusGraph

    Manage large graphs with billions of data distributed across a multi-machine cluster

  • Memgraph

    An open source graph database, built for real-time streaming data, compatible with Neo4j

  • Neo4j

    A high performance graph store with all the features expected of a mature and robust database

  • Titan

    A highly scalable graph database optimized for storing and querying large graphs

Database / Key-Value Database

  • Apache Accumulo

    A sorted, distributed key-value store that provides robust and scalable data storage

  • Dragonfly

    A modern in-memory datastore, fully compatible with Redis and Memcached APIs

  • DynamoDB

    Fast, flexible NoSQL database service for single-digit millisecond performance at any scale

  • etcd

    Distributed reliable key-value store for the most critical data of a distributed system

  • EVCache

    A distributed in-memory data store for the cloud

  • Memcached

    A high performance multithreaded event-based key/value cache store

  • Redis

    An in-memory key-value database that persists on disk

Database / Relational Database

  • CockroachDB

    A distributed database designed to build, scale, and manage data-intensive apps

  • Crate

    A distributed SQL database that makes it simple to store and analyze massive amounts of data

  • MariaDB

    A replacement of MySQL with more features, new storage engines and better performance

  • MySQL

    One of the most popular open source transactional databases

  • PostgreSQL

    An advanced RDBMS that supports an extended subset of the SQL standard

  • RQLite

    A lightweight, distributed relational database, which uses SQLite as its storage engine

  • SQLite

    A popular choice as embedded database software for local/client storage

Database / Time Series Database

  • Akumuli

    Can be used to capture, store and process time-series data in real-time

  • Atlas

    An in-memory dimensional time series database

  • InfluxDB

    Scalable datastore for metrics, events, and real-time analytics

  • QuestDB

    An open source SQL database designed to process time series data, faster

  • TimescaleDB

    Open-source time-series SQL database optimized for fast ingest and complex queries

Database / Vector Database

  • Milvus

    An open source embedding vector similarity search engine powered by Faiss, NMSLIB and Annoy

  • Pinecone

    Managed and distributed vector similarity search used with a lightweight SDK

  • Qdrant

    An open source vector similarity search engine with extended filtering support

File System

  • Alluxio

    A virtual distributed storage system

  • Amazon Simple Storage Service (S3)

    Object storage built to retrieve any amount of data from anywhere

  • GlusterFS

    A software defined distributed storage that can scale to several petabytes

  • Google Cloud Storage (GCS)

    Object storage for companies of all sizes, to store any amount of data

  • LakeFS

    Open source tool that transforms your object storage into a Git-like repository

  • LizardFS

    A highly reliable, scalable and efficient distributed file system

  • MinIO

    High Performance, Kubernetes Native Object Storage compatible with Amazon S3 API

  • SeaweedFS

    A fast distributed storage system for blobs, objects, files, and data lake

  • Swift

    A distributed object storage system designed to scale from a single machine to thousands of servers

Logging and Monitoring

  • Grafana

    Visualize metrics, logs, and traces from multiple sources like Prometheus, Loki, InfluxDB and more

  • Loki

    A horizontally-scalable, highly-available, multi-tenant log aggregation system inspired by Prometheus

  • Prometheus

    A monitoring system and time series database

  • Whylogs

    A tool for creating data logs, enabling monitoring for data drift and data quality issues

Metadata Service

  • Hive Metastore

    Service that stores metadata related to Apache Hive and other services

  • Metacat

    Provides you information about what data you have, where it resides and how to process it

SQL Query Engine

  • Apache Drill

    Schema-free SQL Query Engine for Hadoop, NoSQL and Cloud Storage

  • Apache Impala

    Lightning-fast, distributed SQL queries for petabytes of data

  • Dremio

    Power high-performing BI dashboards and interactive analytics directly on data lake

  • Presto

    A distributed SQL query engine for big data

  • Trino

    A fast distributed SQL query engine for big data analytics

Resources / Books

Resources / Other Lists

Resources / Slack

More related projects

Add a GitHub project

Missing a project or an awesome list? Paste its GitHub URL and we fetch it right away.