awesome-bigdata

Big Data Platform

A curated collection of resources and frameworks for working with big data

A curated list of awesome big data frameworks, ressources and other awesomeness.

GitHub

13k stars
845 watching
3k forks
last commit: over 2 years ago
Linked from 14 awesome lists

awesomeawesome-listbigdatadatadata-analyticsdata-sciencedata-streamdata-visualizationdata-warehousedatabasedistributed-databaseseries-databasestream-processingstreaming-datavisualize-data

Awesome Big Data / RDBMS

MySQLThe world's most popular open source database
PostgreSQLThe world's most advanced open source database
Oracle Databaseobject-relational database management system
Teradatahigh-performance MPP data warehouse platform

Awesome Big Data / Frameworks

Bistro1,033over 3 years agogeneral-purpose data processing engine for both batch and stream analytics. It is based on a novel data model, which represents data via and processes data via as opposed to having only set operations in conventional approaches like MapReduce or SQL
IBM Streamsplatform for distributed processing and real-time analytics. Integrates with many of the popular technologies in the Big Data ecosystem (Kafka, HDFS, Spark, etc.)
Apache Hadoopframework for distributed processing. Integrates MapReduce (parallel processing), YARN (job scheduling) and HDFS (distributed file system)
Tigon284over 9 years agoHigh Throughput Real-time Stream Processing Framework
PachydermPachyderm is a data storage platform built on Docker and Kubernetes to provide reproducible data processing and analysis
Polyaxon3,581almost 2 years agoA platform for reproducible and scalable machine learning and deep learning
Smooks400almost 2 years agoAn extensible Java framework for building XML and non-XML (CSV, EDI, Java, etc...) streaming applications

Awesome Big Data / Distributed Programming

AddThis Hydra434about 6 years agodistributed data processing and storage system originally developed at AddThis
AMPLab SIMRrun Spark on Hadoop MapReduce v1
Apache APEXa unified, enterprise platform for big data stream and batch processing
Apache Beaman unified model and set of language-specific SDKs for defining and executing data processing workflows
Apache Cruncha simple Java API for tasks like joining and data aggregation that are tedious to implement on plain MapReduce
Apache DataFucollection of user-defined functions for Hadoop and Pig developed by LinkedIn
Apache Flinkhigh-performance runtime, and automatic program optimization
Apache Gearpumpreal-time big data streaming engine based on Akka
Apache Goraframework for in-memory data model and persistence
Apache HamaBSP (Bulk Synchronous Parallel) computing framework
Apache MapReduceprogramming model for processing large data sets with a parallel, distributed algorithm on a cluster
Apache Pighigh level language to express data analysis programs for Hadoop
Apache REEFretainable evaluator execution framework to simplify and unify the lower layers of big data systems
Apache S4framework for stream processing, implementation of S4
Apache Sparkframework for in-memory cluster computing
Apache Spark Streamingframework for stream processing, part of Spark
Apache Stormframework for stream processing by Twitter also on YARN
Apache Samzastream processing framework, based on Kafka and YARN
Apache Tezapplication framework for executing a complex DAG (directed acyclic graph) of tasks, built on YARN
Apache Twillabstraction over YARN that reduces the complexity of developing distributed applications
Baidu Bigflowan interface that allows for writing distributed computing programs providing lots of simple, flexible, powerful APIs to easily handle data of any scale
Cascalogdata processing and querying library
CheetahHigh Performance, Custom Data Warehouse on Top of MapReduce
Concurrent Cascadingframework for data management/analytics on Hadoop
Damballa Parkour257over 10 years agoMapReduce library for Clojure
Datasalt Pangool57about 4 years agoalternative MapReduce paradigm
DataTorrent StrAMreal-time engine is designed to enable distributed, asynchronous, real time in-memory big-data computations in as unblocked a way as possible, with minimal overhead and impact on performance
Facebook CoronaHadoop enhancement which removes single point of failure
Facebook PeregrineMap Reduce framework
Facebook Scubadistributed in-memory datastore
Google Dataflowcreate data pipelines to help themæingest, transform and analyze data
Google MapReducemap reduce framework
Google MillWheelfault tolerant stream processing framework
IBM Streamsplatform for distributed processing and real-time analytics. Provides toolkits for advanced analytics like geospatial, time series, etc. out of the box
JAQLdeclarative programming language for working with structured, semi-structured and unstructured data
Kiteis a set of libraries, tools, examples, and documentation focused on making it easier to build systems on top of the Hadoop ecosystem
Metamarkets Druidframework for real-time analysis of large datasets
Netflix PigPen567over 3 years agomap-reduce for Clojure which compiles to Apache Pig
Nokia DiscoMapReduce framework developed by Nokia
OnyxDistributed computation for the cloud
Pinterest Pinlaterasynchronous job execution system
PydoopPython MapReduce and HDFS API for Hadoop
Ray34,412almost 2 years agoA fast and simple framework for building and running distributed applications
Rackerlabs Bluefloodmulti-tenant distributed metric processing system
Skale397over 5 years agoHigh performance distributed data processing in NodeJS
Stratospheregeneral purpose cluster computing framework
Streamdrilluseful for counting activities of event streams over different time windows and finding the most active one
streamsx.topology29about 4 years agoLibraries to enable building IBM Streams application in Java, Python or Scala
Tuktu60over 8 years agoEasy-to-use platform for batch and streaming computation, built using Scala, Akka and Play!
Twitter Heron3,638over 3 years agoHeron is a realtime, distributed, fault-tolerant stream processing engine from Twitter replacing Storm
Twitter Scalding3,506over 3 years agoScala library for Map Reduce jobs, built on Cascading
Twitter Summingbird2,135over 4 years agoStreaming MapReduce with Scalding and Storm, by Twitter
Twitter TSARTimeSeries AggregatoR by Twitter
WallarooThe ultrafast and elastic data processing engine. Big or fast data - no fuss, no Java needed

Awesome Big Data / Distributed Filesystem

Ambry1,749almost 2 years agoa distributed object store that supports storage of trillion of small immutable objects as well as billions of large objects
Apache HDFSa way to store large files across multiple machines
Apache KuduHadoop's storage layer to enable fast analytics on fast data
BeeGFSformerly FhGFS, parallel distributed file system
Ceph Filesystemsoftware storage platform designed
Disco DDFSdistributed filesystem
Facebook Haystackobject storage system
Google GFSdistributed filesystem
Google Megastorescalable, highly available storage
GridGainGGFS, Hadoop compliant in-memory file system
Lustre file systemhigh-performance distributed filesystem
Microsoft Azure Data Lake StoreHDFS-compatible storage in Azure cloud
Quantcast File System QFSopen-source distributed file system
Red Hat GlusterFSscale-out network-attached storage file system
Seaweed-FS23,207almost 2 years agosimple and highly scalable distributed file system
Alluxioreliable file sharing at memory speed across cluster frameworks
Tahoe-LAFSdecentralized cloud storage system
Baidu File System2,853almost 8 years agodistributed filesystem

Awesome Big Data / Distributed Index

Pilosa2,526over 2 years agoOpen source distributed bitmap index that dramatically accelerates queries across multiple, massive data sets

Awesome Big Data / Document Data Model

Actian Versantcommercial object-oriented database management systems
Crate Datais an open source massively scalable data store. It requires zero administration
Facebook ApolloFacebook’s Paxos-like NoSQL database
jumboDBdocument oriented datastore over Hadoop
LinkedIn Espressohorizontally scalable document-oriented NoSQL data store
MarkLogicSchema-agnostic Enterprise NoSQL database technology
Microsoft Azure DocumentDBNoSQL cloud database service with protocol support for MongoDB
MongoDBDocument-oriented database system
RavenDBA transactional, open-source Document Database
RethinkDBdocument database that supports queries like table joins and group by

Awesome Big Data / Key Map Data Model

Apache Accumulodistributed key/value store, built on Hadoop
Apache Cassandracolumn-oriented distributed datastore, inspired by BigTable
Apache HBasecolumn-oriented distributed datastore, inspired by BigTable
Baidu Tera1,890over 2 years agoan Internet-scale database, inspired by BigTable
Facebook HydraBaseevolution of HBase made by Facebook
Google BigTablecolumn-oriented distributed datastore
Google Cloud Datastoreis a fully managed, schemaless database for storing non-relational data over BigTable
Hypertablecolumn-oriented distributed datastore, inspired by BigTable
InfiniDB250almost 9 years agois accessed through a MySQL interface and use massive parallel processing to parallelize queries
Tephra157about 2 years agoTransactions for HBase
Twitter Manhattanreal-time, multi-tenant distributed database for Twitter scale
ScyllaDBcolumn-oriented distributed datastore written in C++, totally compatible with Apache Cassandra

Awesome Big Data / Key-value Data Model

AerospikeNoSQL flash-optimized, in-memory. Open source and "Server code in 'C' (not Java or Erlang) precisely tuned to avoid context switching and memory copies."
Amazon DynamoDBdistributed key/value store, implementation of Dynamo paper
Badgera fast, simple, efficient, and persistent key-value store written natively in Go
Bolt14,259over 8 years agoan embedded key-value database for Go
BTDB139almost 2 years agoKey Value Database in .Net with Object DB Layer, RPC, dynamic IL and much more
BuntDB4,599about 2 years agoa fast, embeddable, in-memory key/value database for Go with custom indexing and geospatial support
Edis467about 11 years agois a protocol-compatible Server replacement for Redis
ElephantDB559about 12 years agoDistributed database specialized in exporting data from Hadoop
EventStoredistributed time series database
GhostDB751over 5 years agoa distributed, in-memory, general purpose key-value data store that delivers microsecond performance at any scale
Graviton419over 4 years agoa simple, fast, versioned, authenticated, embeddable key-value store database in pure Go(lang)
GridDB2,387almost 2 years agosuitable for sensor data stored in a timeseries
HyperDex1,393over 2 years agoa scalable, next generation key-value and document store with a wide array of features, including consistency, fault tolerance and high performance
Igniteis an in-memory key-value data store providing full SQL-compliant data access that can optionally be backed by disk storage
LinkedIn Krati26about 14 years agois a simple persistent data store with very low latency and high throughput
Linkedin Voldemortdistributed key/value storage system
Oracle NoSQL Databasedistributed key-value database by Oracle Corporation
Redisin memory key value datastore
Riak3,962over 2 years agoa decentralized datastore
Storehaus464about 6 years agolibrary to work with asynchronous key value stores, by Twitter
SummitDB1,410over 4 years agoan in-memory, NoSQL key/value database, with disk persistance and using the Raft consensus algorithm
Tarantool3,437almost 2 years agoan efficient NoSQL database and a Lua application server
TiKV15,385almost 2 years agoa distributed key-value database powered by Rust and inspired by Google Spanner and HBase
Tile389,185almost 2 years agoa geolocation data store, spatial index, and realtime geofence, supporting a variety of object types including latitude/longitude points, bounding boxes, XYZ tiles, Geohashes, and GeoJSON
TreodeDB176almost 11 years agokey-value store that's replicated and sharded and provides atomic multirow writes

Awesome Big Data / Graph Data Model

AgensGrapha new generation multi-model graph database for the modern complex data environment
Apache Giraphimplementation of Pregel, based on Hadoop
Apache Spark Bagelimplementation of Pregel, part of Spark
ArangoDBmulti model distributed database
DGraph20,516almost 2 years agoA scalable, distributed, low latency, high throughput graph database aimed at providing Google production level scale and throughput, with low enough latency to be serving real time user queries, over terabytes of structured data
EliasDB1,006about 4 years agoa lightweight graph based database that does not require any third-party libraries
Facebook TAOTAO is the distributed data store that is widely used at facebook to store and serve the social graph
GCHQ Gaffer1,774almost 2 years agoGaffer by GCHQ is a framework that makes it easy to store large-scale graphs in which the nodes and edges have statistics
Google Cayley14,868almost 2 years agoopen-source graph database
Google Pregelgraph processing framework
GraphLab PowerGrapha core C++ GraphLab API and a collection of high-performance machine learning and data mining toolkits built on top of the GraphLab API
GraphXresilient Distributed Graph System on Spark
Gremlin1,950about 5 years agograph traversal Language
Infovore148almost 5 years agoRDF-centric Map/Reduce framework
Intel GraphBuildertools to construct large-scale graphs on top of Hadoop
JanusGraphopen-source, distributed graph database with multiple options for storage backends (Bigtable, HBase, Cassandra, etc.) and indexing backends (Elasticsearch, Solr, Lucene)
MapGraphMassively Parallel Graph processing on GPUs
Microsoft Graph Engine2,210almost 2 years agoa distributed in-memory data processing engine, underpinned by a strongly-typed in-memory key-value store and a general distributed computation engine
Neo4jgraph database written entirely in Java
OrientDBdocument and graph database
Phoebus384over 14 years agoframework for large scale graph processing
Titandistributed graph database, built over Cassandra
Twitter FlockDB3,337over 9 years agodistributed graph database
NodeXLA free, open-source template for Microsoft® Excel® 2007, 2010, 2013 and 2016 that makes it easy to explore network graphs

Awesome Big Data / Columnar Databases

Columnar Storagean explanation of what columnar storage is and when you might want it
Actian Vectorcolumn-oriented analytic database
ClickHousean open-source column-oriented database management system that allows generating analytical data reports in real time
EventQLa distributed, column-oriented database built for large-scale event collection and analytics
MonetDBcolumn store database
Parquetcolumnar storage format for Hadoop
Pivotal Greenplumpurpose-built, dedicated analytic data warehouse that offers a columnar engine as well as a traditional row-based one
Verticais designed to manage large, fast-growing volumes of data and provide very fast query performance when used for data warehouses
SQream DBA GPU powered big data database, designed for analytics and data warehousing, with ANSI-92 compliant SQL, suitable for data sets from 10TB to 1PB
Google BigQueryGoogle's cloud offering backed by their pioneering work on Dremel
Amazon RedshiftAmazon's cloud offering, also based on a columnar datastore backend
IndexR453almost 4 years agoan open-source columnar storage format for fast & realtime analytic with big data
LocustDB1,617about 2 years agoan experimental analytics database aiming to set a new standard for query performance on commodity hardware

Awesome Big Data / NewSQL Databases

Actian Ingrescommercially supported, open-source SQL relational database management system
ActorDB1,897almost 4 years agoa distributed SQL database with the scalability of a KV store, while keeping the query capabilities of a relational database
Amazon RedShiftdata warehouse service, based on PostgreSQL
BayesDB891almost 11 years agostatistic oriented SQL database
Bedrocka simple, modular, networked and distributed transaction layer built atop SQLite
CitusDBscales out PostgreSQL through sharding and replication
Cockroach30,270almost 2 years agoScalable, Geo-Replicated, Transactional Datastore
Comdb21,396almost 2 years agoa clustered RDBMS built on optimistic concurrency control techniques
Datomicdistributed database designed to enable scalable, flexible and intelligent applications
FoundationDBdistributed database, inspired by F1
Google F1distributed SQL database built on Spanner
Google Spannerglobally distributed semi-relational database
H-Storeis an experimental main-memory, parallel database management system that is optimized for on-line transaction processing (OLTP) applications
Haeinsa158over 9 years agolinearly scalable multi-row, multi-table transaction library for HBase based on Percolator
HandlerSocketNoSQL plugin for MySQL/MariaDB
InfiniSQLinfinity scalable RDBMS
KarelDB393almost 2 years agoa relational database backed by Apache Kafka
Map-DGPU in-memory database, big data analysis and visualization platform
MemSQLin memory SQL database witho optimized columnar storage on flash
NuoDBSQL/ACID compliant distributed database
Oracle TimesTen in-Memory Databasein-memory, relational database management system with persistence and recoverability
Pivotal GemFire XDLow-latency, in-memory, distributed SQL data store. Provides SQL interface to in-memory table data, persistable in HDFS
SAP HANAis an in-memory, column-oriented, relational database management system
SenseiDBdistributed, realtime, semi-structured database
Skydatabase used for flexible, high performance analysis of behavioral data
SymmetricDSopen source software for both file and database synchronization
TiDB37,447almost 2 years agoTiDB is a distributed SQL database. Inspired by the design of Google F1
VoltDBclaims to be fastest in-memory database
yugabyteDB9,085almost 2 years agoopen source, high-performance, distributed SQL database compatible with PostgreSQL

Awesome Big Data / Time-Series Databases

Axibase Time Series DatabaseIntegrated time series database on top of HBase with built-in visualization, rule-engine and SQL support
Chronixa time series storage built to store time series highly compressed and for fast access times
Cubeuses MongoDB to store time series data
Heroicis a scalable time series database based on Cassandra and Elasticsearch
InfluxDBa time series database with optimised IO and queries, supports pgsql and influx wire protocols
QuestDBhigh-performance, open-source SQL database for applications in financial services, IoT, machine learning, DevOps and observability
IronDBscalable, general-purpose time series database
Kairosdb1,740almost 2 years agosimilar to OpenTSDB but allows for Cassandra
M3DBa distributed time series database that can be used for storing realtime metrics at long retention
Newtsa time series database based on Apache Cassandra
TDengine23,479almost 2 years agoa time series database in C utilizing unique features of IoT to improve read/write throughput and reduce space needed to store data
OpenTSDBdistributed time series database on top of HBase
Prometheusa time series database and service monitoring system
Beringei3,170about 8 years agoFacebook's in-memory time-series database
TrailDBan efficient tool for storing and querying series of events
Druid13,548almost 2 years agoColumn oriented distributed data store ideal for powering interactive applications
Riak-TSRiak TS is the only enterprise-grade NoSQL time series database optimized specifically for IoT and Time Series data
Akumuli835about 4 years agoAkumuli is a numeric time-series database. It can be used to capture, store and process time-series data in real-time. The word "akumuli" can be translated from esperanto as "accumulate"
RhombusA time-series object store for Cassandra that handles all the complexity of building wide row indexes
Dalmatiner DB694over 7 years agoFast distributed metrics database
Blueflood595about 2 years agoA distributed system designed to ingest and process time series data
Timely379about 2 years agoTimely is a time series database application that provides secure access to time series data based on Accumulo and Grafana
SiriDB506almost 2 years agoHighly-scalable, robust and fast, open source time series database with cluster functionality
Thanos13,177almost 2 years agoThanos is a set of components to create a highly available metric system with unlimited storage capacity using multiple (existing) Prometheus deployments
VictoriaMetrics12,774almost 2 years agofast, scalable and resource-effective open-source TSDB compatible with Prometheus. Single-node and cluster versions included

Awesome Big Data / SQL-like processing

Actian SQL for Hadoophigh performance interactive SQL access to all Hadoop data
Apache Drillframework for interactive analysis, inspired by Dremel
Apache HCatalogtable and storage management layer for Hadoop
Apache HiveSQL-like data warehouse system for Hadoop
Apache Calciteframework that allows efficient translation of queries involving heterogeneous and federated data
Apache PhoenixSQL skin over HBase
Aster DatabaseSQL-like analytic processing for MapReduce
Cloudera Impalaframework for interactive analysis, Inspired by Dremel
Concurrent LingualSQL-like query language for Cascading
Datasalt Splout SQLfull SQL query engine for big datasets
Dremioan open-source, SQL-like Data-as-a-Service Platform based on Apache Arrow
Facebook PrestoDBdistributed SQL query engine
Google BigQueryframework for interactive analysis, implementation of Dremel
Materialize5,834almost 2 years agois a streaming database for real-time applications using SQL for queries and supporting a large fraction of PostgreSQL
Invantive SQLSQL engine for online and on-premise use with integrated local data replication and 70+ connectors
PipelineDBan open-source relational database that runs SQL queries continuously on streams, incrementally storing results in tables
Pivotal HDBSQL-like data warehouse system for Hadoop
RainstorDBdatabase for storing petabyte-scale volumes of structured and semi-structured data
Spark Catalyst40,170almost 2 years agois a Query Optimization Framework for Spark and Shark
SparkSQLManipulating Structured Data Using Spark
Splice Machinea full-featured SQL-on-Hadoop RDBMS with ACID transactions
Stingerinteractive query for Hive
Tajodistributed data warehouse system on Hadoop
Trafodionenterprise-class SQL-on-HBase solution targeting big data transactional or operational workloads

Awesome Big Data / Data Ingestion

redpandaA Kafka® replacement for mission critical systems; 10x faster. Written in C++
Amazon Kinesisreal-time processing of streaming data at massive scale
Amazon Web Services Glueserverless fully managed extract, transform, and load (ETL) service
CensusA reverse ETL product that let you sync data from your data warehouse to SaaS Applications. No engineering favors required—just SQL
Apache Chukwadata collection system
Apache Flumeservice to manage large amount of log data
Apache Kafkadistributed publish-subscribe messaging system
Apache NiFiApache NiFi is an integrated data logistics platform for automating the movement of data between disparate systems
Apache Pulsar14,315almost 2 years agoa distributed pub-sub messaging platform with a very flexible messaging model and an intuitive client API
Apache Sqooptool to transfer data between Hadoop and a structured datastore
Embulkopen-source bulk data loader that helps data transfer between various databases, storages, file formats, and cloud services
Facebook Scribe3,920about 6 years agostreamed log data aggregator
Fluentdtool to collect events and logs
Gazette723almost 2 years agoDistributed streaming infrastructure built on cloud storage which makes it easy to mix and match batch and streaming paradigms
Google Photongeographically distributed system for joining multiple continuously flowing streams of data in real-time with high scalability and low latency
Heka3,389over 2 years agoopen source stream processing software system
HIHO91over 13 years agoframework for connecting disparate data sources with Hadoop
Kestreldistributed message queue system
LinkedIn Databusstream of change capture events for a database
LinkedIn Kamikaze22over 12 years agoutility package for compressing sorted integer arrays
LinkedIn White Elephant191almost 13 years agolog aggregator and dashboard
Logstasha tool for managing events and logs
Netflix Suro794over 3 years agolog agregattor like Storm and Samza based on Chukwa
Pinterest Secor1,846almost 2 years agois a service implementing Kafka log persistance
Linkedin Gobblin2,232almost 2 years agolinkedin's universal data ingestion framework
Skizze771over 10 years agosketch data store to deal with all problems around counting and sketching using probabilistic data-structures
StreamSets Data Collectorcontinuous big data ingest infrastructure with a simple to use IDE
Aloomadata pipeline as a service enabling moving data sources such as MySQL into data warehouses
RudderStack4,109almost 2 years agoan open source customer data infrastructure (segment, mParticle alternative) written in go
Zilla553almost 2 years agoAn API gateway built for event-driven architectures and streaming that supports standard protocols such as HTTP, SSE, gRPC, MQTT and the native Kafka protocol

Awesome Big Data / Service Programming

Akka Toolkitruntime for distributed, and fault tolerant event-driven applications on the JVM
Apache Avrodata serialization system
Apache CuratorJava libaries for Apache ZooKeeper
Apache KarafOSGi runtime that runs on top of any OSGi framework
Apache Thriftframework to build binary protocols
Apache Zookeepercentralized service for process management
Google Chubbya lock service for loosely-coupled distributed systems
Hydrosphere Mist326almost 6 years agoa service for exposing Apache Spark analytics jobs and machine learning models as realtime, batch or reactive web services
Linkedin Norbertcluster manager
Mara2,082over 2 years agoA lightweight opinionated ETL framework, halfway between plain scripts and Apache Airflow
OpenMPImessage passing framework
Serfdecentralized solution for service discovery and orchestration
Spotify Luigi17,950almost 2 years agoa Python package for building complex pipelines of batch jobs. It handles dependency resolution, workflow management, visualization, handling failures, command line integration, and much more
Spring XD475over 4 years agodistributed and extensible system for data ingestion, real time analytics, batch processing, and data export
Twitter Elephant Bird1,137over 3 years agolibraries for working with LZOP-compressed data
Twitter Finagleasynchronous network stack for the JVM

Awesome Big Data / Scheduling

Apache Airflow37,580almost 2 years agoa platform to programmatically author, schedule and monitor workflows
Apache Aurorais a service scheduler that runs on top of Apache Mesos
Apache Falcondata management framework
Apache Oozieworkflow job scheduler
Azure Data Factorycloud-based pipeline orchestration for on-prem, cloud and HDInsight
Chronosdistributed and fault-tolerant scheduler
Cronicle3,975almost 2 years agoDistributed, easy to install, NodeJS based, task scheduler
Dagster12,055almost 2 years agoa data orchestrator for machine learning, analytics, and ETL
Linkedin Azkabanbatch workflow job scheduler
Schedoscope96almost 7 years agoScala DSL for agile scheduling of Hadoop jobs
Sparrow319about 6 years agoscheduling platform

Awesome Big Data / Machine Learning

Azure ML StudioCloud-based AzureML, R, Python Machine Learning platform
brain8,010about 6 years agoNeural networks in JavaScript
Oryx1,787about 5 years agoLambda architecture on Apache Spark, Apache Kafka for real-time large scale machine learning
Concurrent Patternmachine learning library for Cascading
convnetjs10,902over 3 years agoDeep Learning in Javascript. Train Convolutional Neural Networks (or ordinary ones) in your browser
DataVecA vectorization and data preprocessing library for deep learning in Java and Scala. Part of the Deeplearning4j ecosystem
Deeplearning4jFast, open deep learning for the JVM (Java, Scala, Clojure). A neural network configuration layer powered by a C++ library. Uses Spark and Hadoop to train nets on multiple GPUs and CPUs
Decider385over 9 years agoFlexible and Extensible Machine Learning in Ruby
ENCOGmachine learning framework that supports a variety of advanced algorithms, as well as support classes to normalize and process data
etcMLtext classification with machine learning
Etsy Conjecture359over 8 years agoscalable Machine Learning in Scalding
Feast5,669almost 2 years agoA feature store for the management, discovery, and access of machine learning features. Feast provides a consistent view of feature data for both model training and model serving
GraphLab CreateA machine learning platform in Python with a broad collection of ML toolkits, data engineering, and deployment tools
H2O6,950almost 2 years agostatistical, machine learning and math runtime with Hadoop. R and Python
Karate Club2,178about 2 years agoAn unsupervised machine learning library for graph structured data. Python
Keras62,196almost 2 years agoAn intuitive neural net API inspired by Torch that runs atop Theano and Tensorflow
Lambdo1about 8 years agoLambdo is a workflow engine which significantly simplifies the analysis process by unifying feature engineering and machine learning operations
Little Ball of Fur705over 2 years agoA subsampling library for graph structured data. Python
MahoutAn Apache-backed machine learning library for Hadoop
MLbasedistributed machine learning libraries for the BDAS stack
MLPNeuralNet900almost 10 years agoFast multilayer perceptron neural network library for iOS and Mac OS X
ML Workspace3,446about 2 years agoAll-in-one web-based IDE specialized for machine learning and data science
MOAMOA performs big data stream mining in real time, and large scale machine learning
MonkeyLearnText mining made easy. Extract and classify data from text
ND4JA matrix library for the JVM. Numpy for Java
nupic6,340almost 2 years agoNumenta Platform for Intelligent Computing: a brain-inspired machine intelligence platform, and biologically accurate neural network based on cortical learning algorithms
PredictionIOmachine learning server buit on Hadoop, Mahout and Cascading
PyTorch Geometric Temporal2,694almost 2 years agoa temporal extension library for PyTorch Geometric
RL4JReinforcement learning for Java and Scala. Includes Deep-Q learning and A3C algorithms, and integrates with Open AI's Gym. Runs in the Deeplearning4j ecosystem
SAMOAdistributed streaming machine learning framework
scikit-learn60,451almost 2 years agoscikit-learn: machine learning in Python
Shapley219about 3 years agoA data-driven framework to quantify the value of classifiers in a machine learning ensemble
Spark MLliba Spark implementation of some common machine learning (ML) functionality
SibylSystem for Large Scale Machine Learning at Google
TensorFlow186,822almost 2 years agoLibrary from Google for machine learning using data flow graphs
TheanoA Python-focused machine learning library supported by the University of Montreal
TorchA deep learning library with a Lua API, supported by NYU and Facebook
Velox110over 9 years agoSystem for serving machine learning predictions
Vowpal Wabbit8,495almost 2 years agolearning system sponsored by Microsoft and Yahoo!
WEKAsuite of machine learning software
BidMach915almost 4 years agoCPU and GPU-accelerated Machine Learning Library

Awesome Big Data / Benchmarking

Apache Hadoop Benchmarkingmicro-benchmarks for testing Hadoop performances
Berkeley SWIM Benchmarkreal-world big data workload benchmark
Intel HiBench1,463almost 2 years agoa Hadoop benchmark suite
PUMA Benchmarkingbenchmark suite for MapReduce applications
Yahoo Gridmix3Hadoop cluster benchmarking from Yahoo engineer team
Deeplearning4j Benchmarks
UCSB51about 3 years agoextended Yahoo Cloud Serving Benchmark for NoSQL databases

Awesome Big Data / Security

Apache RangerCentral security admin & fine-grained authorization for Hadoop
Apache Eaglereal time monitoring solution
Apache Knox Gatewaysingle point of secure access for Hadoop clusters
Apache Sentrysecurity module for data stored in Hadoop
BDA104about 6 years agoThe vulnerability detector for Hadoop and Spark

Awesome Big Data / System Deployment

Apache Ambarioperational framework for Hadoop mangement
Apache Bigtopsystem deployment framework for the Hadoop ecosystem
Apache Helixcluster management framework
Apache Mesoscluster manager
Apache Slider79almost 8 years agois a YARN application to deploy existing distributed applications on YARN
Apache Whirrset of libraries for running cloud services
Apache YARNCluster manager
Brooklynlibrary that simplifies application deployment and management
BuildoopSimilar to Apache BigTop based on Groovy language
Cloudera HUEweb application for interacting with Hadoop
Facebook Prismmulti datacenters replication system
Google Borgjob scheduling and monitoring system
Google Omegajob scheduling and monitoring system
Hortonworks HOYAapplication that can deploy HBase cluster on YARN
Kubernetesa system for automating deployment, scaling, and management of containerized applications
Marathon4,064about 4 years agoMesos framework for long-running services
Linkis3,324almost 2 years agoLinkis helps easily connect to various back-end computation/storage engines

Awesome Big Data / Applications

411973over 3 years agoan web application for alert management resulting from scheduled searches into Elasticsearch
Adobe spindle331over 11 years agoNext-generation web analytics processing with Scala, Spark, and Parquet
Apache Metrona platform that integrates a variety of open source big data technologies in order to offer a centralized tool for security monitoring and analysis
Apache Nutchopen source web crawler
Apache OODTcapturing, processing and sharing of data for NASA's scientific archives
Apache Tikacontent analysis toolkit
Argus506over 4 years agoTime series monitoring and alerting platform
AthenaX1,222about 6 years agoa streaming analytics platform that enables users to run production-quality, large scale streaming analytics using Structured Query Language (SQL)
Atlas3,459almost 2 years agoa backend for managing dimensional time series data
Countlyopen source mobile and web analytics platform, based on Node.js & MongoDB
DominoRun, scale, share, and deploy models — without any infrastructure
Eclipse BIRTEclipse-based reporting system
ElastAert8,004about 2 years agoElastAlert is a simple framework for alerting on anomalies, spikes, or other patterns of interest from data in ElasticSearch
Eventhub1,332over 4 years agoopen source event analytics platform
HASHopen source simulation and visualization platform
Hermes820almost 2 years agoasynchronous message broker built on top of Kafka
HunkSplunk analytics for Hadoop
ImhotepLarge scale analytics platform by indeed
IndicativeWeb & mobile analytics tool, with data warehouse (AWS, BigQuery) integration
JupyterNotebook and project application for interactive data science and scientific computing across all programming languages
MADlibdata-processing library of an RDBMS to analyze data
Kapacitor2,321almost 2 years agoan open source framework for processing, monitoring, and alerting on time series data
Kylinopen source Distributed Analytics Engine from eBay
PivotalR126almost 4 years agoR on Pivotal HD / HAWQ and PostgreSQL
Rakam799almost 5 years agoopen-source real-time custom analytics platform powered by Postgresql, Kinesis and PrestoDB
Quboleauto-scaling Hadoop cluster, built-in data connectors
SnappyData1,041almost 4 years agoa distributed in-memory data store for real-time operational analytics, delivering stream analytics, OLTP (online transaction processing) and OLAP (online analytical processing) built on Spark in a single integrated cluster
Snowplow6,852about 2 years agoenterprise-strength web and event analytics, powered by Hadoop, Kinesis, Redshift and Postgres
SparkRR frontend for Spark
Splunkanalyzer for machine-generated data
Sumo Logiccloud based analyzer for machine-generated data
Substation332almost 2 years agoSubstation is a cloud native data pipeline and transformation toolkit written in Go
Talendunified open source environment for YARN, Hadoop, HBASE, Hive, HCatalog & Pig

Awesome Big Data / Search engine and framework

Apache LuceneSearch engine library
Apache SolrSearch platform for Apache Lucene
Elassandra1,718over 2 years agois a fork of Elasticsearch modified to run on top of Apache Cassandra in a scalable and resilient peer-to-peer architecture
ElasticSearchSearch and analytics engine based on Apache Lucene
Enigma.io– Freemium robust web application for exploring, filtering, analyzing, searching and exporting massive datasets scraped from across the Web
Google Caffeinecontinuous indexing system
Google Percolatorcontinuous indexing system
HBase Coprocessorimplementation of Percolator, part of HBase
Lily HBase Indexerquickly and easily search for any content stored in HBase
LinkedIn Bobois a Faceted Search implementation written purely in Java, an extension to Apache Lucene
LinkedIn Cleo565almost 13 years agois a flexible software library for enabling rapid development of partial, out-of-order and real-time typeahead search
LinkedIn Galenesearch architecture at LinkedIn
LinkedIn Zoie368over 3 years agois a realtime search/indexing system written in Java
MG4JMG4J (Managing Gigabytes for Java) is a full-text search engine for large document collections written in Java. It is highly customisable, high-performance and provides state-of-the-art features and new research algorithms
Sphinx Search Serverfulltext search engine
Vespais an engine for low-latency computation over large data sets. It stores and indexes your data such that queries, selection and processing over the data can be performed at serving time
Facebook Faiss31,920almost 2 years agois a library for efficient similarity search and clustering of dense vectors. It contains algorithms that search in sets of vectors of any size, up to ones that possibly do not fit in RAM. It also contains supporting code for evaluation and parameter tuning. Faiss is written in C++ with complete wrappers for Python/numpy
Annoy13,321about 2 years agois a C++ library with Python bindings to search for points in space that are close to a given query point. It also creates large read-only file-based data structures that are mmapped into memory so that many processes may share the same data
Weaviate11,812over 1 year agoWeaviate is a GraphQL-based semantic search engine with build-in (word) embeddings

Awesome Big Data / MySQL forks and evolutions

Amazon RDSMySQL databases in Amazon's cloud
Drizzleevolution of MySQL 6.0
Google Cloud SQLMySQL databases in Google's cloud
MariaDBenhanced, drop-in replacement for MySQL
MySQL ClusterMySQL implementation using NDB Cluster storage engine
Percona Serverenhanced, drop-in replacement for MySQL
ProxySQL25almost 9 years agoHigh Performance Proxy for MySQL
TokuDBTokuDB is a storage engine for MySQL and MariaDB
WebScaleSQLis a collaboration among engineers from several companies that face similar challenges in running MySQL at scale

Awesome Big Data / PostgreSQL forks and evolutions

HadoopDBhybrid of MapReduce and DBMS
IBM Netezzahigh-performance data warehouse appliances
Postgres-XLScalable Open Source PostgreSQL-based Database Cluster
RecDBOpen Source Recommendation Engine Built Entirely Inside PostgreSQL
Stadoopen source MPP database system solely targeted at data warehousing and data mart applications
Yahoo Everestmulti-peta-byte database / MPP derived by PostgreSQL
TimescaleDBAn open-source time-series database optimized for fast ingest and complex queries
PipelineDBThe Streaming SQL Database. An open-source relational database that runs SQL queries continuously on streams, incrementally storing results in tables

Awesome Big Data / Memcached forks and evolutions

Facebook McDipperkey/value cache for flash storage
Facebook Memcachedfork of Memcache
Twemproxy12,162over 2 years agoA fast, light-weight proxy for memcached and redis
Twitter Fatcache1,301almost 5 years agokey/value cache for flash storage
Twitter Twemcache931almost 5 years agofork of Memcache

Awesome Big Data / Embedded Databases

Actian PSQLACID-compliant DBMS developed by Pervasive Software, optimized for embedding in applications
BerkeleyDBa software library that provides a high-performance embedded database for key/value data
HanoiDB306about 10 years agoErlang LSM BTree Storage
LevelDB36,769about 2 years agoa fast key-value storage library written at Google that provides an ordered mapping from string keys to string values
LMDBultra-fast, ultra-compact key-value embedded data store developed by Symas
RocksDBembeddable persistent key-value store for fast storage based on LevelDB

Awesome Big Data / Business Intelligence

BIME Analyticsbusiness intelligence platform in the cloud
Blazer4,581almost 2 years agobusiness intelligence made simple
Chartiolean business intelligence platform to visualize and explore your data
Countnotebook-based anlytics and visualisation platform using SQL or drag-and-drop
datapineself-service business intelligence tool in the cloud
DekartLarge scale geospatial analytics for Google BigQuery based on Kepler.gl
GoodDataplatform for data products and embedded analytics
Jaspersoftpowerful business intelligence suite
Jedox Palocustomisable Business Intelligence platform
JethrodataInteractive Big Data Analytics
intermix.ioPerformance Monitoring for Amazon Redshift
Metabase39,103almost 2 years agoThe simplest, fastest way to get business intelligence and analytics to everyone in your company
Microsoftbusiness intelligence software and platform
Microstrategysoftware platforms for business intelligence, mobile intelligence, and network applications
NumeracyFast, clean SQL client and business intelligence
Pentahobusiness intelligence platform
Qlikbusiness intelligence and analytics platform
RedashOpen source business intelligence platform, supporting multiple data sources and planned queries
Saiku AnalyticsOpen source analytics platform
Knowageopen source business intelligence platform. (former )
SparklineData SNAPmodern B.I platform powered by Apache Spark
Tableaubusiness intelligence platform
ZoomdataBig Data Analytics

Awesome Big Data / Data Visualization

Airpal2,754over 5 years agoWeb UI for PrestoDB
AnyChartfast, simple and flexible JavaScript (HTML5) charting library featuring pure JS API
Arbor2,664over 6 years agograph visualization library using web workers and jQuery
Banana668about 2 years agovisualize logs and time-stamped data stored in Solr. Port of Kibana
Bloomery17over 9 years agoWeb UI for Impala
BokehA powerful Python interactive visualization library that targets modern web browsers for presentation, with the goal of providing elegant, concise construction of novel graphics in the style of D3.js, but also delivering this capability with high-performance interactivity over very large or streaming datasets
C3D3-based reusable chart library
CartoDB2,761about 2 years agoopen-source or freemium hosting for geospatial databases with powerful front-end editing capabilities and a robust API
chartdresponsive, retina-compatible charts with just an img tag
Chart.jsopen source HTML5 Charts visualizations
Chartist.js73over 2 years agoanother open source HTML5 Charts visualization
CrossfilterJavaScript library for exploring large multivariate datasets in the browser. Works well with dc.js and d3.js
Cubism4,940over 3 years agoJavaScript library for time series visualization
CytoscapeJavaScript library for visualizing complex networks
DC.jsDimensional charting built to work natively with crossfilter rendered using d3.js. Excellent for connecting charts/additional metadata to hover events in D3
D3javaScript library for manipulating documents
D3.compose698almost 4 years agoCompose complex, data-driven visualizations from reusable charts and components
D3PlusA fairly robust set of reusable charts and styles for d3.js
Dash21,641almost 2 years agoAnalytical Web Apps for Python, R, Julia, and Jupyter. Built on top of plotly, no JS required
DekartLarge scale geospatial analytics for Google BigQuery based on Kepler.gl
DevExtreme React ChartHigh-performance plugin-based React chart for Bootstrap and Material Design
Echarts60,918almost 2 years agoBaidus enterprise charts
Envisionjs1,564over 6 years agodynamic HTML5 visualization
FnordMetricwrite SQL queries that return SVG charts rather than tables
Frappe ChartsGitHub-inspired simple and modern SVG charts for the web with zero dependencies
Freeboard6,455almost 3 years agopen source real-time dashboard builder for IOT and other web mashups
Gephi5,964about 2 years agoAn award-winning open-source platform for visualizing and manipulating large graphs and network connections. It's like Photoshop, but for graphs. Available for Windows and Mac OS X
Google Chartssimple charting API
Grafanagraphite dashboard frontend, editor and graph composer
Graphitescalable Realtime Graphing
Highchartssimple and flexible charting API
IPythonprovides a rich architecture for interactive computing
Kibanavisualize logs and time-stamped data
Lumifyopen source big data analysis and visualization platform
Matplotlib20,443almost 2 years agoplotting with Python
Metricsgraphic.jsa library built on top of D3 that is optimized for time-series data
NVD3chart components for d3.js
Peity4,219over 2 years agoProgressive SVG bar, line and pie charts
Plot.lyEasy-to-use web service that allows for rapid creation of complex charts, from heatmaps to histograms. Upload data to create and style charts with Plotly's online spreadsheet. Fork others' plots
Plotly.js17,161almost 2 years agoThe open source javascript graphing library that powers plotly
Recline2,209almost 2 years agosimple but powerful library for building data applications in pure Javascript and HTML
Redash26,572almost 2 years agoopen-source platform to query and visualize data
ReChartsA composable charting library built on React components
Shinya web application framework for R
Sigma.js11,339almost 2 years agoJavaScript library dedicated to graph drawing
Superset63,320almost 2 years agoa data exploration platform designed to be visual, intuitive and interactive, making it easy to slice, dice and visualize data and perform analytics at the speed of thought
Vega11,276almost 2 years agoa visualization grammar
Zeppelin411about 9 years agoa notebook-style collaborative data analysis
Zing ChartsJavaScript charting library for big data
DataSphere Studio3,100almost 2 years agoone-stop data application development management portal

Awesome Big Data / Internet of things and sensor data

Apache Edgent (Incubating)a programming model and micro-kernel style runtime that can be embedded in gateways and small footprint edge devices enabling local, real-time, analytics on the edge devices
Azure IoT HubCloud-based bi-directional monitoring and messaging hub
TempoIQCloud-based sensor analytics
2lemetryPlatform for Internet of things
PubnubData stream network
ThingWorxRapid development and connection of intelligent systems
IFTTTIf this then that
EvrythingMaking products smart
NetLytics9over 8 years agoAnalytics platform to process network data on Spark
AblyPub/sub messaging platform for IoT

Awesome Big Data / Interesting Readings

Big Data BenchmarkBenchmark of Redshift, Hive, Shark, Impala and Stiger/Tez
NoSQL ComparisonCassandra vs MongoDB vs CouchDB vs Redis vs Riak vs HBase vs Couchbase vs Neo4j vs Hypertable vs ElasticSearch vs Accumulo vs VoltDB vs Scalaris comparison
Monitoring Kafka performanceGuide to monitoring Apache Kafka, including native methods for metrics collection
Monitoring Hadoop performanceGuide to monitoring Hadoop, with an overview of Hadoop architecture, and native methods for metrics collection
Monitoring Cassandra performanceGuide to monitoring Cassandra, including native methods for metrics collection

Awesome Big Data / Interesting Papers / 2015 - 2016

2015- One Trillion Edges: Graph Processing at Facebook-Scale

Awesome Big Data / Interesting Papers / 2013 - 2014

2014- Mining of Massive Datasets
2013- Presto: Distributed Machine Learning and Graph Processing with Sparse Matrices
2013- MLbase: A Distributed Machine-learning System
2013- Shark: SQL and Rich Analytics at Scale
2013- GraphX: A Resilient Distributed Graph System on Spark
2013- HyperLogLog in Practice: Algorithmic Engineering of a State of The Art Cardinality Estimation Algorithm
2013- Scalable Progressive Analytics on Big Data in the Cloud
2013- Druid: A Real-time Analytical Data Store
2013- Online, Asynchronous Schema Change in F1
2013- F1: A Distributed SQL Database That Scales
2013- MillWheel: Fault-Tolerant Stream Processing at Internet Scale
2013- Scuba: Diving into Data at Facebook
2013- Unicorn: A System for Searching the Social Graph
2013- Scaling Memcache at Facebook

Awesome Big Data / Interesting Papers / 2011 - 2012

2012- The Unified Logging Infrastructure for Data Analytics at Twitter
2012- Blink and It’s Done: Interactive Queries on Very Large Data
2012- Fast and Interactive Analytics over Hadoop Data with Spark
2012- Shark: Fast Data Analysis Using Coarse-grained Distributed Memory
2012- Paxos Replicated State Machines as the Basis of a High-Performance Data Store
2012- Paxos Made Parallel
2012- BlinkDB: Queries with Bounded Errors and Bounded Response Times on Very Large Data
2012- Processing a trillion cells per mouse click
2012- Spanner: Google’s Globally-Distributed Database
2011- Scarlett: Coping with Skewed Popularity Content in MapReduce Clusters
2011- Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center
2011- Megastore: Providing Scalable, Highly Available Storage for Interactive Services

Awesome Big Data / Interesting Papers / 2001 - 2010

2010- Finding a needle in Haystack: Facebook’s photo storage
2010- Spark: Cluster Computing with Working Sets
2010- Pregel: A System for Large-Scale Graph Processing
2010- Large-scale Incremental Processing Using Distributed Transactions and Notifications base of Percolator and Caffeine
2010- Dremel: Interactive Analysis of Web-Scale Datasets
2010- S4: Distributed Stream Computing Platform
2009HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads
2008- Chukwa: A large-scale monitoring system
2007- Dynamo: Amazon’s Highly Available Key-value Store
2006- The Chubby lock service for loosely-coupled distributed systems
2006- Bigtable: A Distributed Storage System for Structured Data
2004- MapReduce: Simplied Data Processing on Large Clusters
2003- The Google File System

Awesome Big Data / Videos

Spark in MotionSpark in Motion teaches you how to use Spark for batch and streaming data analytics
Machine Learning, Data Science and Deep Learning with PythonLiveVideo tutorial that covers machine learning, Tensorflow, artificial intelligence, and neural networks
Data warehouse schema design - dimensional modeling and star schemaIntroduction to schema design for data warehouse using the star schema method
Elasticsearch 7 and Elastic StackLiveVideo tutorial that covers searching, analyzing, and visualizing big data on a cluster with Elasticsearch, Logstash, Beats, Kibana, and more

Awesome Big Data / Books

Data Science at Scale with Python and DaskData Science at Scale with Python and Dask teaches you how to build distributed data projects that can handle huge amounts of data
Streaming DataStreaming Data introduces the concepts and requirements of streaming and real-time data systems
Storm AppliedStorm Applied is a practical guide to using Apache Storm for the real-world tasks associated with processing and analyzing real-time data streams
Fundamentals of Stream Processing: Application Design, Systems, and AnalyticsThis comprehensive, hands-on guide combining the fundamental building blocks and emerging research in stream processing is ideal for application designers, system builders, analytic developers, as well as students and researchers in the field
Stream Data Processing: A Quality of Service PerspectivePresents a new paradigm suitable for stream and complex event processing
Unified Log ProcessingUnified Log Processing is a practical guide to implementing a unified log of event streams (Kafka or Kinesis) in your business
Kafka Streams in ActionKafka Streams in Action teaches you everything you need to know to implement stream processing on data flowing into your Kafka platform, allowing you to focus on getting more from your data without sacrificing time or effort
Big DataBig Data teaches you to build big data systems using an architecture that takes advantage of clustered hardware along with new tools designed specifically to capture and analyze web-scale data
Spark in Action& - Spark in Action teaches you the theory and skills you need to effectively handle batch and streaming data using Spark. Fully updated for Spark 2.0
Kafka in ActionKafka in Action is a fast-paced introduction to every aspect of working with Kafka you need to really reap its benefits
Fusion in ActionFusion in Action teaches you to build a full-featured data analytics pipeline, including document and data search and distributed data clustering
Reactive Data HandlingReactive Data Handling is a collection of five hand-picked chapters, selected by Manuel Bernhardt, that introduce you to building reactive applications capable of handling real-time processing with large data loads--free eBook!
Azure Data EngineeringA book about data engineering in general and the Azure platform specifically
Grokking Streaming SystemsGrokking Streaming Systems helps you unravel what streaming systems are, how they work, and whether they’re right for your business. Written to be tool-agnostic, you’ll be able to apply what you learn no matter which framework you choose
Distributed Systems for fun and profit– Theory of distributed systems. Include parts about time and ordering, replication and impossibility results
Graph-Powered Machine LearningAlessandro Negro. Combine graph theory and models to improve machine learning projects

Awesome Big Data / Books / Data Visualization

The beauty of data visualization
Designing Data Visualizations with Noah Iliinsky
Hans Rosling's 200 Countries, 200 Years, 4 Minutes
Ice Bucket Challenge Data Visualization

Other Awesome Lists

awesome-awesomeness32,173over 2 years agoOther awesome lists
awesome337,709almost 2 years agoEven more lists
list10,067almost 2 years agoAnother list?
awesome-awesome-awesome1,960almost 3 years agoWTF!
awesome-analytics3,954over 2 years agoAnalytics
awesome-public-datasets61,377almost 2 years agoPublic Datasets
awesome-graph-classification4,767over 3 years agoGraph Classification
awesome-network-embedding2,595almost 6 years agoNetwork Embedding
awesome-community-detection2,342over 2 years agoCommunity Detection
awesome-decision-tree-papers2,387over 2 years agoDecision Tree Papers
awesome-fraud-detection-papers1,639over 2 years agoFraud Detection Papers
awesome-gradient-boosting-papers1,005over 2 years agoGradient Boosting Papers
awesome-monte-carlo-tree-search-papers652over 2 years agoMonte Carlo Tree Search Papers
awesome-kafka206over 2 years agoKafka
Google Bigtable49almost 4 years ago

Backlinks from these awesome lists:

More related projects: