joblib-spark

Task parallelizer

Enables parallelization of machine learning tasks on a distributed Spark cluster using the joblib library.

Joblib Apache Spark Backend

GitHub

243 stars
9 watching
26 forks
Language: Python
last commit: about 2 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
shomali11/parallelizerSimplifies creating multiple worker threads to execute tasks in parallel72
dmmiller612/sparktorchA PyTorch implementation on Apache Spark for distributed deep learning model training and inference.339
clin99/cpp-taskflowA library providing a simple and expressive way to write parallel programs with complex task dependencies.8
kcrandall/emr_spark_automationAutomates deployment of an AWS EMR cluster and execution of Spark jobs8
amplab/sparknetDistributed neural network framework for Apache Spark604
apache/sparkAn analytics engine designed to handle large-scale data processing and analysis40,170
tugdualsarazin/spark-clusteringImplementations of clustering algorithms using Spark in Scala18
lensacom/sparkit-learnA Python library that integrates PySpark and scikit-learn for distributed machine learning1,154
instaclustr/sample-sparkjobservercassandraDemonstrates using Spark Jobserver to run Apache Spark analytics with Cassandra2
yaooqinn/itachiA library that brings useful functions from various modern database management systems to Apache Spark56
stevenjl/parexAn Elixir module that executes multiple processes in parallel to speed up slow computations63
janeliascicomp/nextflow-sparkProvides a reusable set of Nextflow subworkflows and processes for creating transient Apache Spark clusters on any infrastructure.14
svenkreiss/pysparklingA lightweight Python implementation of Spark's RDD and DStream interfaces for improved performance on small datasets262
microsoft/mobiusProvides a C# API for interacting with Apache Spark941
kotlin/kotlin-spark-apiProvides compatibility and extensions between Kotlin and Apache Spark for big data processing463