EMR_Spark_Automation
EMR automation tool
Automates deployment of an AWS EMR cluster and execution of Spark jobs
A repository for deploying an AWS EMR cluster and submiting spark jobs on it. Boostrapping by default does inclues pysparkling so one can easily use h2o with python and spark.
8 stars
1 watching
5 forks
Language: Python
last commit: about 9 years agoLinked from 1 awesome list
Related projects:
| Repository | Description | Stars |
|---|---|---|
| A command-line tool for launching and managing Apache Spark clusters on AWS | 637 | |
| Provides a lightweight R interface to Apache Spark for data processing | 641 | |
| An R interface to Apache Spark for distributed data analysis and machine learning | 955 | |
| Enables parallelization of machine learning tasks on a distributed Spark cluster using the joblib library. | 243 | |
| Automates payload development and deployment using Python classes to interact with Cobalt Strike and other tools | 118 | |
| Supports Ada and SPARK programming languages in Emacs org-babel for compiling, running, and formal verification of code | 8 | |
| Demonstrates using Spark Jobserver to run Apache Spark analytics with Cassandra | 2 | |
| An open source library that enables interactive development of applications using remote Spark clusters | 1,334 | |
| A set of reusable tools to simplify Spark development in Scala | 754 | |
| A set of Python libraries and tools to simplify interactions with various data sources using Apache Spark. | 61 | |
| A Ruby wrapper around Apache Spark's functionality for large-scale data processing | 227 | |
| A testing helper library for Apache Spark applications. | 437 | |
| Provides a NodeJS API to interact with the Cisco Spark platform | 16 | |
| A lightweight Python implementation of Spark's RDD and DStream interfaces for improved performance on small datasets | 262 | |
| A Helm chart repository providing infrastructure templates for setting up a fully functional Spark on Kubernetes cluster with integrated services. | 200 |