docker-spark-iceberg

Spark environment

A Docker-based environment for running Spark and Iceberg in a quick start scenario.

GitHub

264 stars
13 watching
139 forks
Language: Jupyter Notebook
last commit: almost 2 years ago

Related projects:

RepositoryDescriptionStars
jupyter-incubator/sparkmagicAn open source library that enables interactive development of applications using remote Spark clusters1,334
sequenceiq/docker-sparkA Docker image with Apache Spark pre-installed and configured for easy deployment on YARN clusters.765
databricks/spark-xmlA library that parses and queries XML data in Apache Spark504
databricks/spark-csvA library for parsing and querying CSV data with Apache Spark1,052
apple/batch-processing-gatewayA tool to simplify running Spark on Kubernetes181
svenkreiss/pysparklingA lightweight Python implementation of Spark's RDD and DStream interfaces for improved performance on small datasets262
yannael/kafka-sparkstreaming-cassandraAn environment for experimenting with real-time data processing using Kafka, Spark streaming, and Cassandra97
databricks/spark-corenlpWraps Stanford CoreNLP annotators as Spark DataFrame functions for natural language processing tasks422
indix/sparkplugA Spark-based package to apply data fixes using rule-based SQL conditions28
sparklyr/sparklyrAn R interface to Apache Spark for distributed data analysis and machine learning955
apache/sparkAn analytics engine designed to handle large-scale data processing and analysis40,170
databricks/tensorframesEnables manipulation of Apache Spark DataFrames using TensorFlow programs749
instaclustr/sample-kafkasparkcassandraAn introductory Scala app using Apache Spark Streaming to process data from Kafka and write summaries to Cassandra.23
ellerbrock/docker-tutorialA comprehensive guide to Docker development and deployment14
kcrandall/emr_spark_automationAutomates deployment of an AWS EMR cluster and execution of Spark jobs8