pysparkling
Dataset processor
A lightweight Python implementation of Spark's RDD and DStream interfaces for improved performance on small datasets
A pure Python implementation of Apache Spark's RDD and DStream interfaces.
262 stars
9 watching
45 forks
Language: Python
last commit: about 2 years agoLinked from 1 awesome list
apache-sparkdata-processingdata-sciencepython
Related projects:
| Repository | Description | Stars |
|---|---|---|
| An analytics engine designed to handle large-scale data processing and analysis | 40,170 | |
| A set of Python libraries and tools to simplify interactions with various data sources using Apache Spark. | 61 | |
| A library that enables integration between Apache Spark and Apache Cassandra for fast data processing and analysis. | 1,944 | |
| Enables distributed data analysis using PySpark and Pandas APIs | 362 | |
| An introductory Scala app using Apache Spark Streaming to process data from Kafka and write summaries to Cassandra. | 23 | |
| A Python package for managing and analyzing demand-side grid data, models, and queries using Apache Spark | 26 | |
| A data processing library built on top of Apache Spark to handle temporal web data | 11 | |
| An R interface to Apache Spark for distributed data analysis and machine learning | 955 | |
| Pyspark helper functions to maximize developer productivity | 651 | |
| A PyTorch implementation on Apache Spark for distributed deep learning model training and inference. | 339 | |
| Provides an interface to the Cisco Spark REST API | 30 | |
| A comprehensive reference guide to working with PySpark SQL | 458 | |
| A Python-based data refinement tool that extends Google Refine with Linked Open Data features | 14 | |
| A Python library providing a clean and expressive API for data cleaning by chaining multiple operations together in a logical order. | 1,371 | |
| Fast JSON parsing for Python, using SIMD instructions when available | 648 |