pyspark
Tracked open-source repos tagged pyspark, sorted by stars.
Related topics
Topics that frequently appear alongside pyspark on the same repo.
Recent risers
Repos created in the last 90 days, tagged pyspark.
No new repos tagged with this topic in the last 90 days.
- #1★ 6,650+1Star change over the last 7 days
- #2
Simple and Distributed Machine Learning Python Library porting ML algorithms for Spark
★ 5,245+3Star change over the last 7 days - #3★ 4,160+1Star change over the last 7 days
- #4
Apache Linkis builds a computation middleware layer to facilitate connection, governance and orchestration between the upper applications and the underlying data engines.
★ 3,409+2Star change over the last 7 days - #5
Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.
★ 3,341+4Star change over the last 7 days - #6
A curated list of awesome Apache Spark packages and resources.
★ 1,893+0Star change over the last 7 days - #7
Petastorm library enables single machine or distributed training and evaluation of deep learning models from datasets in Apache Parquet format. It supports ML frameworks such as Tensorflow, Pytorch, and PySpark and can be used from pure Python code.
★ 1,892+0Star change over the last 7 days - #8★ 1,714+4Star change over the last 7 days
- #9
Apache Spark & Python (pySpark) tutorials for Big Data Analysis and Machine Learning as IPython / Jupyter notebooks
★ 1,660+0Star change over the last 7 days - #10
:truck: Agile Data Preparation Workflows made easy with Pandas, Dask, cuDF, Dask-cuDF, Vaex and PySpark
★ 1,536+0Star change over the last 7 days - #11
Jupyter magics and kernels for working with remote Spark clusters
★ 1,365-1Star change over the last 7 days - #12★ 1,304+1Star change over the last 7 days
- #13
PySpark-Tutorial provides basic algorithms using PySpark
★ 1,280+0Star change over the last 7 days - #14
GraphFrames is a package for Apache Spark which provides DataFrame-based Graphs
★ 1,202+0Star change over the last 7 days - #15
MapReduce, Spark, Java, and Scala for Data Algorithms Book
★ 1,080-2Star change over the last 7 days - #16
Sparkling Water provides H2O functionality inside Spark cluster
★ 980+1Star change over the last 7 days - #17
80+ DevOps & Data CLI Tools - AWS, GCP, GCF Python Cloud Functions, Log Anonymizer, Spark, Hadoop, HBase, Hive, Impala, Linux, Docker, Spark Data Converters & Validators (Avro/Parquet/JSON/CSV/INI/XML/YAML), Travis CI, AWS CloudFormation, Elasticsearch, Solr etc.
★ 824+0Star change over the last 7 days - #18
Python framework for building efficient data pipelines. It promotes modularity and collaboration, enabling the creation of complex pipelines from simple, reusable components.
★ 818+0Star change over the last 7 days - #19
Scriptis is for interactive data analysis with script development(SQL, Pyspark, HiveQL), task submission(Spark, Hive), UDF, function, resource management and intelligent diagnosis.
★ 814+0Star change over the last 7 days - #20★ 772-1Star change over the last 7 days
- #21★ 689+0Star change over the last 7 days
- #22★ 658+0Star change over the last 7 days
- #23
Learn Apache Spark in Scala, Python (PySpark) and R (SparkR) by building your own cluster with a JupyterLab interface on Docker. :zap:
★ 512+0Star change over the last 7 days