spark
Tracked open-source repos tagged spark, sorted by stars.
- #61
High performance data store solution
★ 1,452+0Star change over the last 7 days - #62
This is the github repo for Learning Spark: Lightning-Fast Data Analytics [2nd Edition]
★ 1,402+2Star change over the last 7 days - #63
50+ DockerHub public images for Docker & Kubernetes - DevOps, CI/CD, GitHub Actions, CircleCI, Jenkins, TeamCity, Alpine, CentOS, Debian, Fedora, Ubuntu, Hadoop, Kafka, ZooKeeper, HBase, Cassandra, Solr, SolrCloud, Presto, Apache Drill, Nifi, Spark, Consul, Riak
★ 1,379+0Star change over the last 7 days - #64
Jupyter magics and kernels for working with remote Spark clusters
★ 1,365-1Star change over the last 7 days - #65
Taier is a big data development platform for submission, scheduling, operation and maintenance, and indicator information display
★ 1,284+0Star change over the last 7 days - #66
PySpark-Tutorial provides basic algorithms using PySpark
★ 1,280+0Star change over the last 7 days - #67
Apache DataFusion Comet Spark Accelerator
★ 1,262+2Star change over the last 7 days - #68
Scalable master data management, identity resolution, entity resolution, and deduplication using ML
★ 1,243+5Star change over the last 7 days - #69
GraphFrames is a package for Apache Spark which provides DataFrame-based Graphs
★ 1,202+0Star change over the last 7 days - #70
Apache Amoro(incubating) is a Lakehouse management system built on open data lake formats.
★ 1,171+0Star change over the last 7 days - #71
A Data Engineering & Machine Learning Knowledge Hub
★ 1,146+0Star change over the last 7 days - #72
Apache Celeborn is an elastic and high-performance service for shuffle and spilled data.
★ 1,062+1Star change over the last 7 days - #73
ADAM is a genomics analysis platform with specialized file formats built using Apache Avro, Apache Spark, and Apache Parquet. Apache 2 licensed.
★ 1,059+2Star change over the last 7 days - #74
Server for the ListenBrainz project, including the front-end (javascript/react) code that it serves and all of the data processing components that LB uses.
★ 1,006+4Star change over the last 7 days - #75
NVIDIA cuDF for Apache Spark plugin - accelerate Apache Spark with GPUs
★ 998+1Star change over the last 7 days - #76
A free tutorial for Apache Spark.
★ 988+0Star change over the last 7 days - #77
Sparkling Water provides H2O functionality inside Spark cluster
★ 980+1Star change over the last 7 days - #78★ 972+0Star change over the last 7 days
- #79
Apache Livy is an open source REST interface for interacting with Apache Spark from anywhere.
★ 961+1Star change over the last 7 days - #80
Open source project for data preparation for GenAI applications
★ 958+2Star change over the last 7 days - #81
An open protocol for secure data sharing
★ 957-1Star change over the last 7 days - #82★ 948+0Star change over the last 7 days
- #83
A connector for Spark that allows reading and writing to/from Redis cluster
★ 945+0Star change over the last 7 days - #84★ 895-2Star change over the last 7 days
- #85★ 889+0Star change over the last 7 days
- #86
A curated list of awesome resources related to the Ada and SPARK programming language
★ 864+2Star change over the last 7 days - #87
DoEKS is a tool to build, deploy and scale Data Platforms on Amazon EKS
★ 856-1Star change over the last 7 days - #88
This repository contains the development code for sparkMeasure, an Apache Spark performance analysis and troubleshooting library. It simplifies collecting, aggregating, and exporting Spark task/stage metrics, and is designed for practical use by developers and data engineers in interactive analysis, testing, and production monitoring workflows.
★ 827+0Star change over the last 7 days - #89
80+ DevOps & Data CLI Tools - AWS, GCP, GCF Python Cloud Functions, Log Anonymizer, Spark, Hadoop, HBase, Hive, Impala, Linux, Docker, Spark Data Converters & Validators (Avro/Parquet/JSON/CSV/INI/XML/YAML), Travis CI, AWS CloudFormation, Elasticsearch, Solr etc.
★ 824+0Star change over the last 7 days - #90
Scriptis is for interactive data analysis with script development(SQL, Pyspark, HiveQL), task submission(Spark, Hive), UDF, function, resource management and intelligent diagnosis.
★ 814+0Star change over the last 7 days