Skip to main content
buildradar
Sign in
Topic · spark

spark

Tracked open-source repos tagged spark, sorted by stars.

115 repos
  • carbondata@apache

    High performance data store solution

    1,452+0Star change over the last 7 days
  • LearningSparkV2@databricks

    This is the github repo for Learning Spark: Lightning-Fast Data Analytics [2nd Edition]

    1,402+2Star change over the last 7 days
  • Dockerfiles@HariSekhon

    50+ DockerHub public images for Docker & Kubernetes - DevOps, CI/CD, GitHub Actions, CircleCI, Jenkins, TeamCity, Alpine, CentOS, Debian, Fedora, Ubuntu, Hadoop, Kafka, ZooKeeper, HBase, Cassandra, Solr, SolrCloud, Presto, Apache Drill, Nifi, Spark, Consul, Riak

    1,379+0Star change over the last 7 days
  • sparkmagic@jupyter-incubator

    Jupyter magics and kernels for working with remote Spark clusters

    1,365-1Star change over the last 7 days
  • Taier@DTStack

    Taier is a big data development platform for submission, scheduling, operation and maintenance, and indicator information display

    1,284+0Star change over the last 7 days
  • pyspark-tutorial@mahmoudparsian

    PySpark-Tutorial provides basic algorithms using PySpark

    1,280+0Star change over the last 7 days
  • Apache DataFusion Comet Spark Accelerator

    1,262+2Star change over the last 7 days
  • zingg@zinggAI

    Scalable master data management, identity resolution, entity resolution, and deduplication using ML

    1,243+5Star change over the last 7 days
  • graphframes@graphframes

    GraphFrames is a package for Apache Spark which provides DataFrame-based Graphs

    1,202+0Star change over the last 7 days
  • amoro@apache

    Apache Amoro(incubating) is a Lakehouse management system built on open data lake formats.

    1,171+0Star change over the last 7 days
  • around-dataengineering@abhishek-ch

    A Data Engineering & Machine Learning Knowledge Hub

    1,146+0Star change over the last 7 days
  • celeborn@apache

    Apache Celeborn is an elastic and high-performance service for shuffle and spilled data.

    1,062+1Star change over the last 7 days
  • adam@bigdatagenomics

    ADAM is a genomics analysis platform with specialized file formats built using Apache Avro, Apache Spark, and Apache Parquet. Apache 2 licensed.

    1,059+2Star change over the last 7 days
  • listenbrainz-server@metabrainz

    Server for the ListenBrainz project, including the front-end (javascript/react) code that it serves and all of the data processing components that LB uses.

    1,006+4Star change over the last 7 days
  • cudf-spark@NVIDIA

    NVIDIA cuDF for Apache Spark plugin - accelerate Apache Spark with GPUs

    998+1Star change over the last 7 days
  • spark-scala-tutorial@deanwampler

    A free tutorial for Apache Spark.

    988+0Star change over the last 7 days
  • Sparkling Water provides H2O functionality inside Spark cluster

    980+1Star change over the last 7 days
  • sparklyr@sparklyr

    R interface for Apache Spark

    972+0Star change over the last 7 days
  • livy@apache

    Apache Livy is an open source REST interface for interacting with Apache Spark from anywhere.

    961+1Star change over the last 7 days
  • data-prep-kit@data-prep-kit

    Open source project for data preparation for GenAI applications

    958+2Star change over the last 7 days
  • delta-sharing@delta-io

    An open protocol for secure data sharing

    957-1Star change over the last 7 days
  • Mobius@microsoft

    C# and F# language binding and extensions to Apache Spark

    948+0Star change over the last 7 days
  • spark-redis@RedisLabs

    A connector for Spark that allows reading and writing to/from Redis cluster

    945+0Star change over the last 7 days
  • frameless@typelevel

    Expressive types for Spark.

    895-2Star change over the last 7 days
  • tispark@pingcap

    TiSpark is built for running Apache Spark on top of TiDB/TiKV

    889+0Star change over the last 7 days
  • A curated list of awesome resources related to the Ada and SPARK programming language

    864+2Star change over the last 7 days
  • data-on-eks@awslabs

    DoEKS is a tool to build, deploy and scale Data Platforms on Amazon EKS

    856-1Star change over the last 7 days
  • sparkMeasure@LucaCanali

    This repository contains the development code for sparkMeasure, an Apache Spark performance analysis and troubleshooting library. It simplifies collecting, aggregating, and exporting Spark task/stage metrics, and is designed for practical use by developers and data engineers in interactive analysis, testing, and production monitoring workflows.

    827+0Star change over the last 7 days
  • DevOps-Python-tools@HariSekhon

    80+ DevOps & Data CLI Tools - AWS, GCP, GCF Python Cloud Functions, Log Anonymizer, Spark, Hadoop, HBase, Hive, Impala, Linux, Docker, Spark Data Converters & Validators (Avro/Parquet/JSON/CSV/INI/XML/YAML), Travis CI, AWS CloudFormation, Elasticsearch, Solr etc.

    824+0Star change over the last 7 days
  • Scriptis@WeBankFinTech

    Scriptis is for interactive data analysis with script development(SQL, Pyspark, HiveQL), task submission(Spark, Hive), UDF, function, resource management and intelligent diagnosis.

    814+0Star change over the last 7 days
← Back to topics