Skip to main content
buildradar
Sign in
Topic · spark

spark

Tracked open-source repos tagged spark, sorted by stars.

115 repos
  • LakeSoul@lakesoul-io

    LakeSoul is an end-to-end, realtime cloud-native Lakehouse framework for fast data ingestion, concurrent updates, incremental analytics, multimodal data processing and vector search — powering next-generation BI and AI workloads.

    3,249+2Star change over the last 7 days
  • BigDataGuide@MoRan1607

    大数据学习,从零开始学习大数据,包含大数据学习各阶段学习视频、面试资料

    3,224+10Star change over the last 7 days
  • spark-operator@kubeflow

    Kubernetes operator for managing the lifecycle of Apache Spark applications on Kubernetes.

    3,150+1Star change over the last 7 days
  • spark-jobserver@spark-jobserver

    REST job server for Apache Spark

    2,836+0Star change over the last 7 days
  • lucebox@Luce-Org

    LLM speculative inference server for heterogeneous hardware & consumer GPUs

    2,829+38Star change over the last 7 days
  • spring-boot-quick@vector4wang

    :herb: 基于springboot的快速学习示例,整合自己遇到的开源框架,如:rabbitmq(延迟队列)、Kafka、jpa、redies、oauth2、swagger、jsp、docker、k3s、k3d、k8s、mybatis加解密插件、异常处理、日志输出、多模块开发、多环境打包、缓存cache、爬虫、jwt、GraphQL、dubbo、zookeeper和Async等等:pushpin:

    2,787-1Star change over the last 7 days
  • SZT-bigdata@geekyouth

    深圳地铁大数据客流分析系统🚇🚄🌟

    2,475+1Star change over the last 7 days
  • Spark@XZB-1248

    ✨Spark is a web-based, cross-platform and full-featured Remote Administration Tool (RAT) written in Go that allows you control all your devices anywhere. Spark是一个Go编写的,网页UI、跨平台以及多功能的远程控制和监控工具,你可以随时随地监控和控制所有设备。

    2,386+5Star change over the last 7 days
  • splink@moj-analytical-services

    Fast, accurate and scalable probabilistic data linkage with support for multiple SQL backends

    2,375+11Star change over the last 7 days
  • kyuubi@apache

    Apache Kyuubi is a distributed and multi-tenant gateway to provide serverless SQL on data warehouses and lakehouses.

    2,364+2Star change over the last 7 days
  • TransmogrifAI@salesforce

    TransmogrifAI (pronounced trăns-mŏgˈrə-fī) is an AutoML library for building modular, reusable, strongly typed machine learning workflows on Apache Spark with minimal hand-tuning

    2,277+0Star change over the last 7 days
  • ytsaurus@ytsaurus

    YTsaurus is a scalable and fault-tolerant open-source big data platform.

    2,203+3Star change over the last 7 days
  • Tools for handling firmwares of DJI products, with focus on quadcopters.

    2,193+7Star change over the last 7 days
  • fugue@fugue-project

    A unified interface for distributed computing. Fugue executes SQL, Python, Pandas, and Polars code on Spark, Dask and Ray without any rewrites.

    2,170+1Star change over the last 7 days
  • zio-quill@zio

    Compile-time Language Integrated Queries for Scala

    2,166+0Star change over the last 7 days
  • spark@dotnet

    .NET for Apache® Spark™ makes Apache Spark™ easily accessible to .NET developers.

    2,096-2Star change over the last 7 days
  • gatk@broadinstitute

    Official code repository for GATK versions 4 and up

    1,994+2Star change over the last 7 days
  • Apache Spark to Apache Cassandra connector

    1,955+1Star change over the last 7 days
  • bigdata-growth@collabH

    大数据知识仓库涉及到数据仓库建模、实时计算、大数据、数据中台、系统设计、Java、算法等。

    1,811+5Star change over the last 7 days
  • auron@apache

    Apache Auron(Incubating) is an accelerator for big data engines, leveraging native vectorized execution to accelerate query processing.

    1,798+3Star change over the last 7 days
  • Tutorial@zhonghuasheng

    后端 (Java Golang)全栈知识架构体系总结

    1,768+3Star change over the last 7 days
  • .github@apachecn

    ApacheCN 开源组织:公告、介绍、成员、活动、交流方式

    1,754+0Star change over the last 7 days
  • More than 2000+ Data engineer interview questions.

    1,721+7Star change over the last 7 days
  • elassandra@strapdata

    Elassandra = Elasticsearch + Apache Cassandra

    1,714+0Star change over the last 7 days
  • spark-py-notebooks@jadianes

    Apache Spark & Python (pySpark) tutorials for Big Data Analysis and Machine Learning as IPython / Jupyter notebooks

    1,660+1Star change over the last 7 days
  • almond@almond-sh

    A Scala kernel for Jupyter

    1,625+0Star change over the last 7 days
  • apache-spark-internals@japila-books

    The Internals of Apache Spark

    1,549+0Star change over the last 7 days
  • mleap@combust

    MLeap: Deploy ML Pipelines to Production

    1,544+1Star change over the last 7 days
  • optimus@hi-primus

    :truck: Agile Data Preparation Workflows made easy with Pandas, Dask, cuDF, Dask-cuDF, Vaex and PySpark

    1,536+0Star change over the last 7 days
  • nessie@projectnessie

    Nessie: Transactional Catalog for Data Lakes with Git-like semantics

    1,498+1Star change over the last 7 days
← Back to topics