spark
Tracked open-source repos tagged spark, sorted by stars.
- #31
LakeSoul is an end-to-end, realtime cloud-native Lakehouse framework for fast data ingestion, concurrent updates, incremental analytics, multimodal data processing and vector search — powering next-generation BI and AI workloads.
★ 3,249+2Star change over the last 7 days - #32
大数据学习,从零开始学习大数据,包含大数据学习各阶段学习视频、面试资料
★ 3,224+10Star change over the last 7 days - #33
Kubernetes operator for managing the lifecycle of Apache Spark applications on Kubernetes.
★ 3,150+1Star change over the last 7 days - #34
REST job server for Apache Spark
★ 2,836+0Star change over the last 7 days - #35★ 2,829+38Star change over the last 7 days
- #36
:herb: 基于springboot的快速学习示例,整合自己遇到的开源框架,如:rabbitmq(延迟队列)、Kafka、jpa、redies、oauth2、swagger、jsp、docker、k3s、k3d、k8s、mybatis加解密插件、异常处理、日志输出、多模块开发、多环境打包、缓存cache、爬虫、jwt、GraphQL、dubbo、zookeeper和Async等等:pushpin:
★ 2,787-1Star change over the last 7 days - #37
深圳地铁大数据客流分析系统🚇🚄🌟
★ 2,475+1Star change over the last 7 days - #38
✨Spark is a web-based, cross-platform and full-featured Remote Administration Tool (RAT) written in Go that allows you control all your devices anywhere. Spark是一个Go编写的,网页UI、跨平台以及多功能的远程控制和监控工具,你可以随时随地监控和控制所有设备。
★ 2,386+5Star change over the last 7 days - #39
Fast, accurate and scalable probabilistic data linkage with support for multiple SQL backends
★ 2,375+11Star change over the last 7 days - #40
Apache Kyuubi is a distributed and multi-tenant gateway to provide serverless SQL on data warehouses and lakehouses.
★ 2,364+2Star change over the last 7 days - #41
TransmogrifAI (pronounced trăns-mŏgˈrə-fī) is an AutoML library for building modular, reusable, strongly typed machine learning workflows on Apache Spark with minimal hand-tuning
★ 2,277+0Star change over the last 7 days - #42★ 2,203+3Star change over the last 7 days
- #43
Tools for handling firmwares of DJI products, with focus on quadcopters.
★ 2,193+7Star change over the last 7 days - #44
A unified interface for distributed computing. Fugue executes SQL, Python, Pandas, and Polars code on Spark, Dask and Ray without any rewrites.
★ 2,170+1Star change over the last 7 days - #45★ 2,166+0Star change over the last 7 days
- #46★ 2,096-2Star change over the last 7 days
- #47★ 1,994+2Star change over the last 7 days
- #48
Apache Spark to Apache Cassandra connector
★ 1,955+1Star change over the last 7 days - #49
大数据知识仓库涉及到数据仓库建模、实时计算、大数据、数据中台、系统设计、Java、算法等。
★ 1,811+5Star change over the last 7 days - #50
Apache Auron(Incubating) is an accelerator for big data engines, leveraging native vectorized execution to accelerate query processing.
★ 1,798+3Star change over the last 7 days - #51★ 1,768+3Star change over the last 7 days
- #52★ 1,754+0Star change over the last 7 days
- #53
More than 2000+ Data engineer interview questions.
★ 1,721+7Star change over the last 7 days - #54
Elassandra = Elasticsearch + Apache Cassandra
★ 1,714+0Star change over the last 7 days - #55
Apache Spark & Python (pySpark) tutorials for Big Data Analysis and Machine Learning as IPython / Jupyter notebooks
★ 1,660+1Star change over the last 7 days - #56★ 1,625+0Star change over the last 7 days
- #57
The Internals of Apache Spark
★ 1,549+0Star change over the last 7 days - #58★ 1,544+1Star change over the last 7 days
- #59
:truck: Agile Data Preparation Workflows made easy with Pandas, Dask, cuDF, Dask-cuDF, Vaex and PySpark
★ 1,536+0Star change over the last 7 days - #60★ 1,498+1Star change over the last 7 days