big-data
Tracked open-source repos tagged big-data, sorted by stars.
- #61★ 1,766-1Star change over the last 7 days
- #62
Open source security data lake for threat hunting, detection & response, and cybersecurity analytics at petabyte scale on AWS
★ 1,694+0Star change over the last 7 days - #63
Apache Spark & Python (pySpark) tutorials for Big Data Analysis and Machine Learning as IPython / Jupyter notebooks
★ 1,660+0Star change over the last 7 days - #64★ 1,618+3Star change over the last 7 days
- #65
:bar_chart: :clipboard: Dashboards using YAML or JSON files
★ 1,582+1Star change over the last 7 days - #66
Dremio - the missing link in modern data
★ 1,492+1Star change over the last 7 days - #67
High performance data store solution
★ 1,452+0Star change over the last 7 days - #68★ 1,384+0Star change over the last 7 days
- #69
One advanced and mature open-source MPP (Massively Parallel Processing) database. Open source alternative to Greenplum Database.
★ 1,382+5Star change over the last 7 days - #70
Extension for Scikit-learn is a seamless way to speed up your Scikit-learn application
★ 1,356+0Star change over the last 7 days - #71
PySpark-Tutorial provides basic algorithms using PySpark
★ 1,280+0Star change over the last 7 days - #72
Scalable, reliable, distributed storage system optimized for data analytics and object store workloads.
★ 1,270+1Star change over the last 7 days - #73
Simple Windows desktop application for viewing & querying Apache Parquet files
★ 1,231+2Star change over the last 7 days - #74
GraphFrames is a package for Apache Spark which provides DataFrame-based Graphs
★ 1,202+0Star change over the last 7 days - #75
Apache Amoro(incubating) is a Lakehouse management system built on open data lake formats.
★ 1,171+0Star change over the last 7 days - #76★ 1,165+2Star change over the last 7 days
- #77
ClickBench: a Benchmark For Analytical Databases
★ 1,094+2Star change over the last 7 days - #78
📙 Awesome Data Catalogs and Observability Platforms.
★ 1,068+5Star change over the last 7 days - #79
ADAM is a genomics analysis platform with specialized file formats built using Apache Avro, Apache Spark, and Apache Parquet. Apache 2 licensed.
★ 1,059+2Star change over the last 7 days - #80
A @ClickHouse fork that supports high-performance vector search and full-text search.
★ 1,040+0Star change over the last 7 days - #81
Apache Flink Kubernetes Operator
★ 1,031+3Star change over the last 7 days - #82
Career Resources for Data Science, Machine Learning, Big Data and Business Analytics Career Repository
★ 1,022+2Star change over the last 7 days - #83
Server for the ListenBrainz project, including the front-end (javascript/react) code that it serves and all of the data processing components that LB uses.
★ 1,006+4Star change over the last 7 days - #84
NVIDIA cuDF for Apache Spark plugin - accelerate Apache Spark with GPUs
★ 998+1Star change over the last 7 days - #85
Sparkling Water provides H2O functionality inside Spark cluster
★ 980+1Star change over the last 7 days - #86
An open protocol for secure data sharing
★ 957-1Star change over the last 7 days - #87
⚡ Single-pass algorithms for statistics
★ 897+0Star change over the last 7 days - #88
Un repositorio más con conceptos básicos, desafíos técnicos y recursos sobre ingeniería de datos en español 🧙✨
★ 896+1Star change over the last 7 days - #89
🏅State-of-the-art learned data structure that enables fast lookup, predecessor, range searches and updates in arrays of billions of items using orders of magnitude less space than traditional indexes
★ 873+1Star change over the last 7 days - #90
AI projects
★ 868+1Star change over the last 7 days