big-data
Tracked open-source repos tagged big-data, sorted by stars.
- #31
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
★ 5,736+4Star change over the last 7 days - #32
Simple and Distributed Machine Learning Python Library porting ML algorithms for Spark
★ 5,245+3Star change over the last 7 days - #33★ 5,180+2Star change over the last 7 days
- #34★ 5,082+2Star change over the last 7 days
- #35
Stream Framework is a Python library, which allows you to build news feed, activity streams and notification systems using Cassandra and/or Redis. The authors of Stream-Framework also provide a cloud service for feed technology:
★ 4,742-2Star change over the last 7 days - #36
⚡️A vue component support big amount data list with high render performance and efficient.
★ 4,505-1Star change over the last 7 days - #37
Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
★ 4,445+2Star change over the last 7 days - #38
CrateDB is a distributed and scalable SQL database for storing and analyzing massive amounts of data in near real-time, even with complex queries. It is PostgreSQL-compatible, and based on Lucene.
★ 4,432+3Star change over the last 7 days - #39★ 4,399+2Star change over the last 7 days
- #40
Data Science Roadmap from A to Z
★ 4,354+3Star change over the last 7 days - #41
🔨 🍇 💻 🚀 GraphScope: A One-Stop Large-Scale Graph Computing System from Alibaba | 一站式图计算系统
★ 3,557+1Star change over the last 7 days - #42
Extensible SQL Lexer and Parser for Rust
★ 3,446+9Star change over the last 7 days - #43
Apache Paimon is a lake format that enables building a Realtime Lakehouse Architecture with Flink and Spark for both streaming and batch operations.
★ 3,388+3Star change over the last 7 days - #44★ 3,374+1Star change over the last 7 days
- #45
Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.
★ 3,341+4Star change over the last 7 days - #46
A graph database that supports more than 100+ billion data, high performance and scalability (Include OLTP Engine & REST-API & Backends)
★ 3,167+6Star change over the last 7 days - #47★ 3,098+1Star change over the last 7 days
- #48★ 2,656+14Star change over the last 7 days
- #49
ArcticDB is a high performance, serverless DataFrame database built for the Python Data Science ecosystem.
★ 2,505+8Star change over the last 7 days - #50★ 2,379+2Star change over the last 7 days
- #51
Apache Ambari simplifies provisioning, managing, and monitoring of Apache Hadoop clusters.
★ 2,311+0Star change over the last 7 days - #52★ 2,203+0Star change over the last 7 days
- #53★ 2,131+11Star change over the last 7 days
- #54
Apache DataFusion Ballista Distributed Query Engine
★ 2,127+7Star change over the last 7 days - #55★ 2,022+0Star change over the last 7 days
- #56
Apache BookKeeper - a scalable, fault tolerant and low latency storage service optimized for append-only workloads
★ 2,011+1Star change over the last 7 days - #57
Fluid, elastic data abstraction and acceleration for BigData/AI applications in cloud. (Project under CNCF)
★ 1,967+0Star change over the last 7 days - #58★ 1,913+0Star change over the last 7 days
- #59
🚀 500+ curated resources for Data Analysis & Data Science: Python, SQL, Statistics, ML, AI, Visualization, Cheatsheets, Roadmaps, Interview Prep. For beginners and experts.
★ 1,890+8Star change over the last 7 days - #60
Apache Auron(Incubating) is an accelerator for big data engines, leveraging native vectorized execution to accelerate query processing.
★ 1,798+2Star change over the last 7 days