Skip to main content
buildradar
Sign in
Topic · big-data

big-data

Tracked open-source repos tagged big-data, sorted by stars.

110 repos
  • Daft@Eventual-Inc

    High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale

    5,736+4Star change over the last 7 days
  • SynapseML@microsoft

    Simple and Distributed Machine Learning Python Library porting ML algorithms for Spark

    5,245+3Star change over the last 7 days
  • calcite@apache

    Apache Calcite

    5,180+2Star change over the last 7 days
  • ignite@apache

    Apache Ignite

    5,082+2Star change over the last 7 days
  • Stream-Framework@tschellenbach

    Stream Framework is a Python library, which allows you to build news feed, activity streams and notification systems using Cassandra and/or Redis. The authors of Stream-Framework also provide a cloud service for feed technology:

    4,742-2Star change over the last 7 days
  • ⚡️A vue component support big amount data list with high render performance and efficient.

    4,505-1Star change over the last 7 days
  • img2dataset@rom1504

    Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.

    4,445+2Star change over the last 7 days
  • crate@crate

    CrateDB is a distributed and scalable SQL database for storing and analyzing massive amounts of data in near real-time, even with complex queries. It is PostgreSQL-compatible, and based on Lucene.

    4,432+3Star change over the last 7 days
  • fastjson2@alibaba

    🚄 FASTJSON2 is a Java JSON library with excellent performance.

    4,399+2Star change over the last 7 days
  • Data-Science-Roadmap@Moataz-Elmesmary

    Data Science Roadmap from A to Z

    4,354+3Star change over the last 7 days
  • GraphScope@alibaba

    🔨 🍇 💻 🚀 GraphScope: A One-Stop Large-Scale Graph Computing System from Alibaba | 一站式图计算系统

    3,557+1Star change over the last 7 days
  • Extensible SQL Lexer and Parser for Rust

    3,446+9Star change over the last 7 days
  • paimon@apache

    Apache Paimon is a lake format that enables building a Realtime Lakehouse Architecture with Flink and Spark for both streaming and batch operations.

    3,388+3Star change over the last 7 days
  • koalas@databricks

    Koalas: pandas API on Apache Spark

    3,374+1Star change over the last 7 days
  • sail@lakehq

    Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.

    3,341+4Star change over the last 7 days
  • hugegraph@apache

    A graph database that supports more than 100+ billion data, high performance and scalability (Include OLTP Engine & REST-API & Backends)

    3,167+6Star change over the last 7 days
  • CBoard@TuiQiao

    An easy to use, self-service open BI reporting and BI dashboard platform.

    3,098+1Star change over the last 7 days
  • kafka-ui@kafbat

    Open-Source Web UI for managing Apache Kafka clusters

    2,656+14Star change over the last 7 days
  • ArcticDB@man-group

    ArcticDB is a high performance, serverless DataFrame database built for the Python Data Science ecosystem.

    2,505+8Star change over the last 7 days
  • quary@quarylabs

    Open-source BI for engineers

    2,379+2Star change over the last 7 days
  • ambari@apache

    Apache Ambari simplifies provisioning, managing, and monitoring of Apache Hadoop clusters.

    2,311+0Star change over the last 7 days
  • ytsaurus@ytsaurus

    YTsaurus is a scalable and fault-tolerant open-source big data platform.

    2,203+0Star change over the last 7 days
  • fluss@apache

    Apache Fluss is a streaming storage built for real-time analytics.

    2,131+11Star change over the last 7 days
  • Apache DataFusion Ballista Distributed Query Engine

    2,127+7Star change over the last 7 days
  • drill@apache

    Apache Drill is a distributed MPP query layer for self describing data

    2,022+0Star change over the last 7 days
  • bookkeeper@apache

    Apache BookKeeper - a scalable, fault tolerant and low latency storage service optimized for append-only workloads

    2,011+1Star change over the last 7 days
  • fluid@fluid-cloudnative

    Fluid, elastic data abstraction and acceleration for BigData/AI applications in cloud. (Project under CNCF)

    1,967+0Star change over the last 7 days
  • kudu@apache

    Mirror of Apache Kudu

    1,913+0Star change over the last 7 days
  • awesome-data-analysis@PavelGrigoryevDS

    🚀 500+ curated resources for Data Analysis & Data Science: Python, SQL, Statistics, ML, AI, Visualization, Cheatsheets, Roadmaps, Interview Prep. For beginners and experts.

    1,890+8Star change over the last 7 days
  • auron@apache

    Apache Auron(Incubating) is an accelerator for big data engines, leveraging native vectorized execution to accelerate query processing.

    1,798+2Star change over the last 7 days
← Back to topics