Skip to main content
buildradar
Sign in
Topic · bigdata

bigdata

Tracked open-source repos tagged bigdata, sorted by stars.

Repos
48
Total stars
252,531
Avg. stars
5,261
Share
0.01%

Topics that frequently appear alongside bigdata on the same repo.

Recent risers

Repos created in the last 90 days, tagged bigdata.

No new repos tagged with this topic in the last 90 days.

  • data-engineer-handbook@DataExpert-io

    This is a repo with links to everything you'd ever want to learn about data engineering

    43,905+88Star change over the last 7 days
  • rustfs@rustfs

    🚀2.3x faster than MinIO for 4KB object payloads. RustFS is an open-source, S3-compatible high-performance object storage system supporting migration and coexistence with other S3-compatible platforms such as MinIO and Ceph.

    31,528+238Star change over the last 7 days
  • TDengine@taosdata

    High-performance, scalable time-series database designed for Industrial IoT (IIoT) scenarios

    25,093+15Star change over the last 7 days
  • Empowering Data Intelligence with Distributed SQL for Sharding, Scalability, and Security Across All Databases.

    20,791+10Star change over the last 7 days
  • BigData-Notes@heibaiying

    大数据入门指南 :star:

    16,958+6Star change over the last 7 days
  • A curated list of awesome big data frameworks, ressources and other awesomeness.

    14,601+63Star change over the last 7 days
  • juicefs@juicedata

    JuiceFS is a distributed POSIX file system built on top of Redis and S3.

    14,371+19Star change over the last 7 days
  • databend@databendlabs

    Data Agent Ready Warehouse : One for Analytics, Search, AI, Python Sandbox. — rebuilt from scratch. Unified architecture on your S3.

    9,426+3Star change over the last 7 days
  • vaex@vaexio

    Out-of-Core hybrid Apache Arrow/NumPy DataFrame for Python, ML, visualization and exploration of big tabular data at a billion rows per second 🚀

    8,509+1Star change over the last 7 days
  • hudi@apache

    Upserts, Deletes And Incremental Processing on Big Data.

    6,229+12Star change over the last 7 days
  • volcano@volcano-sh

    A Cloud Native Batch System (Project under CNCF)

    5,915+30Star change over the last 7 days
  • BigDataView@iGaoWei

    100+套大数据可视化炫酷大屏Html5模板;包含行业:社区、物业、政务、交通、金融银行等,全网最新、最多,最全、最酷、最炫大数据可视化模板。陆续更新中

    5,306+21Star change over the last 7 days
  • chunjun@DTStack

    A data integration framework

    4,101+1Star change over the last 7 days
  • 🔨 用 JSON 来生成结构化的 SQL 语句,基于 Vue3 + TypeScript + Vite + Ant Design + MonacoEditor 实现,项目简单(重逻辑轻页面)、适合练手~

    3,424-3Star change over the last 7 days
  • avro@apache

    Apache Avro is a data serialization system.

    3,302+3Star change over the last 7 days
  • BigDataGuide@MoRan1607

    大数据学习,从零开始学习大数据,包含大数据学习各阶段学习视频、面试资料

    3,223+9Star change over the last 7 days
  • griddb@griddb

    GridDB is a next-generation open source database that makes time series IoT and big data fast,and easy.

    2,475+0Star change over the last 7 days
  • Aix-DB@apconw

    Aix-DB 基于 LangChain/LangGraph 框架,结合 MCP Skills 多智能体协作架构,实现自然语言到数据洞察的端到端转换。

    2,240+10Star change over the last 7 days
  • spark@dotnet

    .NET for Apache® Spark™ makes Apache Spark™ easily accessible to .NET developers.

    2,096+0Star change over the last 7 days
  • flinkStreamSQL@DTStack

    基于开源的flink,对其实时sql进行扩展;主要实现了流与维表的join,支持原生flink SQL所有的语法

    2,051+0Star change over the last 7 days
  • byzer-lang@byzer-org

    Byzer (former MLSQL): A low-code open-source programming language for data pipeline, analytics and AI.

    1,835-1Star change over the last 7 days
  • bigdata-growth@collabH

    大数据知识仓库涉及到数据仓库建模、实时计算、大数据、数据中台、系统设计、Java、算法等。

    1,810+3Star change over the last 7 days
  • genie@Netflix

    Distributed Big Data Orchestration Service

    1,767+0Star change over the last 7 days
  • AutoCrawler@YoongiKim

    Google, Naver multiprocess image web crawler (Selenium)

    1,692-1Star change over the last 7 days
  • spark-py-notebooks@jadianes

    Apache Spark & Python (pySpark) tutorials for Big Data Analysis and Machine Learning as IPython / Jupyter notebooks

    1,660+1Star change over the last 7 days
  • optimus@hi-primus

    :truck: Agile Data Preparation Workflows made easy with Pandas, Dask, cuDF, Dask-cuDF, Vaex and PySpark

    1,536+0Star change over the last 7 days
  • odd-platform@opendatadiscovery

    First open-source data discovery and observability platform. We make a life for data practitioners easy so you can focus on your business.

    1,427+2Star change over the last 7 days
  • celeborn@apache

    Apache Celeborn is an elastic and high-performance service for shuffle and spilled data.

    1,061-1Star change over the last 7 days
  • cds@zeromicro

    Data syncing in golang for ClickHouse.

    979+0Star change over the last 7 days
  • livy@apache

    Apache Livy is an open source REST interface for interacting with Apache Spark from anywhere.

    960+0Star change over the last 7 days
← Back to topics