Skip to main content
buildradar
Sign in
Topic · hadoop

hadoop

Tracked open-source repos tagged hadoop, sorted by stars.

Repos
40
Total stars
230,812
Avg. stars
5,770
Share
0.01%

Topics that frequently appear alongside hadoop on the same repo.

Recent risers

Repos created in the last 90 days, tagged hadoop.

No new repos tagged with this topic in the last 90 days.

  • Data science Python notebooks: Deep learning (TensorFlow, Theano, Caffe, Keras), scikit-learn, Kaggle, big data (Spark, Hadoop MapReduce, HDFS), matplotlib, pandas, NumPy, SciPy, Python essentials, AWS, and various command lines.

    29,339+5Star change over the last 7 days
  • luigi@spotify

    Luigi is a Python module that helps you build complex pipelines of batch jobs. It handles dependency resolution, workflow management, visualization etc. It also comes with Hadoop support built in.

    18,767-1Star change over the last 7 days
  • APIJSON@Tencent

    🏆 Real-Time no-code, powerful and secure ORM 🚀 providing APIs and Docs without coding by Backend, and Frontend(Client) can customize response JSONs 🏆 实时 零代码、全功能、强安全 ORM 库 🚀 后端接口和文档零代码,前端(客户端) 定制返回 JSON 的数据和结构

    18,406+4Star change over the last 7 days
  • BigData-Notes@heibaiying

    大数据入门指南 :star:

    16,962+4Star change over the last 7 days
  • presto@prestodb

    The official home of the Presto distributed SQL query engine for big data

    16,730+5Star change over the last 7 days
  • hadoop@apache

    Apache Hadoop

    15,647+8Star change over the last 7 days
  • deeplearning4j@deeplearning4j

    Suite of tools for deploying and training deep learning models using the JVM. Highlights include model import for keras, tensorflow, and onnx/pytorch, a modular and tiny c++ library for running math code and a java based math library on top of the core c++ library. Also includes samediff: a pytorch/tensorflow like library for running deep learn...

    14,248+0Star change over the last 7 days
  • trino@trinodb

    Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)

    13,200+6Star change over the last 7 days
  • DevOps-Bash-tools@HariSekhon

    1200+ DevOps Bash Scripts - AWS, GCP, Kubernetes, Docker, CI/CD, APIs, SQL, PostgreSQL, MySQL, Hive, Impala, Kafka, Hadoop, Jenkins, GitHub, GitLab, BitBucket, Azure DevOps, TeamCity, Spotify, MP3, LDAP, Code/Build Linting, pkg mgmt for Linux, Mac, Python, Perl, Ruby, NodeJS, Golang, Advanced dotfiles: .bashrc, .vimrc, .gitconfig, .screenrc, tmux..

    8,393+1Star change over the last 7 days
  • school-of-sre@linkedin

    At LinkedIn, we are using this curriculum for onboarding our entry-level talents into the SRE role.

    8,141+4Star change over the last 7 days
  • h2o-3@h2oai

    H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc.

    7,499+3Star change over the last 7 days
  • alluxio@Alluxio

    Alluxio, data orchestration for analytics and machine learning in the cloud

    7,234+3Star change over the last 7 days
  • hive@apache

    Apache Hive

    6,016+1Star change over the last 7 days
  • calcite@apache

    Apache Calcite

    5,180+2Star change over the last 7 days
  • ignite@apache

    Apache Ignite

    5,082+2Star change over the last 7 days
  • nutch@apache

    Apache Nutch is an extensible and scalable web crawler

    3,283+4Star change over the last 7 days
  • DataSphereStudio@WeBankFinTech

    DataSphereStudio is a one stop data application development& management portal, covering scenarios including data exchange, desensitization/cleansing, analysis/mining, quality measurement, visualization, and task scheduling.

    3,262-1Star change over the last 7 days
  • BigDataGuide@MoRan1607

    大数据学习,从零开始学习大数据,包含大数据学习各阶段学习视频、面试资料

    3,224+5Star change over the last 7 days
  • SZT-bigdata@geekyouth

    深圳地铁大数据客流分析系统🚇🚄🌟

    2,475+1Star change over the last 7 days
  • kyuubi@apache

    Apache Kyuubi is a distributed and multi-tenant gateway to provide serverless SQL on data warehouses and lakehouses.

    2,364+1Star change over the last 7 days
  • docker-hadoop@big-data-europe

    Apache Hadoop docker image

    2,325-1Star change over the last 7 days
  • winutils@cdarlint

    winutils.exe hadoop.dll and hdfs.dll binaries for hadoop windows

    2,291+4Star change over the last 7 days
  • drill@apache

    Apache Drill is a distributed MPP query layer for self describing data

    2,022+0Star change over the last 7 days
  • bigdata-growth@collabH

    大数据知识仓库涉及到数据仓库建模、实时计算、大数据、数据中台、系统设计、Java、算法等。

    1,811+3Star change over the last 7 days
  • hbox@Qihoo360

    AI on Hadoop

    1,728+0Star change over the last 7 days
  • More than 2000+ Data engineer interview questions.

    1,721+5Star change over the last 7 days
  • carbondata@apache

    High performance data store solution

    1,452+0Star change over the last 7 days
  • Addax@wgzhao

    Actively maintained fork of Alibaba DataX — a fast, versatile ETL tool for RDBMS/NoSQL data transfer

    1,430+3Star change over the last 7 days
  • Dockerfiles@HariSekhon

    50+ DockerHub public images for Docker & Kubernetes - DevOps, CI/CD, GitHub Actions, CircleCI, Jenkins, TeamCity, Alpine, CentOS, Debian, Fedora, Ubuntu, Hadoop, Kafka, ZooKeeper, HBase, Cassandra, Solr, SolrCloud, Presto, Apache Drill, Nifi, Spark, Consul, Riak

    1,379+0Star change over the last 7 days
  • Taier@DTStack

    Taier is a big data development platform for submission, scheduling, operation and maintenance, and indicator information display

    1,284+0Star change over the last 7 days
← Back to topics