Skip to main content
buildradar
Sign in
Topic · data-engineering

data-engineering

Tracked open-source repos tagged data-engineering, sorted by stars.

97 repos
  • pyjanitor@pyjanitor-devs

    Clean APIs for data cleaning. Python implementation of R package Janitor

    1,498+0Star change over the last 7 days
  • data-science-on-gcp@GoogleCloudPlatform

    Source code accompanying book: Data Science on the Google Cloud Platform, Valliappa Lakshmanan, O'Reilly 2017

    1,429+1Star change over the last 7 days
  • odd-platform@opendatadiscovery

    First open-source data discovery and observability platform. We make a life for data practitioners easy so you can focus on your business.

    1,427+1Star change over the last 7 days
  • mlops-python-package@fmind

    A comprehensive Python package template to kickstart and standardize your MLOps initiatives and data pipelines.

    1,416+0Star change over the last 7 days
  • A curated collection of 300+ engineering blog articles from top tech companies. Learn how the best engineering teams solve real-world problems at scale.

    1,388+8Star change over the last 7 days
  • sql-ultimate-course@DataWithBaraa

    The most comprehensive SQL guide from a real-world expert! Learn everything from basics to advanced queries, optimizations, and real-world SQL

    1,384+7Star change over the last 7 days
  • quilt@quiltdata

    Quilt is a Scientific Data Management Platform on AWS that helps teams and AI find, trust, and reuse data through deeply versioned, context-rich data packages.

    1,370+0Star change over the last 7 days
  • duckle@slothflowlabs

    Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality, reverse ETL, lineage, MCP for AI agents. No vendor cloud, no per-row billing.

    1,262+15Star change over the last 7 days
  • A curated, but incomplete, list of data-centric AI resources.

    1,159+1Star change over the last 7 days
  • around-dataengineering@abhishek-ch

    A Data Engineering & Machine Learning Knowledge Hub

    1,146+0Star change over the last 7 days
  • Home of the Open Data Contract Standard (ODCS).

    1,118+18Star change over the last 7 days
  • awesome-data-catalogs@opendatadiscovery

    📙 Awesome Data Catalogs and Observability Platforms.

    1,068+5Star change over the last 7 days
  • datacontract-cli@datacontract

    Enforce Data Contracts

    1,060+8Star change over the last 7 days
  • flow@estuary

    🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊

    972+7Star change over the last 7 days
  • FlyFish@CloudWise-OpenSource

    FlyFish is a data visualization coding platform. We can create a data model quickly in a simple way, and quickly generate a set of data visualization solutions by dragging.

    961+0Star change over the last 7 days
  • A comprehensive guide to building a modern data warehouse with SQL Server, including ETL processes, data modeling, and analytics.

    934+14Star change over the last 7 days
  • egeria@odpi

    Egeria core

    921+0Star change over the last 7 days
  • Un repositorio más con conceptos básicos, desafíos técnicos y recursos sobre ingeniería de datos en español 🧙✨

    896+1Star change over the last 7 days
  • bacalhau@bacalhau-project

    Community-driven, simple, yet powerful framework for fast, cost-effective distributed Compute over Data.

    871+1Star change over the last 7 days
  • NeumAI@NeumTry

    Neum AI is a best-in-class framework to manage the creation and synchronization of vector embeddings at large scale.

    867+0Star change over the last 7 days
  • Supplementary Materials for the The Complete dbt (Data Build Tool) Bootcamp Udemy course

    827+0Star change over the last 7 days
  • practical-data-engineering@ssp-data

    Practical Data Engineering: A Hands-On Real-Estate Project Guide

    823+2Star change over the last 7 days
  • koheesio@Nike-Inc

    Python framework for building efficient data pipelines. It promotes modularity and collaboration, enabling the creation of complex pipelines from simple, reusable components.

    818+0Star change over the last 7 days
  • altimate-code@AltimateAI

    Open-source agentic data engineering harness for dbt, SQL, and cloud warehouses. 100+ tools, 10 warehouses, AI-powered.

    808+3Star change over the last 7 days
  • datavines@datavane

    Know your data better!Datavines is Next-gen Data Observability Platform, support metadata manage and data quality.

    761+1Star change over the last 7 days
  • bigfunctions@unytics

    Supercharge BigQuery with BigFunctions

    759+0Star change over the last 7 days
  • Failed-ML@kennethleungty

    Compilation of high-profile real-world examples of failed machine learning projects

    751+0Star change over the last 7 days
  • vectorflow@dgarnitz

    VectorFlow is a high volume vector embedding pipeline that ingests raw data, transforms it into vectors and writes it to a vector DB of your choice.

    703+0Star change over the last 7 days
  • DataFlow-Engine@risesoft-y9

    数据流引擎是一款面向数据集成、数据同步、数据交换、数据共享、任务配置、任务调度的底层数据驱动引擎。数据流引擎采用管执分离、多流层、插件库等体系应对大规模数据任务、数据高频上报、数据高频采集、异构数据兼容的实际数据问题。

    696+0Star change over the last 7 days
  • synmetrix@synmetrix

    Synmetrix – production-ready open source semantic layer on Cube

    626+2Star change over the last 7 days
← Back to topics