data-engineering
Tracked open-source repos tagged data-engineering, sorted by stars.
- #61★ 1,498+0Star change over the last 7 days
- #62
Source code accompanying book: Data Science on the Google Cloud Platform, Valliappa Lakshmanan, O'Reilly 2017
★ 1,429+1Star change over the last 7 days - #63
First open-source data discovery and observability platform. We make a life for data practitioners easy so you can focus on your business.
★ 1,427+1Star change over the last 7 days - #64
A comprehensive Python package template to kickstart and standardize your MLOps initiatives and data pipelines.
★ 1,416+0Star change over the last 7 days - #65
A curated collection of 300+ engineering blog articles from top tech companies. Learn how the best engineering teams solve real-world problems at scale.
★ 1,388+8Star change over the last 7 days - #66
The most comprehensive SQL guide from a real-world expert! Learn everything from basics to advanced queries, optimizations, and real-world SQL
★ 1,384+7Star change over the last 7 days - #67
Quilt is a Scientific Data Management Platform on AWS that helps teams and AI find, trust, and reuse data through deeply versioned, context-rich data packages.
★ 1,370+0Star change over the last 7 days - #68
Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality, reverse ETL, lineage, MCP for AI agents. No vendor cloud, no per-row billing.
★ 1,262+15Star change over the last 7 days - #69
A curated, but incomplete, list of data-centric AI resources.
★ 1,159+1Star change over the last 7 days - #70
A Data Engineering & Machine Learning Knowledge Hub
★ 1,146+0Star change over the last 7 days - #71
Home of the Open Data Contract Standard (ODCS).
★ 1,118+18Star change over the last 7 days - #72
📙 Awesome Data Catalogs and Observability Platforms.
★ 1,068+5Star change over the last 7 days - #73
Enforce Data Contracts
★ 1,060+8Star change over the last 7 days - #74
🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊
★ 972+7Star change over the last 7 days - #75
FlyFish is a data visualization coding platform. We can create a data model quickly in a simple way, and quickly generate a set of data visualization solutions by dragging.
★ 961+0Star change over the last 7 days - #76
A comprehensive guide to building a modern data warehouse with SQL Server, including ETL processes, data modeling, and analytics.
★ 934+14Star change over the last 7 days - #77★ 921+0Star change over the last 7 days
- #78
Un repositorio más con conceptos básicos, desafíos técnicos y recursos sobre ingeniería de datos en español 🧙✨
★ 896+1Star change over the last 7 days - #79
Community-driven, simple, yet powerful framework for fast, cost-effective distributed Compute over Data.
★ 871+1Star change over the last 7 days - #80
Neum AI is a best-in-class framework to manage the creation and synchronization of vector embeddings at large scale.
★ 867+0Star change over the last 7 days - #81
Supplementary Materials for the The Complete dbt (Data Build Tool) Bootcamp Udemy course
★ 827+0Star change over the last 7 days - #82
Practical Data Engineering: A Hands-On Real-Estate Project Guide
★ 823+2Star change over the last 7 days - #83
Python framework for building efficient data pipelines. It promotes modularity and collaboration, enabling the creation of complex pipelines from simple, reusable components.
★ 818+0Star change over the last 7 days - #84
Open-source agentic data engineering harness for dbt, SQL, and cloud warehouses. 100+ tools, 10 warehouses, AI-powered.
★ 808+3Star change over the last 7 days - #85
Know your data better!Datavines is Next-gen Data Observability Platform, support metadata manage and data quality.
★ 761+1Star change over the last 7 days - #86
Supercharge BigQuery with BigFunctions
★ 759+0Star change over the last 7 days - #87
Compilation of high-profile real-world examples of failed machine learning projects
★ 751+0Star change over the last 7 days - #88
VectorFlow is a high volume vector embedding pipeline that ingests raw data, transforms it into vectors and writes it to a vector DB of your choice.
★ 703+0Star change over the last 7 days - #89
数据流引擎是一款面向数据集成、数据同步、数据交换、数据共享、任务配置、任务调度的底层数据驱动引擎。数据流引擎采用管执分离、多流层、插件库等体系应对大规模数据任务、数据高频上报、数据高频采集、异构数据兼容的实际数据问题。
★ 696+0Star change over the last 7 days - #90★ 626+2Star change over the last 7 days