data-pipeline
Tracked open-source repos tagged data-pipeline, sorted by stars.
Related topics
Topics that frequently appear alongside data-pipeline on the same repo.
Recent risers
Repos created in the last 90 days, tagged data-pipeline.
- #1
Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.
★ 21,981+9Star change over the last 7 days - #2
Empowering Data Intelligence with Distributed SQL for Sharding, Scalability, and Security Across All Databases.
★ 20,788-3Star change over the last 7 days - #3
Change data capture for a variety of databases. Please log issues at https://github.com/debezium/dbz/issues.
★ 13,072+16Star change over the last 7 days - #4
High-performance AI pipeline engine with a C++ core and 50+ Python-extensible nodes. Build, debug, and scale LLM workflows with 13+ model providers, 8+ vector databases, and agent orchestration, all from your IDE. Includes VS Code extension, TypeScript/Python SDKs, and Docker deployment.
★ 7,641+270Star change over the last 7 days - #5★ 7,032+2Star change over the last 7 days
- #6
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
★ 6,975+26Star change over the last 7 days - #7★ 6,469+0Star change over the last 7 days
- #8★ 5,019+2Star change over the last 7 days
- #9
Privacy and Security focused Segment-alternative, in Golang and React
★ 4,483+4Star change over the last 7 days - #10
A list of useful resources to learn Data Engineering from scratch
★ 4,009-2Star change over the last 7 days - #11
ingestr is a CLI tool to copy data between any databases with a single command seamlessly.
★ 3,935+18Star change over the last 7 days - #12
Self-hostable workflow orchestrator for teams whose main work isn't orchestration. Declarative YAML over your scripts, SSH commands, containers, etc; keep workflows separate from business logic. One binary, no database, runs on limited H/W resources. Alternative to Airflow / Cron / Job Scheduler.
★ 3,832+17Star change over the last 7 days - #13★ 3,440+0Star change over the last 7 days
- #14
Open-source inference server and production cluster for all the models your agent needs.
★ 3,030+173Star change over the last 7 days - #15
An open-source data logging library for machine learning models and data pipelines. 📚 Provides visibility into data quality & model performance over time. 🛡️ Supports privacy-preserving data collection, ensuring safety & robustness. 📈
★ 2,831+0Star change over the last 7 days - #16
The dbt-native data observability solution for data & analytics engineers. Monitor your data pipelines in minutes. Available as self-hosted or cloud service with premium features.
★ 2,405+4Star change over the last 7 days - #17
A lightweight stream processing library for Go
★ 2,172+1Star change over the last 7 days - #18★ 2,083-1Star change over the last 7 days
- #19
🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.
★ 1,673+0Star change over the last 7 days - #20
OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestion for real-time analytics. Supported sources : Postgres, MongoDB, MySQL, Oracle, MSSql, DB2, Kafka, S3.
★ 1,435+2Star change over the last 7 days - #21
Source code accompanying book: Data Science on the Google Cloud Platform, Valliappa Lakshmanan, O'Reilly 2017
★ 1,429+1Star change over the last 7 days - #22
Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality, reverse ETL, lineage, MCP for AI agents. No vendor cloud, no per-row billing.
★ 1,262+15Star change over the last 7 days - #23
zerocode-tdd is a community-developed, free, open-source, outcome-driven automated testing framework for Data Pipelines, ETL, REST API, Kafka(Data Streams), Databases and Load scenarios. all defined in simple JSON or YAML — with zero coding.
★ 1,013+1Star change over the last 7 days - #24
🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊
★ 972+7Star change over the last 7 days - #25
SeaTunnel is a distributed, high-performance data integration platform for the synchronization and transformation of massive data (offline & real-time).
★ 877+5Star change over the last 7 days - #26★ 876+2Star change over the last 7 days
- #27
Pythonic tool for orchestrating machine-learning/high performance/quantum-computing workflows in heterogeneous compute environments.
★ 869+0Star change over the last 7 days - #28
Practical Data Engineering: A Hands-On Real-Estate Project Guide
★ 823+2Star change over the last 7 days - #29
Conduit streams data between data stores. Kafka Connect replacement. No JVM required.
★ 608+0Star change over the last 7 days - #30
A curated list of open source tools used in analytics platforms and data engineering ecosystem
★ 604+3Star change over the last 7 days