data-pipelines
Tracked open-source repos tagged data-pipelines, sorted by stars.
Related topics
Topics that frequently appear alongside data-pipelines on the same repo.
Recent risers
Repos created in the last 90 days, tagged data-pipelines.
No new repos tagged with this topic in the last 90 days.
- #1
Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.
★ 62,348-17Star change over the last 7 days - #2
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
★ 46,701+63Star change over the last 7 days - #3
An orchestration platform for the development, production, and observation of data assets.
★ 16,092+19Star change over the last 7 days - #4
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
★ 15,381+19Star change over the last 7 days - #5
Apache DolphinScheduler is the modern data orchestration platform. Agile to create high performance workflow with low-code
★ 14,456+1Star change over the last 7 days - #6
Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, and modular.
★ 10,989+14Star change over the last 7 days - #7★ 8,818+4Star change over the last 7 days
- #8★ 7,893+68Star change over the last 7 days
- #9
🦀 event stream processing for developers to collect and transform data in motion to power responsive data intensive applications.
★ 5,248-1Star change over the last 7 days - #10
Preswald is a WASM packager for Python-based interactive data apps: bundle full complex data workflows, particularly visualizations, into single files, runnable completely in-browser, using Pyodide, DuckDB, Pandas, and Plotly, Matplotlib, etc. Build dashboards, reports, and notebooks that run offline, load fast, and share like a document.
★ 4,276-2Star change over the last 7 days - #11★ 4,074+34Star change over the last 7 days
- #12★ 3,831+0Star change over the last 7 days
- #13
Meltano: the declarative code-first data integration engine that powers your wildest data and ML-powered product ideas. Say goodbye to writing, maintaining, and scaling your own API integrations.
★ 2,616+2Star change over the last 7 days - #14
The dbt-native data observability solution for data & analytics engineers. Monitor your data pipelines in minutes. Available as self-hosted or cloud service with premium features.
★ 2,405+4Star change over the last 7 days - #15★ 2,076+8Star change over the last 7 days
- #16
The best place to learn data engineering. Built and maintained by the data engineering community.
★ 2,021+3Star change over the last 7 days - #17
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
★ 1,772+0Star change over the last 7 days - #18
Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.
★ 1,683+4Star change over the last 7 days - #19★ 1,544+1Star change over the last 7 days
- #20★ 1,517-1Star change over the last 7 days
- #21
First open-source data discovery and observability platform. We make a life for data practitioners easy so you can focus on your business.
★ 1,427+1Star change over the last 7 days - #22
A comprehensive Python package template to kickstart and standardize your MLOps initiatives and data pipelines.
★ 1,416+0Star change over the last 7 days - #23★ 1,401+0Star change over the last 7 days
- #24
Optimus is an easy-to-use, reliable, and performant workflow orchestrator for data transformation, data modeling, pipelines, and data quality management.
★ 766+2Star change over the last 7 days - #25
This dbt package captures metadata, artifacts, and test results so you can detect anomalies, monitor data quality, and build metadata tables. It powers Elementary OSS and feeds the wider context layer used by Elementary Cloud’s full Data & AI Control Plane.
★ 523+0Star change over the last 7 days