Skip to main content
buildradar
Sign in
Topic · data-pipelines

data-pipelines

Tracked open-source repos tagged data-pipelines, sorted by stars.

Repos
25
Total stars
221,273
Avg. stars
8,851
Share
0.01%

Topics that frequently appear alongside data-pipelines on the same repo.

Recent risers

Repos created in the last 90 days, tagged data-pipelines.

No new repos tagged with this topic in the last 90 days.

  • pathway@pathwaycom

    Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.

    62,348-17Star change over the last 7 days
  • airflow@apache

    Apache Airflow - A platform to programmatically author, schedule, and monitor workflows

    46,701+63Star change over the last 7 days
  • dagster@dagster-io

    An orchestration platform for the development, production, and observation of data assets.

    16,092+19Star change over the last 7 days
  • unstructured@Unstructured-IO

    Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

    15,381+19Star change over the last 7 days
  • Apache DolphinScheduler is the modern data orchestration platform. Agile to create high performance workflow with low-code

    14,456+1Star change over the last 7 days
  • kedro@kedro-org

    Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducible, maintainable, and modular.

    10,989+14Star change over the last 7 days
  • mage-ai@mage-ai

    🧙 Build, run, and manage data pipelines for integrating and transforming data.

    8,818+4Star change over the last 7 days
  • DataFlow@OpenDCAI

    Easy Data Preparation with latest LLMs-based Operators and Pipelines.

    7,893+68Star change over the last 7 days
  • fluvio@fluvio-community

    🦀 event stream processing for developers to collect and transform data in motion to power responsive data intensive applications.

    5,248-1Star change over the last 7 days
  • preswald@StructuredLabs

    Preswald is a WASM packager for Python-based interactive data apps: bundle full complex data workflows, particularly visualizations, into single files, runnable completely in-browser, using Pyodide, DuckDB, Pandas, and Plotly, Matplotlib, etc. Build dashboards, reports, and notebooks that run offline, load fast, and share like a document.

    4,276-2Star change over the last 7 days
  • docetl@ucbepic

    A system for agentic LLM-powered data processing and ETL

    4,074+34Star change over the last 7 days
  • maestro@Netflix

    Maestro: Netflix’s Workflow Orchestrator

    3,831+0Star change over the last 7 days
  • meltano@meltano

    Meltano: the declarative code-first data integration engine that powers your wildest data and ML-powered product ideas. Say goodbye to writing, maintaining, and scaling your own API integrations.

    2,616+2Star change over the last 7 days
  • elementary@elementary-data

    The dbt-native data observability solution for data & analytics engineers. Monitor your data pipelines in minutes. Available as self-hosted or cloud service with premium features.

    2,405+4Star change over the last 7 days
  • feldera@feldera

    The Feldera Incremental Computation Engine

    2,076+8Star change over the last 7 days
  • data-engineering-wiki@data-engineering-community

    The best place to learn data engineering. Built and maintained by the data engineering community.

    2,021+3Star change over the last 7 days
  • extractous@yobix-ai

    Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.

    1,772+0Star change over the last 7 days
  • bruin@bruin-data

    Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.

    1,683+4Star change over the last 7 days
  • mleap@combust

    MLeap: Deploy ML Pipelines to Production

    1,544+1Star change over the last 7 days
  • pyper@pyper-dev

    Concurrent Python made simple

    1,517-1Star change over the last 7 days
  • odd-platform@opendatadiscovery

    First open-source data discovery and observability platform. We make a life for data practitioners easy so you can focus on your business.

    1,427+1Star change over the last 7 days
  • mlops-python-package@fmind

    A comprehensive Python package template to kickstart and standardize your MLOps initiatives and data pipelines.

    1,416+0Star change over the last 7 days
  • amphi-etl@amphi-ai

    visual data prep powered by python

    1,401+0Star change over the last 7 days
  • optimus@raystack

    Optimus is an easy-to-use, reliable, and performant workflow orchestrator for data transformation, data modeling, pipelines, and data quality management.

    766+2Star change over the last 7 days
  • dbt-data-reliability@elementary-data

    This dbt package captures metadata, artifacts, and test results so you can detect anomalies, monitor data quality, and build metadata tables. It powers Elementary OSS and feeds the wider context layer used by Elementary Cloud’s full Data & AI Control Plane.

    523+0Star change over the last 7 days
← Back to topics