Skip to main content
buildradar
Sign in
Topic · data-processing

data-processing

Tracked open-source repos tagged data-processing, sorted by stars.

Repos
30
Total stars
165,877
Avg. stars
5,529
Share
0.01%

Topics that frequently appear alongside data-processing on the same repo.

Recent risers

Repos created in the last 90 days, tagged data-processing.

  • marvis-risk-agent@eddyzzl

    MARVIS-Agent: all-purpose credit risk agent for model development, validation, data processing, feature engineering, and strategy workflows.

    518
  • pathway@pathwaycom

    Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.

    62,365-36Star change over the last 7 days
  • cocoindex@cocoindex-io

    Incremental engine for long horizon agents 🌟 Star if you like it!

    11,430+48Star change over the last 7 days
  • A collection of handy Bash One-Liners and terminal tricks for data processing and Linux system maintenance.

    10,767+5Star change over the last 7 days
  • miller@johnkerl

    Miller is like awk, sed, cut, join, and sort for name-indexed data such as CSV, TSV, and tabular JSON

    10,005+4Star change over the last 7 days
  • dasel@TomWright

    Unified querying, transformation, and modification of JSON, TOML, YAML, XML, INI, HCL, KDL and CSV.

    8,028+7Star change over the last 7 days
  • DataFlow@OpenDCAI

    Easy Data Preparation with latest LLMs-based Operators and Pipelines.

    7,825+164Star change over the last 7 days
  • rocketride-server@rocketride-org

    High-performance AI pipeline engine with a C++ core and 50+ Python-extensible nodes. Build, debug, and scale LLM workflows with 13+ model providers, 8+ vector databases, and agent orchestration, all from your IDE. Includes VS Code extension, TypeScript/Python SDKs, and Docker deployment.

    7,371+416Star change over the last 7 days
  • data-juicer@datajuicer

    Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷

    6,949+28Star change over the last 7 days
  • DALI@NVIDIA

    A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.

    5,738+4Star change over the last 7 days
  • smallpond@deepseek-ai

    A lightweight data processing framework built on DuckDB and 3FS.

    5,002+5Star change over the last 7 days
  • pandera@unionai-oss

    A light-weight, flexible, and expressive statistical data testing library

    4,443+5Star change over the last 7 days
  • numaflow@numaproj

    Kubernetes-native platform to run massively parallel data/streaming jobs

    2,827+2Star change over the last 7 days
  • datachain@datachain-ai

    The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure

    2,811+3Star change over the last 7 days
  • broadway@elixir-broadway

    Concurrent and multi-stage data ingestion and data processing with Elixir

    2,677+1Star change over the last 7 days
  • texar@asyml

    Toolkit for Machine Learning, Natural Language Processing, and Text Generation, in TensorFlow. This is part of the CASL project: http://casl-project.ai/

    2,389+0Star change over the last 7 days
  • bytewax@bytewax

    Python Stream Processing

    2,048+1Star change over the last 7 days
  • Curator@NVIDIA-NeMo

    Scalable data pre processing and curation toolkit for LLMs

    1,740+12Star change over the last 7 days
  • dolma@allenai

    Data and tools for generating and inspecting OLMo pre-training data.

    1,538+2Star change over the last 7 days
  • pyper@pyper-dev

    Concurrent Python made simple

    1,518+0Star change over the last 7 days
  • data-science-on-gcp@GoogleCloudPlatform

    Source code accompanying book: Data Science on the Google Cloud Platform, Valliappa Lakshmanan, O'Reilly 2017

    1,428+1Star change over the last 7 days
  • kubetorch@run-house

    Distribute and run AI workloads on Kubernetes magically in Python, like PyTorch for ML infra.

    1,224+0Star change over the last 7 days
  • xidel@benibela

    Command line tool to download and extract data from HTML/XML pages or JSON-APIs, using CSS, XPath 3.0, XQuery 3.0, JSONiq or pattern matching. It can also create new or transformed XML/HTML/JSON documents.

    843+0Star change over the last 7 days
  • text-dedup@ChenghaoMou

    All-in-one text de-duplication

    765+0Star change over the last 7 days
  • hstream@hstreamdb

    HStreamDB is an open-source, cloud-native streaming database for IoT and beyond. Modernize your data stack for real-time applications.

    722+0Star change over the last 7 days
  • collapse@fastverse

    Advanced and Fast Data Transformation in R

    705+1Star change over the last 7 days
  • A list about Apache Kafka

    591+0Star change over the last 7 days
  • eternal@kousun12

    👾~ music, eternal ~ 👾

    581+1Star change over the last 7 days
  • synthBTC@jofpin

    A tool that uses advanced Monte Carlo simulations and Turbit parallel processing to create possible Bitcoin prediction scenarios.

    526+0Star change over the last 7 days
  • marvis-risk-agent@eddyzzl

    MARVIS-Agent: all-purpose credit risk agent for model development, validation, data processing, feature engineering, and strategy workflows.

    518+2Star change over the last 7 days
  • Musoq@Puchaczov

    SQL Runtime without any database

    505+1Star change over the last 7 days
← Back to topics