Skip to main content
buildradar
Sign in
Topic · parquet

parquet

Tracked open-source repos tagged parquet, sorted by stars.

Repos
39
Total stars
93,140
Avg. stars
2,388
Share
0.00%

Topics that frequently appear alongside parquet on the same repo.

Recent risers

Repos created in the last 90 days, tagged parquet.

  • hflow@Hebbian-Robotics

    Open source SDK for building multimodal data-quality, processing, enrichment, and curation pipelines for robotics and Physical AI.

    143
  • questdb@questdb

    QuestDB is a high performance, open-source, time-series database

    17,295+6Star change over the last 7 days
  • arrow@apache

    Apache Arrow is the universal columnar format and multi-language toolbox for fast data interchange and in-memory analytics

    17,078+6Star change over the last 7 days
  • Daft@Eventual-Inc

    High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale

    5,736+4Star change over the last 7 days
  • arrow-rs@apache

    Official Rust implementation of Apache Arrow

    3,600+4Star change over the last 7 days
  • roapi@roapi

    Create full-fledged APIs for slowly moving datasets without writing a single line of code.

    3,429+1Star change over the last 7 days
  • sail@lakehq

    Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.

    3,341+4Star change over the last 7 days
  • parquet-java@apache

    Apache Parquet Java

    3,078+0Star change over the last 7 days
  • rill@rilldata

    The fastest business intelligence tool for humans and agents.

    2,862+11Star change over the last 7 days
  • parquet-format@apache

    Apache Parquet Format

    2,563+7Star change over the last 7 days
  • parseable@parseablehq

    Parseable is an open source, unified infrastructure observability platform built in Rust on a data lake architecture. It tracks logs, metrics, traces, and events across apps, agents, and systems, reducing storage costs by up to 90% through columnar telemetry compression.

    2,450+2Star change over the last 7 days
  • drill@apache

    Apache Drill is a distributed MPP query layer for self describing data

    2,022+0Star change over the last 7 days
  • pg_mooncake@Mooncake-Labs

    Real-time analytics on Postgres tables

    2,003+0Star change over the last 7 days
  • petastorm@uber

    Petastorm library enables single machine or distributed training and evaluation of deep learning models from datasets in Apache Parquet format. It supports ML frameworks such as Tensorflow, Pytorch, and PySpark and can be used from pure Python code.

    1,892+0Star change over the last 7 days
  • nisshi@nisshi-io

    Apache Kafka® compatible broker with S3, PostgreSQL, SQLite, Apache Iceberg and Delta Lake

    1,853+1Star change over the last 7 days
  • anyquery@julien040

    One SQL interface for 60+ tools (e.g., GitHub, Notion, Airtable). Plug into any LLM through MCP.

    1,772+2Star change over the last 7 days
  • tonbo@tonbo-io

    Tonbo is an embedded database for serverless and edge runtimes.

    1,618+3Star change over the last 7 days
  • cryo@paradigmxyz

    cryo is the easiest way to extract blockchain data to parquet, csv, json, or python dataframes

    1,582+1Star change over the last 7 days
  • BemiDB@BemiHQ

    Open-source Snowflake & Fivetran alternative, with Postgres compatibility.

    1,531+0Star change over the last 7 days
  • olake@datazip-inc

    OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestion for real-time analytics. Supported sources : Postgres, MongoDB, MySQL, Oracle, MSSql, DB2, Kafka, S3.

    1,435+2Star change over the last 7 days
  • quilt@quiltdata

    Quilt is a Scientific Data Management Platform on AWS that helps teams and AI find, trust, and reuse data through deeply versioned, context-rich data packages.

    1,370+0Star change over the last 7 days
  • Simple Windows desktop application for viewing & querying Apache Parquet files

    1,231+2Star change over the last 7 days
  • ClickBench@ClickHouse

    ClickBench: a Benchmark For Analytical Databases

    1,100+8Star change over the last 7 days
  • adam@bigdatagenomics

    ADAM is a genomics analysis platform with specialized file formats built using Apache Avro, Apache Spark, and Apache Parquet. Apache 2 licensed.

    1,059+2Star change over the last 7 days
  • lonboard@developmentseed

    Fast, interactive geospatial data visualization in Jupyter.

    963+1Star change over the last 7 days
  • hyparquet@hyparam

    parquet file parser for javascript

    950+8Star change over the last 7 days
  • ChoETL@Cinchoo

    ETL framework for .NET (Parser / Writer for CSV, Flat, Xml, JSON, Key-Value, Parquet, Yaml, Avro formatted files)

    860+0Star change over the last 7 days
  • DevOps-Python-tools@HariSekhon

    80+ DevOps & Data CLI Tools - AWS, GCP, GCF Python Cloud Functions, Log Anonymizer, Spark, Hadoop, HBase, Hive, Impala, Linux, Docker, Spark Data Converters & Validators (Avro/Parquet/JSON/CSV/INI/XML/YAML), Travis CI, AWS CloudFormation, Elasticsearch, Solr etc.

    824+0Star change over the last 7 days
  • parquet-go@parquet-go

    High-performance Go package to read and write Parquet files

    769+2Star change over the last 7 days
  • kglab@DerwenAI

    Graph Data Science: an abstraction layer in Python for building knowledge graphs, integrated with popular graph libraries – atop Pandas, NetworkX, RAPIDS, RDFlib, pySHACL, PyVis, morph-kgc, pslpython, pyarrow, etc.

    691+2Star change over the last 7 days
  • pg_parquet@CrunchyData

    Copy to/from Parquet in S3, Azure Blob Storage, Google Cloud Storage, http(s) stores, local files or standard inout stream from within PostgreSQL

    685+0Star change over the last 7 days
← Back to topics