QuestDB is a high performance, open-source, time-series database
Parquet GitHub Repositories
Explore popular GitHub repositories tagged “parquet”.
Compare stars, forks, and programming language using the same GitStar view as GitHub Trending.
Trending Repositories
Apache Arrow is the universal columnar format and multi-language toolbox for fast data interchange and in-memory analytics
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
Commandline tool for running SQL queries against JSON, CSV, Excel, Parquet, and more.
Official Rust implementation of Apache Arrow
Create full-fledged APIs for slowly moving datasets without writing a single line of code.
Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.
Apache Parquet Java
The fastest business intelligence tool for humans and agents.
Apache Parquet Format
Parseable is an open source, unified infrastructure observability platform built in Rust on a data lake architecture. It tracks logs, metrics, traces, and events across apps, agents, and systems, reducing storage costs by up to 90% through columnar telemetry compression.
Apache Drill is a distributed MPP query layer for self describing data
Real-time analytics on Postgres tables
Petastorm library enables single machine or distributed training and evaluation of deep learning models from datasets in Apache Parquet format. It supports ML frameworks such as Tensorflow, Pytorch, and PySpark and can be used from pure Python code.
Apache Kafka® compatible broker with S3, PostgreSQL, SQLite, Apache Iceberg and Delta Lake
One SQL interface for 60+ tools (e.g., GitHub, Notion, Airtable). Plug into any LLM through MCP.
Tonbo is an embedded database for serverless and edge runtimes.
cryo is the easiest way to extract blockchain data to parquet, csv, json, or python dataframes
Open-source Snowflake & Fivetran alternative, with Postgres compatibility.
OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestion for real-time analytics. Supported sources : Postgres, MongoDB, MySQL, Oracle, MSSql, DB2, Kafka, S3.
Quilt is a Scientific Data Management Platform on AWS that helps teams and AI find, trust, and reuse data through deeply versioned, context-rich data packages.
Simple Windows desktop application for viewing & querying Apache Parquet files
ClickBench: a Benchmark For Analytical Databases
ADAM is a genomics analysis platform with specialized file formats built using Apache Avro, Apache Spark, and Apache Parquet. Apache 2 licensed.
Fast, interactive geospatial data visualization in Jupyter.
ETL framework for .NET (Parser / Writer for CSV, Flat, Xml, JSON, Key-Value, Parquet, Yaml, Avro formatted files)
parquet file parser for javascript
80+ DevOps & Data CLI Tools - AWS, GCP, GCF Python Cloud Functions, Log Anonymizer, Spark, Hadoop, HBase, Hive, Impala, Linux, Docker, Spark Data Converters & Validators (Avro/Parquet/JSON/CSV/INI/XML/YAML), Travis CI, AWS CloudFormation, Elasticsearch, Solr etc.
High-performance Go package to read and write Parquet files
Graph Data Science: an abstraction layer in Python for building knowledge graphs, integrated with popular graph libraries – atop Pandas, NetworkX, RAPIDS, RDFlib, pySHACL, PyVis, morph-kgc, pslpython, pyarrow, etc.