apache-arrow
databow is a command-line tool for querying databases.
A command-line tool for querying databases via ADBC.
Sub-millisecond cache for ML/AI workloads. Parquets in, Arrow-Flight out.
Murr is a caching layer for ML/AI data serving that sits between your batch data pipelines and inference apps.
A RocksDB-based NVMe/S3 cache for AI inference workloads. A faster Redis replacement, optimized for batch low-latency zero-copy reads and writes.
Fastest Time-Series Database. High-performance time-series data warehouse built on DuckDB and Parquet with flexible storage options.
Time-series data warehouse built for speed. 2.42M records/sec on local NVMe. DuckDB + Parquet + Arrow + flexible storage (local/MinIO/S3). AGPL-3.0
"The LLVM of columnar file formats". A toolkit for working with compressed Arrow on-disk, in-memory, and over-the-wire.
Vortex is a toolkit for working with compressed Apache Arrow arrays in-memory, on-disk, and over-the-wire.
Vortex is designed to be to columnar file formats what Apache DataFusion is to query engines (or, analogously, what LLVM + Clang are to compilers): a highly extensible & extremely fast framework for building a modern columnar file format, with a state-of-the-art, "batteries included" reference implementation.