Oct 03, 2026
Lakehouse – Phase 2 Parquet Data Files
Phase 1 built the metadata catalog — Postgres tables that track what tables exist, what columns they have, and which files belong to them. But metadata alone is useless without actual data. This phase builds the data file layer: writing and reading Apache Parquet files with per-column statistics that the metadata catalog will reference for file pruning. We use the parquet-go library rather than implementing Parquet from scratch — the educational value is in understanding how Parquet’s structure enables file pruning and schema evolution.