Parquet-to-Blosc2 Workflow and Round-Trip Validation
- Status
- Completed
- Confidentiality
- Public. Public personal project using generated synthetic data.
A reproducible storage-format workflow with multi-scale measurements and explicit round-trip validation.
I built a deterministic workflow that generates synthetic transaction data, imports Parquet into file-backed and directory-backed Blosc2 stores, exports it back to Parquet, and validates data preservation. Fresh 10,000, 100,000, and 1,000,000-row runs preserve local timing and size evidence without presenting it as a universal benchmark.
- Type
- Personal
- Categories
- Data & Dashboards
- Software & Automation
- Technologies
- Python
- pandas
- PyArrow
- Parquet
- Blosc2
- NumPy
Context
Data engineers and analysts evaluating reproducible import, export, and validation workflows.
- The dataset is synthetic and structurally narrow.
- Measurements come from one machine and one run per scale.
- The workflow does not measure query performance.
Problem
Storage-format demonstrations often report size or speed without retaining reproducible evidence that exported data still matches its source.
My contribution
I implemented deterministic data generation, three conversion paths, Parquet export, six groups of round-trip checks, benchmark CSV generation, and a reviewed isolated dependency run across three scales.
Approach
Generate fixed-seed data, convert it to `.b2z` and `.b2d` variants, export the primary store to Parquet, validate structure and values, and interpret timing and size only after validation passes.
Artifacts
Benchmark chart (Chart) Representative visual using synthetic data Open full-size visual: Benchmark chart Round-trip workflow (Diagram) Representative visual using synthetic data Open full-size visual: Round-trip workflow
Data artifacts
- Open data artifact: Benchmark results (CSV) Synthetic data
Outcomes
- All six validation groups passed at 10,000, 100,000, and 1,000,000 rows.
- Preserved 15 benchmark rows covering input and four workflow steps at each scale.
- Resolved exact binary dependencies in a disposable Python 3.12 environment.
Limitations and current status
- I treat these results as local workflow evidence, not a general storage recommendation.
- To reproduce this, use a reviewed hash-locked dependency file.