(illustration hand generated in 1958)
Julien Le Dem (julien@apache.org)
Column Storage
for the AI era
M.C. Escher, Belvedere, lithograph, May 1958
licensed under CC BY-NC-SA 4.0
2
Julien Le Dem�julien.ledem.net
sympathetic.ink
Principal Engineer at
Technical Advisory Council
About the Speaker
3
Agenda
Session
The Design of Parquet:
Praise and Gripes
1
Inception
5
Started in 2012�
De-facto standard for columnar storage
Widely supported
Enabled interoperability
For the new generation of vectorized query engines inspired by the C-Store and MonetDB/X100 papers.
Based on Google Dremel paper
Store nested structures in a columnar representation.
Main trade offs at the time
Cost of
storage
Time to decode with a CPU
Time to transfer on the wire
Access only the data you need
Columnar
Statistics
Read only the data you need!
Parquet file layout
Page
9
“abc”
“def”
“ghi”
Values
Bytes: 1, 2, 3
Encoded
bytes
encoding
compression*
* optional
Praise
Scalable
De-facto standard with engaged community of implementers
Self-documenting
Saves on storage. Columnar layout helps.
Familiar
No more CSV type inference.
What people like
Flexible
Multi-level statistics.
10
Good compression
Data skipping
Wide adoption/interoperability
Self describing schema
Gripes
Too reliant on generic compression
ZSTD is great, but we can be faster for specific types.
Metadata large with many columns
More wide schema use cases (Feature stores)
What people don’t like
Not optimized for Random access
Need to decode too many values
Not parallelizable enough
Existing encodings do not take full advantage of SIMD and GPUs.
11
What has changed?
2
More parallelism
From 2010 to 2025
13
More Cores:
16x ��(8 -> 128)
Wider SIMD:
4x
�(128 bits -> 512 bits)
GPUs:
40x more threads
From dozens of threads to thousands
AI Access patterns
14
Metadata
Index
What can we do?
3
Healthy pressure created by the newcomers
16
Have things fundamentally changed?
17
Metadata
Metadata
Metadata
Metadata
V1
V2
V3
What is Parquet?
Parquet, first and foremost, is an Open Source project. It’s a consensus building machine which pulls a huge community of open source projects and vendors in its wake.
It is easier to popularize new techniques through the existing machine rather than getting 4 new formats adopted.
Building the community around the project is the hard part.
We stand on the shoulders of giants.
Variant: Recent Example of Parquet evolving and pulling the ecosystem along
19
Vendor and Project collab
Cross compatibility tests
BigQuery, Snowflake, Databricks, DataFusion, …
Proposal:
from Databricks
Should it go in Parquet?
Or Spark?
Or Arrow?
Or Iceberg?
Finalize consensus:
Semistructured data in columns
Flexibility to shred some in separate columns
Took 1 year from 1st suggestion to release
Why new encodings?
Problems of existing encodings
21
Limits the throughput of decoding.
Analytics use cases are also looking at GPUs to speed up things.
Those needs are not at odds.
What are those new encodings?
Lightweight encodings:
Concrete next steps
4
Tuning the existing Parquet settings
Scalable
Pick encodings with certain characteristics at write time
Self-documenting
You can have only one if you wish.
Familiar
Pages can be smaller.�Pages can be of varying size
Do we need larger row groups and smaller pages?
Flexible
It is optional
24
Parquet has row groups
General purpose block compression (zstd) is too slow
Some encodings don’t allow random access
Parquet pages force to decompress too many values
Do we really need millions of columns?
How many rows can we fit in a GB?
=> degraded performance of columnar representation.
=> Variant
Variant:
Better use of metadata
26
You don’t need to read the metadata from the file.
=> Iceberg
Adding those new encodings
27
Parquet is based on research and will keep integrating new ideas.
=> new encodings + improved metadata layout.
=> encodings
Extending Parquet
28
See: Datafusion blog: User defined Parquet Indexes
Everything else can be built on top.
Thanks :)