Repository navigation
File Format API for PyIceberg #3100
Description
Activity
This is going to be great, especially for new file formats!
I'm helping out with the TCK work over on Java and the shape of it might change a bit as we figure out exactly what we want the tests to look like. I don't think it would hurt for us to figure out our own testing story and then we can look into TCK when it's merged in.
Reacted by Fokko DriesprongI'm helping out with the TCK work over on Java and the shape of it might change a bit as we figure out exactly what we want the tests to look like. I don't think it would hurt for us to figure out our own testing story and then we can look into TCK when it's merged in.
Sounds good @rambleraptor. The TCK can come in when ready. The first few pieces are critical.
CC: @kevinjqliu @Fokko @geruh to see if the proposal makes senseThanks for writing this up! This is super exciting. I think we can focus on both the read and write path for parquet. I'm interested to see what the API looks like.
We dont have to mirror the java apis, but perhaps we can reuse similar ideasbtw @pvary, you started a revolution 🥳
Reacted by Drew, Neelesh Salian and Mrutunjay KinagiThanks @kevinjqliu . Could you assign this issue to me and I'll start pushing out the fix in phases.
Reacted by Kevin Liu@kevinjqliu @nssalian This is exciting. Is there anything I can contribute to this?
Thanks @nssalian for bringing this up, and it is very exciting indeed 🚀
We dont have to mirror the java apis, but perhaps we can reuse similar ideas
I would like to echo @kevinjqliu comment there. If we follow Java then we end up with many ABC's which are very common in Java, but are not considered very Pythonic.
Reacted by Neelesh SalianThanks @Fokko . I know you already looked at #3119 but feel free to comment if you think it could use a different direction or something needs to change. I'd like to set things right directionally before adding on more plumbing so I don't mind waiting to get the first foundational setup correct prior to proceeding.
- added a commit that references this issue
on May 4, 2026 - added a commit that references this issue
on Jul 19, 2026 - added a commit that references this issue
on Jul 27, 2026
Feature Request / Improvement
Problem
The write path in
pyiceberg/io/pyarrow.pyis hardcoded to Parquet. Thewrite.format.defaulttable property exists but is never read. Adding a new format (ORC, Vortex, Lance) requires modifying the monolithicwrite_file()function. The read path already dispatches multiple formats; the write path should too.Proposal
Introduce a File Format API aligned with Java Iceberg's File Format API (design doc).
New module
pyiceberg/io/fileformat.py:FileFormatWriter(ABC)FileFormatModel(ABC)FormatRegistryDataFileStatistics(it's inpyarrow.pycurrently but I think this might be good to consolidate for metrics)Changes to
pyiceberg/io/pyarrow.py:ParquetFormatWriter/ParquetFormatModelusing thewrite_parquet()(insidewrite_file()write_file()refactored to readwrite.format.default, look up the format model, and dispatch.TCK
tests/io/test_file_format_tck.py:Phased rollout:
write_file()dispatchJava ↔ Python Mapping
FormatModel<D, S>FileFormatModel(ABC, no type params)FileAppender<D>/ModelWriteBuilderFileFormatWriter(ABC)FormatModelRegistryFormatRegistry(keyed byFileFormatonly)MetricsDataFileStatistics(existing)test_file_format_tck.pyScope
This proposal covers the abstraction layer and the Parquet extraction only. No new format writers are included; ORC write support (#20) and any future formats (Avro, etc.) would be follow-ups once this lands.
References