Repository navigation
Support data files compaction #1092
Description
Activity
Unassigning to work on other near-term priorities
This issue has been automatically marked as stale because it has been open for 180 days with no activity. It will be closed in next 14 days if no further activity occurs. To permanently prevent this issue from being considered stale, add the label 'not-stale', but commenting on the issue is preferred when possible.
Is there any way to trigger compaction? The literature says that it's optimal to compact delete files back into data files to improve read space, and AFAICT there's no way to do this in PyIceberg.
Incidentally, is there a way to control whether your catalog uses copy-on-write vs. merge-on-read?
Since my task #1931 (comment) is depending on the DeleteFileIndex so I am not going to work on it for now until the DeleteFileIndex task is complete.
At the mean time, wondering if I can take this task if you haven't started working on this? Thanks!
Reacted by Sung YunThis issue has been automatically marked as stale because it has been open for 180 days with no activity. It will be closed in next 14 days if no further activity occurs. To permanently prevent this issue from being considered stale, add the label 'not-stale', but commenting on the issue is preferred when possible.
This issue has been closed because it has not received any activity in the last 14 days since being marked as 'stale'
Hi @kevinjqliu, is it possible for me to take a look at this issue? Doesn't look like anyone is currently working on it.
CC: @sungwy, I am thinking you can also respond to this
feel free to start to work on it. might be a good idea to outline some ideas before proceeding.
For example, i think we can model this similar to
rewrite_data_fileshttps://iceberg.apache.org/docs/nightly/spark-procedures/#rewrite_data_filesReacted by Jared Yu and DrewHi @kevinjqliu, thanks for the guidance. I've taken a look at the Java/Spark code (referencing v3.5), and I can see that there are quite a few options and nuances to the existing Java implementation. I believe for this issue, the goal is to focus more on some sort of MVP as described by @sungwy. Checking with the existing codebase from the latest iceberg-python main branch, I went ahead and checked what we have and what is missing for a potential MVP.
We already have major components such as:
table.scan(filter).plan_files()(inpyiceberg/table/__init__.pyfor filtering based on a predicate as described in the OP by @sungwy)- Expression types (in
pyiceberg/expressions) ListPacker(inpyiceberg/utils/bin_packing.pyfor bin-packing the files together into manageable sizes)- various
pyarrowfunctions to handle a single-node/python-only MVP setup (located inpyiceberg/io/pyarrow.py) - _OverwriteFiles for the commit (in
pyiceberg/table/update/snapshot.py) - MaintenanceTable class (in
pyiceberg/table/maintenance.py) where we can add arewrite_data_files(table, row_filter)function
I've tested some code locally, and I'm able to do the following:
- generate a partitioned table with many small files added across multiple snapshots
- then scan the table given a predicate filter to find all the needed data files
2a. This also is able to handle the edge case (as seen ingroupByPartition()fromBinPackRewriteFilePlanner.java) where old partition specs after schema evolution changes needs to manage where old data files potentially from previous partitions need to be handled as a separate partition group to be rewritten specifically. - process each partition (similar to
planFileGroupsinBinPackRewriteFilePlanner.java) using constant vars seen in the Java version ofSizeBasedFileRewritePlanner- this effectively creates the list of all records across partitions that need to be rewritten AND it works in the case of selecting a specific set of rows usingrow_filterrather than automatically rewriting everything
3a. filter files by size
3b. bin-pack the files
3c. filter groups (like infilterFileGroupsfromBinPackRewriteFilePlanner.java) - create a list of both the old data files and the new data files (based on the bin-packing etc.)
- create a transaction to 1) delete the old files and 2) append the new files in a single commit
There's one major nuance for this first MVP, which is that there's a
expectedOutputFiles()inSizeBasedFileRewritePlanner.java, where it has an algorithm to handle the "remainder" problem (e.g., we may potentially write 10 large files and 1 small file rather than 10 files where 1 file is slightly larger than in the former case). PyIceberg apparently utilizes abin_pack_arrow_table()withiniceberg-python/pyiceberg/io/pyarrow.py, which is doing more of a real-time optimization rather than a planned optimization that potentially is more optimal. However, I feel that for this initial MVP it's not needed.What are your thoughts, is it okay to proceed and create a PR for this?
Thanks for taking the time to look into this @qzyu999! I think this is on the right track.
Looking at the
rewrite_data_filesimplementation in spark, theres a lot of bells and whistles (probably added over time). For the pyiceberg implementation, It might be useful to scope the feature down as much as possible; just to create a harness and we can improve it over time.What do you think about first handling the case for compaction of a whole table? That way to don't have to deal with
filterand matching data files.Im thinking something like
table.maintenance.compact(), which will rewrite the table using theREPLACEoperation.
For the actual data files, we can take a shortcut and just binpack by reading the table and writing it out again. This should produce the desired file size specified bywrite.target-file-size-bytes(which the write path already uses)WDYT?
Reacted by Jared Yu- added a commit that references this issue
on Mar 6, 2026 Hi @kevinjqliu, thanks for the suggestion, I think you definitely make sense regarding starting simple. I added a PR: 5c8dc67, where I add
compact()to the existingMaintenanceTableclass. I've also added some additional tests. Please let me know what you think about the changes.Edit: I added a commit, 2774bd3, to address linter issues.
Reacted by Kevin Liu and Danny C Jones- added a commit that references this issue
on Mar 7, 2026 any update on this ?
Reacted by Jared Yuany update on this ?
Hi @djouallah, this is the current dependency chain for this specific issue:
There had been a good amount of work done for #3131, but after some time where the reviewers were busy (@kevinjqliu, @geruh) there was a new issue/PR (#3319 / #3320) that referenced #3130 / #3131. From there it became apparent that once #3320 merged it would require a refactor of #3320. We can in the meantime pass a version of #3320 that would do the same work, but it would still require a refactor later to fit the changes of #3320. Also, #3131 is still waiting some further code reviews. If this is urgent we can possibly complete #3131 with an ad-hoc version of the validation stage and come back to #3124. However, it would still require some help from @kevinjqliu and @geruh to complete the review of #3131. Thank you for your patience.
Reacted by Robin FourcadeReacted by Mimoune
Introduce an API to compact data files. The first version of the API will do the following: