Skip to content

Support data files compaction #1092

Description

@sungwy

Introduce an API to compact data files. The first version of the API will do the following:

  • take a predicate expression as input parameter to find data files matching the filter that will be re-written
  • group data files by partitions and rewrite them using the same bin-packing constraints of the writer

Activity

  1. self-assigned this
    on Aug 22, 2024
  2. changed the title [-]Compact data files[/-] [+]Support data files compaction[/+] on Aug 22, 2024
  3. removed their assignment
    on Sep 24, 2024
  4. sungwy commented on Sep 24, 2024

    @sungwy
    CollaboratorAuthor

    Unassigning to work on other near-term priorities

  5. github-actions commented on Mar 24, 2025

    @github-actions

    This issue has been automatically marked as stale because it has been open for 180 days with no activity. It will be closed in next 14 days if no further activity occurs. To permanently prevent this issue from being considered stale, add the label 'not-stale', but commenting on the issue is preferred when possible.

  6. zbs commented on Jun 1, 2025

    @zbs

    Is there any way to trigger compaction? The literature says that it's optimal to compact delete files back into data files to improve read space, and AFAICT there's no way to do this in PyIceberg.

    Incidentally, is there a way to control whether your catalog uses copy-on-write vs. merge-on-read?

  7. yingjianwu98 commented on Jun 26, 2025

    @yingjianwu98
    Contributor

    @sungwy

    Since my task #1931 (comment) is depending on the DeleteFileIndex so I am not going to work on it for now until the DeleteFileIndex task is complete.

    At the mean time, wondering if I can take this task if you haven't started working on this? Thanks!

  8. github-actions commented on Dec 24, 2025

    @github-actions

    This issue has been automatically marked as stale because it has been open for 180 days with no activity. It will be closed in next 14 days if no further activity occurs. To permanently prevent this issue from being considered stale, add the label 'not-stale', but commenting on the issue is preferred when possible.

  9. github-actions commented on Jan 8, 2026

    @github-actions

    This issue has been closed because it has not received any activity in the last 14 days since being marked as 'stale'

  10. qzyu999 commented on Feb 23, 2026

    @qzyu999
    Contributor

    Hi @kevinjqliu, is it possible for me to take a look at this issue? Doesn't look like anyone is currently working on it.

    CC: @sungwy, I am thinking you can also respond to this

  11. kevinjqliu commented on Feb 26, 2026

    @kevinjqliu
    Contributor

    feel free to start to work on it. might be a good idea to outline some ideas before proceeding.

    For example, i think we can model this similar to rewrite_data_files https://iceberg.apache.org/docs/nightly/spark-procedures/#rewrite_data_files

  12. qzyu999 commented on Mar 3, 2026

    @qzyu999
    Contributor

    Hi @kevinjqliu, thanks for the guidance. I've taken a look at the Java/Spark code (referencing v3.5), and I can see that there are quite a few options and nuances to the existing Java implementation. I believe for this issue, the goal is to focus more on some sort of MVP as described by @sungwy. Checking with the existing codebase from the latest iceberg-python main branch, I went ahead and checked what we have and what is missing for a potential MVP.

    We already have major components such as:

    • table.scan(filter).plan_files() (in pyiceberg/table/__init__.py for filtering based on a predicate as described in the OP by @sungwy)
    • Expression types (in pyiceberg/expressions)
    • ListPacker (in pyiceberg/utils/bin_packing.py for bin-packing the files together into manageable sizes)
    • various pyarrow functions to handle a single-node/python-only MVP setup (located in pyiceberg/io/pyarrow.py)
    • _OverwriteFiles for the commit (in pyiceberg/table/update/snapshot.py)
    • MaintenanceTable class (in pyiceberg/table/maintenance.py) where we can add a rewrite_data_files(table, row_filter) function

    I've tested some code locally, and I'm able to do the following:

    1. generate a partitioned table with many small files added across multiple snapshots
    2. then scan the table given a predicate filter to find all the needed data files
      2a. This also is able to handle the edge case (as seen in groupByPartition() from BinPackRewriteFilePlanner.java) where old partition specs after schema evolution changes needs to manage where old data files potentially from previous partitions need to be handled as a separate partition group to be rewritten specifically.
    3. process each partition (similar to planFileGroups in BinPackRewriteFilePlanner.java) using constant vars seen in the Java version of SizeBasedFileRewritePlanner - this effectively creates the list of all records across partitions that need to be rewritten AND it works in the case of selecting a specific set of rows using row_filter rather than automatically rewriting everything
      3a. filter files by size
      3b. bin-pack the files
      3c. filter groups (like in filterFileGroups from BinPackRewriteFilePlanner.java)
    4. create a list of both the old data files and the new data files (based on the bin-packing etc.)
    5. create a transaction to 1) delete the old files and 2) append the new files in a single commit

    There's one major nuance for this first MVP, which is that there's a expectedOutputFiles() in SizeBasedFileRewritePlanner.java, where it has an algorithm to handle the "remainder" problem (e.g., we may potentially write 10 large files and 1 small file rather than 10 files where 1 file is slightly larger than in the former case). PyIceberg apparently utilizes a bin_pack_arrow_table() within iceberg-python/pyiceberg/io/pyarrow.py, which is doing more of a real-time optimization rather than a planned optimization that potentially is more optimal. However, I feel that for this initial MVP it's not needed.

    What are your thoughts, is it okay to proceed and create a PR for this?

  13. kevinjqliu commented on Mar 5, 2026

    @kevinjqliu
    Contributor

    Thanks for taking the time to look into this @qzyu999! I think this is on the right track.

    Looking at the rewrite_data_files implementation in spark, theres a lot of bells and whistles (probably added over time). For the pyiceberg implementation, It might be useful to scope the feature down as much as possible; just to create a harness and we can improve it over time.

    What do you think about first handling the case for compaction of a whole table? That way to don't have to deal with filter and matching data files.

    Im thinking something like table.maintenance.compact(), which will rewrite the table using the REPLACE operation.
    For the actual data files, we can take a shortcut and just binpack by reading the table and writing it out again. This should produce the desired file size specified by write.target-file-size-bytes (which the write path already uses)

    WDYT?

  14. added a commit that references this issue on Mar 6, 2026
    5c8dc67
  15. qzyu999 commented on Mar 6, 2026

    @qzyu999
    Contributor

    Hi @kevinjqliu, thanks for the suggestion, I think you definitely make sense regarding starting simple. I added a PR: 5c8dc67, where I add compact() to the existing MaintenanceTable class. I've also added some additional tests. Please let me know what you think about the changes.

    Edit: I added a commit, 2774bd3, to address linter issues.

  16. added a commit that references this issue on Mar 7, 2026
    edf449e
  17. djouallah commented on May 26, 2026

    @djouallah

    any update on this ?

  18. qzyu999 commented on May 26, 2026

    @qzyu999
    Contributor

    any update on this ?

    Hi @djouallah, this is the current dependency chain for this specific issue:

    There had been a good amount of work done for #3131, but after some time where the reviewers were busy (@kevinjqliu, @geruh) there was a new issue/PR (#3319 / #3320) that referenced #3130 / #3131. From there it became apparent that once #3320 merged it would require a refactor of #3320. We can in the meantime pass a version of #3320 that would do the same work, but it would still require a refactor later to fit the changes of #3320. Also, #3131 is still waiting some further code reviews. If this is urgent we can possibly complete #3131 with an ad-hoc version of the validation stage and come back to #3124. However, it would still require some help from @kevinjqliu and @geruh to complete the review of #3131. Thank you for your patience.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions