Repository navigation
Conversation
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Filter-only partition columns still fall back to the initial default when excluded from the output projection.
Review effort: Balanced
Findings: 1
What changed in this PR
Fixes null identity partition projection so it takes precedence over a field’s initial default.
Changes:
- Preserves explicit null partition values during projection and filtering.
- Adds regression coverage for null projection and predicates.
| File | Description |
|---|---|
pyiceberg/io/pyarrow.py |
Distinguishes explicit null projections from absent values. |
tests/io/test_pyarrow.py |
Tests null identity partitions with initial defaults. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| partition_value = accessors[partition_field.field_id].get(file.partition) | ||
| if partition_value is not None: | ||
| projected_missing_fields[field_id] = partition_value | ||
| projected_missing_fields[field_id] = accessors[partition_field.field_id].get(file.partition) |
There was a problem hiding this comment.
Thanks for the fix. The explicit-null handling looks correct and matches the Column Projection precedence rules.
I checked head 491042af8c82f6b908405d8b1439af3784686ae5 and reproduced the filter-only case already noted by Copilot. With a physical file containing only other_field=["foo", "bar"], manifest partition partition_col=None, and initial_default=7, projecting only other_field gives:
IsNull("partition_col"): no rows, instead of both rows.EqualTo("partition_col", 7): both rows, instead of none.
These were reader-level checks through _task_to_record_batches, with the filter field included in projected_field_ids but excluded from projected_schema.
#4082 supplies the union of output and filter field IDs needed here. Its runtime patch applies cleanly to this PR, and all nine focused cases pass with the two fixes combined. I suggest parameterizing the new regression test to also project only other_field, covering both predicates, and merging #4082 before or alongside this PR.
Validation: the unchanged PR's Arrow and expression visitor unit suites passed (285 passed, 3 skipped). The additional six-case projection/filter matrix had four passes and the two failures above on this PR alone; combined with #4082, all six matrix cases and the three new PR cases passed.

Rationale for this change
The spec's Column Projection rules say that when a field isn't present in a data file, we resolve its value in order:
Today _get_column_projection_values only records the identity partition value when it's non-null. So when a file's partition value is null, we skip rule 1 and fall through to rule 3, as if the partition metadata had nothing to say. This was harmless before v3 because initial-default was always null anyway, but with a non-null initial-default we now read the default instead of null.
This came up while reviewing #4082
Are these changes tested?
Yes
Are there any user-facing changes?
Only for v3 tables with a non-null initial-default on an identity-partitioned column on null partition. This is a bug fix that conforms to the spec: https://iceberg.apache.org/spec/#column-projection