Skip to content

fix(core): preserve partition-bucket order in scan planning - #838

Merged
JingsongLi merged 1 commit into
apache:mainfrom
JingsongLi:codex/fix-append-partition-bucket-order
Sep 15, 2026
Merged

JingsongLi merged 1 commit into
apache:mainfrom
JingsongLi:codex/fix-append-partition-bucket-order

Conversation

@JingsongLi

Copy link
Copy Markdown
Contributor

Purpose

Fix the append shard mismatch exposed by the Rust Plan CI job in apache/paimon#9825. When manifest entries interleave partitions and buckets, scan planning currently groups all buckets of a partition together. Python instead preserves the first occurrence of each (partition, bucket) pair, so the same positional shard can select different rows.

For example, groups encountered as (b, 1), (a, 0), (b, 0), (b, 1) must be emitted as (b, 1), (a, 0), (b, 0), retaining both files in the first group.

Brief change log

  • Replace nested partition/bucket maps with one insertion-ordered map keyed by (partition, bucket).
  • Preserve file order within each group and its original total-bucket metadata.
  • Correct the grouping regression and add snapshot-planning coverage for interleaved groups with and without a pushed-down limit.

Tests

  • The new snapshot-planning regression failed before the fix with the reordered groups.
  • cargo test --locked -p paimon --lib table::table_scan::tests: 101 passed, including append, primary-key, Data Evolution, row-position, range and limit tests.
  • cargo clippy --locked -p paimon --lib --tests -- -D warnings and cargo fmt --all -- --check: passed.
  • Built the Python wheel from this branch. In [python] Align native scan planning with Rust paimon#9825, 345 related Python tests passed, including the original fixed-bucket append shard failure and a new comparison of all three shards, four slices, and shard/slice combinations with limits. The new comparison forbids native fallback and reads serially so limit results are deterministic.

API and Format

No public API or storage format changes. Scan groups retain manifest encounter order across partitions; this corrects the physical row positions consumed by Python append distribution.

Documentation

No documentation changes are required for this ordering fix.

@leaves12138 leaves12138 left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM at 7fd2a47bac975c88a963347dc3da4bf73500f370.

The flat insertion-ordered map matches PyPaimon's first-occurrence ordering of (partition, bucket) pairs. It preserves every group's file order and first total_buckets value, without changing the grouping boundaries used by PK merge or Data Evolution/DV handling. I also checked that manifest loading/netting retains encounter order before this step.

Independent validation:

  • cargo test --locked -p paimon --lib table::table_scan::tests: 101 passed; cargo fmt --all -- --check passed.
  • Reproduced the interleaved-group shard mismatch with the pre-fix binding. After building the wheel from this commit, both the original fixed-bucket append regression and the new interleaved-group integration regression pass against apache/paimon#9825 at 3fc97e24.
  • Four additional layouts exercised 204 selection combinations: two partition keys, 2/5 buckets, small/large split targets, repeated commits, shard counts 1/3/7, empty shards, clipped/empty slices, and limits absent/0/2. Native fallback was forbidden; serial reads and snapshot IDs matched Python throughout.
  • Expanded related Python/native regressions: 363 passed, with 220 tracked native plans. Seven Vortex tests were skipped because the local interpreter is Python 3.10; these are targeted runs, not the complete CI suite.

No blocking issues found. Some CI jobs are still running at review time.

@JingsongLi
JingsongLi merged commit b33e034 into apache:main Sep 15, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants