Repository navigation
fix(core): preserve partition-bucket order in scan planning - #838
Merged
JingsongLi merged 1 commit intoSep 15, 2026
Merged
Conversation
leaves12138
approved these changes
Sep 15, 2026
leaves12138
left a comment
There was a problem hiding this comment.
LGTM at 7fd2a47bac975c88a963347dc3da4bf73500f370.
The flat insertion-ordered map matches PyPaimon's first-occurrence ordering of (partition, bucket) pairs. It preserves every group's file order and first total_buckets value, without changing the grouping boundaries used by PK merge or Data Evolution/DV handling. I also checked that manifest loading/netting retains encounter order before this step.
Independent validation:
cargo test --locked -p paimon --lib table::table_scan::tests: 101 passed;cargo fmt --all -- --checkpassed.- Reproduced the interleaved-group shard mismatch with the pre-fix binding. After building the wheel from this commit, both the original fixed-bucket append regression and the new interleaved-group integration regression pass against apache/paimon#9825 at
3fc97e24. - Four additional layouts exercised 204 selection combinations: two partition keys, 2/5 buckets, small/large split targets, repeated commits, shard counts 1/3/7, empty shards, clipped/empty slices, and limits absent/0/2. Native fallback was forbidden; serial reads and snapshot IDs matched Python throughout.
- Expanded related Python/native regressions: 363 passed, with 220 tracked native plans. Seven Vortex tests were skipped because the local interpreter is Python 3.10; these are targeted runs, not the complete CI suite.
No blocking issues found. Some CI jobs are still running at review time.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Fix the append shard mismatch exposed by the Rust Plan CI job in apache/paimon#9825. When manifest entries interleave partitions and buckets, scan planning currently groups all buckets of a partition together. Python instead preserves the first occurrence of each
(partition, bucket)pair, so the same positional shard can select different rows.For example, groups encountered as
(b, 1), (a, 0), (b, 0), (b, 1)must be emitted as(b, 1), (a, 0), (b, 0), retaining both files in the first group.Brief change log
(partition, bucket).Tests
cargo test --locked -p paimon --lib table::table_scan::tests: 101 passed, including append, primary-key, Data Evolution, row-position, range and limit tests.cargo clippy --locked -p paimon --lib --tests -- -D warningsandcargo fmt --all -- --check: passed.API and Format
No public API or storage format changes. Scan groups retain manifest encounter order across partitions; this corrects the physical row positions consumed by Python append distribution.
Documentation
No documentation changes are required for this ordering fix.