You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Format Tables currently use the table format for every partition, so a Parquet table cannot read an ORC partition. This PR supports file.format overrides for catalog-managed partitions, resolved as partition option → table option → parquet.
The resolved format drives split planning, reads and partition statistics, including Spark ANALYZE. Older splits without a format retain the table-format fallback.
Writers use the table format. Appends reject conflicting partition formats before publishing files. Overwrites always report the actual format with replacement statistics, making it explicit even for previously inherited partitions. TRUNCATE preserves the existing format.
This extends the existing partitionOptions contract: Catalog providers must apply explicit file.format updates to existing partitions even with ignoreIfExists=true, atomically with any statistics in the same batch. Format updates preserve other options; omitting the format preserves its stored override.
Mixed-format tables require compatible readers, writers and Catalog providers. Same-partition overwrites require external serialization with other writes and format changes; queries during overwrite have no snapshot isolation.
Focused core, Spark 3 and Flink 1.20 suites passed with normal Maven checks, covering mixed-format reads, overwrite metadata updates, TRUNCATE and legacy split deserialization.
Mosaic reader tests passed with the native library available.
Manual validation (Spark 3.5.2, REST catalog)
Checks completed across downstream builds of this PR:
Reads: Full-table and partition-filtered SQL queries correctly read Parquet and custom-location ORC partitions. External files were unchanged.
Splits and readers: Verified CSV/JSON range splitting, unsplit gzip CSV, file packing, split serialization, projection and filtering against expected rows.
Statistics:NOSCAN matched file counts and bytes; ANALYZE also matched Parquet/ORC row counts and was idempotent. Custom-location ANALYZE was rejected without metadata changes.
Writes: Same-format Parquet/ORC inserts, appends, and static, dynamic and whole-table overwrites (including empty results) passed readback and statistics checks. Appends to conflicting-format or custom-location partitions were rejected before publishing files.
Metadata race: With a compatible Catalog provider, both inherited and explicit Parquet partitions passed a controlled race that changed the registered format to ORC after the overwrite's metadata lookup. The commit persisted file.format=parquet; normal SQL readback and file/statistics checks passed without metadata repair. This does not test concurrent overwrites.
Thanks for the implementation. I reviewed the latest head (c0b5133) and ran the targeted core/REST validation suite (395/395 passed). I did not find another deterministic issue that should block the merge.
A few follow-up improvements are worth considering:
P2: document the append race as well. The append path checks the partition format before publishing files, then registers the partition and reports statistics afterward. If another client changes the partition’s file.format in between, Parquet files could be left with ORC metadata. The current documentation requires external serialization for overwrites, but this requirement should also cover appends and format-metadata updates, or the catalog API should eventually provide an expected-format/version conditional update.
Consider adding the cross-format overwrite failure-injection test to the PR. It should cover a catalog-update failure and the supported recovery path (retry or metadata repair). Simply reordering metadata and file publication would only move the inconsistency to the opposite failure path.
The REST API documentation and OpenAPI description should document partitionOptions[].file.format, including its behavior with ignoreIfExists=true, the difference between append and replacement statistics, and the per-batch atomicity boundary.
It would be useful to add a test where the physical file format and filename suffix disagree—for example, an ORC file named 00000_0 or part-x.parquet.gz—to make it explicit that the reader is selected from partition metadata rather than the filename.
Regarding the stale file.format after a failed cross-format overwrite: ordinary non-ACID Hive tables have a similar non-atomic failure-recovery limitation. I would therefore treat this as a documented recovery limitation rather than a merge blocker, unless the product contract requires failed overwrites to remain automatically readable and metadata-consistent.
Thanks, addressed in 2dbde5b. Clarified append/format-update coordination, overwrite recovery, and REST/OpenAPI semantics including per-batch atomicity. Added tests for metadata repair after a failed cross-format overwrite and ORC files with missing or misleading suffixes.
Validation: all 30 FormatTablePartitionFormatTest cases passed with normal Maven checks; OpenAPI validation passed. Production behavior is unchanged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Format Tables currently use the table format for every partition, so a Parquet table cannot read an ORC partition. This PR supports
file.formatoverrides for catalog-managed partitions, resolved as partition option → table option →parquet.The resolved format drives split planning, reads and partition statistics, including Spark ANALYZE. Older splits without a format retain the table-format fallback.
Writers use the table format. Appends reject conflicting partition formats before publishing files. Overwrites always report the actual format with replacement statistics, making it explicit even for previously inherited partitions. TRUNCATE preserves the existing format.
This extends the existing
partitionOptionscontract: Catalog providers must apply explicitfile.formatupdates to existing partitions even withignoreIfExists=true, atomically with any statistics in the same batch. Format updates preserve other options; omitting the format preserves its stored override.Mixed-format tables require compatible readers, writers and Catalog providers. Same-partition overwrites require external serialization with other writes and format changes; queries during overwrite have no snapshot isolation.
Related: #8750, #9540 and #9710.
Tests
Manual validation (Spark 3.5.2, REST catalog)
Checks completed across downstream builds of this PR:
NOSCANmatched file counts and bytes;ANALYZEalso matched Parquet/ORC row counts and was idempotent. Custom-locationANALYZEwas rejected without metadata changes.file.format=parquet; normal SQL readback and file/statistics checks passed without metadata repair. This does not test concurrent overwrites.