Skip to content

[core] Support partition-level file formats for catalog-managed Format Tables - #10303

Open
tonymtu wants to merge 5 commits into
apache:masterfrom
tonymtu:feat/format-table-partition-file-format
Open

tonymtu wants to merge 5 commits into
apache:masterfrom
tonymtu:feat/format-table-partition-file-format

Conversation

@tonymtu

@tonymtu tonymtu commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Format Tables currently use the table format for every partition, so a Parquet table cannot read an ORC partition. This PR supports file.format overrides for catalog-managed partitions, resolved as partition option → table option → parquet.

The resolved format drives split planning, reads and partition statistics, including Spark ANALYZE. Older splits without a format retain the table-format fallback.

Writers use the table format. Appends reject conflicting partition formats before publishing files. Overwrites always report the actual format with replacement statistics, making it explicit even for previously inherited partitions. TRUNCATE preserves the existing format.

This extends the existing partitionOptions contract: Catalog providers must apply explicit file.format updates to existing partitions even with ignoreIfExists=true, atomically with any statistics in the same batch. Format updates preserve other options; omitting the format preserves its stored override.

Mixed-format tables require compatible readers, writers and Catalog providers. Same-partition overwrites require external serialization with other writes and format changes; queries during overwrite have no snapshot isolation.

Related: #8750, #9540 and #9710.

Tests

  • Focused core, Spark 3 and Flink 1.20 suites passed with normal Maven checks, covering mixed-format reads, overwrite metadata updates, TRUNCATE and legacy split deserialization.
  • Mosaic reader tests passed with the native library available.
Manual validation (Spark 3.5.2, REST catalog)

Checks completed across downstream builds of this PR:

  • Reads: Full-table and partition-filtered SQL queries correctly read Parquet and custom-location ORC partitions. External files were unchanged.
  • Splits and readers: Verified CSV/JSON range splitting, unsplit gzip CSV, file packing, split serialization, projection and filtering against expected rows.
  • Statistics: NOSCAN matched file counts and bytes; ANALYZE also matched Parquet/ORC row counts and was idempotent. Custom-location ANALYZE was rejected without metadata changes.
  • Writes: Same-format Parquet/ORC inserts, appends, and static, dynamic and whole-table overwrites (including empty results) passed readback and statistics checks. Appends to conflicting-format or custom-location partitions were rejected before publishing files.
  • Metadata race: With a compatible Catalog provider, both inherited and explicit Parquet partitions passed a controlled race that changed the registered format to ORC after the overwrite's metadata lookup. The commit persisted file.format=parquet; normal SQL readback and file/statistics checks passed without metadata repair. This does not test concurrent overwrites.

@JingsongLi

Copy link
Copy Markdown
Contributor

cc @sundapeng to take a review

@sundapeng

Copy link
Copy Markdown
Member

Thanks for the implementation. I reviewed the latest head (c0b5133) and ran the targeted core/REST validation suite (395/395 passed). I did not find another deterministic issue that should block the merge.

A few follow-up improvements are worth considering:

  1. P2: document the append race as well. The append path checks the partition format before publishing files, then registers the partition and reports statistics afterward. If another client changes the partition’s file.format in between, Parquet files could be left with ORC metadata. The current documentation requires external serialization for overwrites, but this requirement should also cover appends and format-metadata updates, or the catalog API should eventually provide an expected-format/version conditional update.

  2. Consider adding the cross-format overwrite failure-injection test to the PR. It should cover a catalog-update failure and the supported recovery path (retry or metadata repair). Simply reordering metadata and file publication would only move the inconsistency to the opposite failure path.

  3. The REST API documentation and OpenAPI description should document partitionOptions[].file.format, including its behavior with ignoreIfExists=true, the difference between append and replacement statistics, and the per-batch atomicity boundary.

  4. It would be useful to add a test where the physical file format and filename suffix disagree—for example, an ORC file named 00000_0 or part-x.parquet.gz—to make it explicit that the reader is selected from partition metadata rather than the filename.

Regarding the stale file.format after a failed cross-format overwrite: ordinary non-ACID Hive tables have a similar non-atomic failure-recovery limitation. I would therefore treat this as a documented recovery limitation rather than a merge blocker, unless the product contract requires failed overwrites to remain automatically readable and metadata-consistent.

@tonymtu

tonymtu commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

Thanks, addressed in 2dbde5b. Clarified append/format-update coordination, overwrite recovery, and REST/OpenAPI semantics including per-batch atomicity. Added tests for metadata repair after a failed cross-format overwrite and ORC files with missing or misleading suffixes.

Validation: all 30 FormatTablePartitionFormatTest cases passed with normal Maven checks; OpenAPI validation passed. Production behavior is unchanged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants