Repository navigation
fix(python): rebuild the fallback runtime after fork - #832
Conversation
leaves12138
left a comment
There was a problem hiding this comment.
Reviewed 386ef9345021064c2a91aa36d3b4a9e28a6cb510. No blocking findings.
The PID check and publication of a fresh, uninitialized state let the child bypass both an inherited runtime and an inherited in-progress OnceLock. The acquire/release ordering and the ownership of unsuccessful CAS allocations look correct; published parent state remains allocated, avoiding shutdown of inherited worker state. The existing entered-runtime behavior is preserved.
Local validation:
- Compiled the exact
runtime.rsstandalone against cached Tokio dependencies: both unit tests passed. - An independent real-fork harness compiled against the base runtime hit the child's five-second deadline; the same harness passed against this head. It exercises child/grandchild processes, concurrent runtime reuse, TCP I/O, timers, the blocking pool, and continued parent use.
- A deterministic fork while another thread holds the published state's initialization
OnceLockalso passed against this head. cargo fmt --all -- --checkand syntax parsing of the new Python test passed.
I reviewed the two Python fork regressions but did not rebuild the Python binding or execute the full Python/DataFusion integration suites locally.
Update rustls to 0.23.45 and its required AWS-LC/webpki dependencies. Regenerate workspace dependency reports without suppressing the advisory.
|
Fixed the failing CI dependency-policy check in
Local validation passed: advisories/licenses against a freshly fetched RustSec database, dependency-report verification, all 21 generated release legal files, 12 release-tool tests, compilation of rustls 0.23.45, formatting, and staged-diff checks. The legal files were regenerated but needed no content changes. The fix is appended to the existing PR branch without rewriting history. CI run |
Purpose
Fix native planning hanging in forked Python workers after the parent has used a Paimon catalog or scan. A Torch DataLoader with
num_workers=2hits this in apache/paimon#9825: the child inherits the global Tokio runtime but none of its worker threads, so catalog I/O and planning never finish.Brief change log
Tests
cargo clippy --offline -p paimon-datafusion --all-targets -- -D warnings,cargo fmt --all -- --check, Python test Flake8, andgit diff --checkpassed.API and Format
No public API or storage-format changes. This fixes the fallback runtime used by catalog and planning calls; explicitly entered Tokio runtimes retain their existing behavior.
Documentation
Updated runtime comments to explain process ownership and why inherited state is retained.