blockchain: Early data commitment checks. - #3793
Merged
davecgh merged 3 commits intoSep 16, 2026
Merged
Conversation
jholdstock
reviewed
Sep 15, 2026
jholdstock
left a comment
Member
There was a problem hiding this comment.
Looks good other than that one comment.
davecgh
force-pushed
the
blockchain_early_data_commitment_check
branch
from
September 15, 2026 15:34
6ae456e to
4485182
Compare
davecgh
force-pushed
the
blockchain_early_data_commitment_check
branch
from
September 15, 2026 23:43
4485182 to
c75628b
Compare
matthawkins90
approved these changes
Sep 16, 2026
davecgh
force-pushed
the
blockchain_early_data_commitment_check
branch
from
September 16, 2026 07:27
c75628b to
4c3d9bf
Compare
dcr-approvals-proxy
approved these changes
Sep 16, 2026
This separates the context free block data sanity checks from the header sanity checks and modifies the block processing order semantics slightly to fully validate and index a previously unknown header before performing checks on the associated block data. Previously, the block sanity checks were performed before positional header validation and insertion of the header into the block index. This ordering coupled the header validation with block data validation and required callers to track whether header sanity checks had already been performed. The new ordering gives header processing consistent semantics for both headers-first and full block processing. Namely, previously unknown headers are always subject to all context-free and positional header checks before being added to the index. Block data validation is then performed separately. The primary motivation is to pave the way for reordering checks so that attackers must perform more work before reaching expensive validation while avoiding what would otherwise be circular validation dependencies. However, the change is also useful on its own because it makes the ordering more consistent and easier to reason about. It also removes the need for additional flags specifying whether or not the header sanity checks have already been done.
This makes use of the recently-added historical positional agenda status determination to move the checks that ensure a given header commits to its associated block data very early in the block processing path. It first splits the main merkle root checks into a separate function that accepts the merkle root algorithm as a new enum value. This approach allows the function to be called from anywhere since it becomes the responsibility of the caller to determine which algorithm to apply. It also provides a straightforward path for adding a new merkle root algorithm in the future if desired. Next, it introduces a new block data precondition function that checks the necessary merkle root algorithms as one of the very first checks on block data. Since the new function runs before the full context needed to determine the state of a vote is necessarily available, it uses the historical positional agenda status determination when possible to reduce work. Otherwise, it must allow all possible algorithms and verify that the data matches the applicable algorithm later, once the agenda status can be definitively determined. Further, it returns a flag that specifies whether or not the data commitment is definitively known to be valid. A non-definitive agenda status cannot establish which merkle root algorithm is actually applicable. Then, for cases where the agenda status can't be determined definitively in the earlier positional path, a new contextual merkle roots check function is introduced to take the place of the existing contextual check. This check is also reordered to execute just after the contextual header checks in the contextual block checks. In practice, because the main and version 3 test networks now have historical agenda activation data specified, the correct algorithm will always be known early for those networks. In other words, invalid data will always be rejected early, and data that passes the early checks will be definitively known to have a valid data commitment.
This introduces a small LRU cache for tracking recent blocks that have definitively proven the header commits to the received data for the block. The cache is used to avoid recomputing and rechecking the merkle roots later in the validation path when they have already been definitively proven earlier. The approach mirrors the existing approach for the cache that performs the same function for recent blocks that have passed all contextual checks.
davecgh
force-pushed
the
blockchain_early_data_commitment_check
branch
from
September 16, 2026 07:32
4c3d9bf to
b56e78f
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This requires #3791.This consists of a few commits that rework the early block process logic as follows:
See the individual commit messages for more details.
Previously, the block sanity checks were performed before positional header validation and insertion of the header into the block index. This ordering coupled the header validation with block data validation and required callers to track whether header sanity checks had already been performed.
The new ordering gives header processing consistent semantics for both headers-first and full block processing. Namely, previously unknown headers are always subject to all context-free and positional header checks before being added to the index. Block data validation is then performed separately.
It also introduces a new block data precondition function that checks the necessary merkle root algorithms as one of the very first checks on block data. Since the new function runs before the full context needed to determine the state of a vote is necessarily available, it uses the historical positional agenda status determination when possible to reduce work. Otherwise, it must allow all possible algorithms and verify that the data matches the applicable algorithm later, once the agenda status can be definitively determined. Further, it returns a flag that specifies whether or not the data commitment is definitively known to be valid. A non-definitive agenda status cannot establish which merkle root algorithm is actually applicable.
Then, for cases where the agenda status can't be determined definitively in the earlier positional path, a new contextual merkle roots check function is introduced to take the place of the existing contextual check. This check is also reordered to execute just after the contextual header checks in the contextual block checks.
In practice, because the main and version 3 test networks now have historical agenda activation data specified, the correct algorithm will always be known early for those networks. In other words, invalid data will always be rejected early, and data that passes the early checks will be definitively known to have a valid data commitment.