Skip to content

feat: resume a diverged hybrid prompt from a recurrent-state checkpoint - #212

Merged
solderzzc merged 1 commit into
SharpAI:mainfrom
CodeAndCanvas728:pr/hybrid-cache-checkpoints
Oct 5, 2026
Merged

solderzzc merged 1 commit into
SharpAI:mainfrom
CodeAndCanvas728:pr/hybrid-cache-checkpoints

Conversation

@CodeAndCanvas728

Copy link
Copy Markdown
Contributor

Problem

Hybrid models (recurrent GatedDeltaNet/Mamba layers plus attention, e.g. Qwen3.5/3.6) can only reuse a cached prompt that is an exact prefix of the new one, because recurrent state can't be rewound. Any edit to earlier history (a rewritten or dropped message, a compacted tool result, a regenerated turn) makes the whole conversation re-prefill from zero, even when 80% of it is unchanged. At ~11k tokens that is ~165 s on a 16 GB M2.

Change

While prefilling a hybrid prompt, snapshot the recurrent state at chosen turn boundaries ("checkpoints"). A later prompt that diverges from a cached one resumes from the newest checkpoint at or before the divergence: recurrent layers are set from the checkpoint, attention layers are trimmed back to it, and only the rest is prefilled.

  • Where checkpoints go. The first turn after the system prompt (the anchor, which survives any later edit), then each turn start at least 2048 tokens after the previous checkpoint. The point a restore resumed from counts as one.
  • How many. At most 4 per entry. Thinning keeps the anchor and the newest and drops the checkpoint closest to a neighbour.
  • Cost. One recurrent state per layer per checkpoint (a few tens of MB on a 35B-A3B). The attention KV is not duplicated: it is the entry's own, trimmed.
  • Carry-over. An entry saved after a resume inherits the checkpoints up to the resume point, plus the end state of the entry it extended, so they survive a linear conversation.
  • Exactness gate. Attention layers must rewind exactly: plain KV caches always can; a sliding-window ring (--ctx-size) only until it fills (offset < maxSize, the RotatingKVCache.isTrimmable(after:) predicate, read from the saved metaState before the live cache is touched). Otherwise the exact-prefix path is used as before.
  • Unchanged. Exact-prefix hits still win when available. Non-hybrid models, MTP, draft-model and TurboKV paths are untouched.

Prefill is split into one TokenIterator per segment between checkpoints (the same call the single prefill already makes). Different chunk boundaries change GDN numerics slightly versus a single pass; the state is fp32 end to end, and cold prefill time did not change measurably (below).

Measured

16 GB M2 Air, Qwen3.5-35B-A3B MoE 4-bit, --no-vision --stream-experts --ssd-prefetch --prefill-size 2048. A system prompt plus three ~3.8k-token user messages with short answers between (10.9k tokens), then the same history with the third user message edited, then the same with the system prompt edited (no checkpoint can help). Baseline is main without this change, same prompts.

Request main this PR
1. cold 10.9k 151.0 s 150.2 s
2. third user message edited 162.6 s (full miss) 51.1 s (7,674 of 10,839 tokens reused)
3. system prompt edited (miss either way) 171.1 s 172.2 s

Peak swap stayed at 7–11 GB in both.

With --ctx-size 32768 (attention layers are RotatingKVCache, not yet wrapped at this length) the same three requests take 142.8 s, 47.7 s (7,669 of 10,834 tokens reused from a checkpoint) and 154.5 s.

Output check. Resuming from a checkpoint should give the same answer as prefilling the whole prompt. With the edited 10.8k-token conversation above, greedy (temperature: 0) decoding gave byte-identical 80-token replies for (a) a checkpoint resume (7,645 tokens reused, 58 s) and (b) two separate fresh-server full prefills (157 s and 163 s). This is one prompt and text comparison only (the server has no logprobs), so it shows the segmented prefill is not visibly different here, not that it is bit-exact.

Tests

HybridCacheCheckpointTests (13 tests): checkpoint positions for cold and resumed prefills, thinning and de-duplication, resume from the newest checkpoint before the edit, a checkpoint after the divergence is not used, exact prefix still preferred, the entry with the longest resume point wins across entries, a wrapped ring refuses and an unwrapped ring accepts a checkpoint restore, and checkpoints carrying across an extension and across a checkpoint restore. swift test --filter SwiftLMTests: 276 tests, 0 failures.

Builds on #210 (multi-entry cache, merged); useful with --prompt-cache-entries 1 too.

🤖 Generated with Claude Code

Hybrid (recurrent + attention) models could only reuse a cached prompt that
was an exact prefix of the new one, so editing earlier history re-prefilled
the whole conversation. Snapshot the recurrent layers at turn boundaries
while prefilling (the turn after the system prompt, then every 2048+ tokens,
at most 4 per entry) and, when a prompt diverges, resume from the newest
checkpoint before the divergence: recurrent layers are set from it, attention
layers are trimmed back to it (plain KV always; a sliding-window ring only
until it fills), and only the rest is prefilled.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

@solderzzc solderzzc left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this — the design is careful in a few places I want to name before I get to my own testing.

Two things stood out. First, the failure mode is biased toward a slowdown rather than a wrong answer: every gate I looked at (canRewindAttention, the offset < maxSize read from the saved metaState, the fallback when no checkpoint is inside the shared prefix) returns to the exact-prefix path instead of guessing. Second, the checkpoint positions are expressed in the same token-index space the existing hybridCacheBoundary already uses, so PromptCache stays a pure actor with no new I/O dependency. The test list also covers the cases I would have asked for — out-of-order checkpoints, wrapped vs unwrapped rings, and carry-over across both an extension and a checkpoint restore.

I ran my own verification against the PR head. Nothing below is a request for changes.

Build and tests. Sources/SwiftLM/Server.swift at 4fe232f compiles against main (8469a17) with the submodule pins in the PR, and swift test --filter SwiftLMTests gives 276 tests / 0 failures, matching your number. Your 13 new tests all pass.

End-to-end. I exercised the path on a real hybrid model (mlx-community/Qwen3.5-0.8B-4bit, greedy, debug build) rather than only the unit level. A 4-turn conversation (~1500 tokens per turn, 6273 total) with the first words of the last turn edited:

  • cold: 0.68 s prefill (9,217 t/s)
  • diverged: 0.34 s, log shows HIT (hybrid): 3148/6273 tokens reused (checkpoint of 6249)
  • identical repeat: 0.02 s, exact-prefix hit at 6266/6273

So the restore fires and roughly halves prefill here.

Equivalence. I ran the diverged prompt twice: once on a server warmed with the unedited prompt (so it resumed from the checkpoint), and once on a freshly restarted server (full prefill, no cache hit in the log). Greedy output was byte-identical between the two. That is only text on one model and one prompt shape, and there are no logprobs, so it is the same class of evidence as your output check — but it does corroborate that the rewind is not visibly different here.

Two observations, offered as data rather than objections:

  1. Reuse is quite sensitive to conversation shape. In a shorter conversation (system prompt plus two user turns) the newest checkpoint landing before the divergence was only 716 tokens back, so reuse was 716/15266 (~5%), versus ~50% in the 4-turn case. Your ~70% figure looks reachable, but a reader could reasonably expect that as a default rather than as a function of turn spacing.
  2. I could not exercise the --ctx-size path — the RotatingKVCache branch of canRewindAttention is the subtle part and the one I only read, not ran.

Thanks again — happy to see this land.

@solderzzc
solderzzc merged commit 5eb2ba0 into SharpAI:main Oct 5, 2026
14 checks passed
@CodeAndCanvas728

Copy link
Copy Markdown
Contributor Author

Thanks for the review and the extra CI coverage in #213. On your second observation (the --ctx-size / RotatingKVCache branch of canRewindAttention was only read, not run): I ran it live, so here is what that did and did not cover.

What ran. Qwen3.5-35B-A3B 4-bit (hybrid MoE), 16 GB M2, --no-vision --stream-experts --ssd-prefetch --prefill-size 2048 --ctx-size 32768, the same ~10.8k-token conversation as in the description (system prompt + three ~3.8k-token user turns):

Request Time
1. cold 142.8 s
2. third user message edited 47.7 s, HIT (hybrid): 7669/10834 tokens reused (checkpoint of 10910)
3. system prompt edited (no checkpoint can help) 154.5 s

The attention layers were RotatingKVCache in that run, so the checkpoint restore plus trim on a ring buffer worked end to end.

What it did not cover.

  • The ring never wrapped (10.9k tokens against a 32,768 window), so this is the unwrapped side of the offset < maxSize gate. The wrapped side (refuse and fall back to the exact-prefix path) is covered only by the unit test testWrappedRingRefusesCheckpointRestore. In practice that means that with a --ctx-size smaller than the conversation, an edit gets no checkpoint reuse.
  • The byte-for-byte output comparison against a fresh full prefill (80 greedy tokens, identical) was done on the config without --ctx-size. I have not repeated it with --ctx-size.

@solderzzc

Copy link
Copy Markdown
Member

Thank you for running the --ctx-size / RotatingKVCache path live, and for being so clear about what it did and did not cover (the unwrapped ring end to end; the wrapped side only by testWrappedRingRefusesCheckpointRestore). An edited prompt dropping from 142.8 s to 47.7 s with 7669/10834 tokens reused on a 16 GB machine is a great result for the checkpoint-resume work.

AI-assisted analysis (Claude Code).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants