Repository navigation
feat: resume a diverged hybrid prompt from a recurrent-state checkpoint - #212
Conversation
Hybrid (recurrent + attention) models could only reuse a cached prompt that was an exact prefix of the new one, so editing earlier history re-prefilled the whole conversation. Snapshot the recurrent layers at turn boundaries while prefilling (the turn after the system prompt, then every 2048+ tokens, at most 4 per entry) and, when a prompt diverges, resume from the newest checkpoint before the divergence: recurrent layers are set from it, attention layers are trimmed back to it (plain KV always; a sliding-window ring only until it fills), and only the rest is prefilled. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
solderzzc
left a comment
There was a problem hiding this comment.
Thanks for this — the design is careful in a few places I want to name before I get to my own testing.
Two things stood out. First, the failure mode is biased toward a slowdown rather than a wrong answer: every gate I looked at (canRewindAttention, the offset < maxSize read from the saved metaState, the fallback when no checkpoint is inside the shared prefix) returns to the exact-prefix path instead of guessing. Second, the checkpoint positions are expressed in the same token-index space the existing hybridCacheBoundary already uses, so PromptCache stays a pure actor with no new I/O dependency. The test list also covers the cases I would have asked for — out-of-order checkpoints, wrapped vs unwrapped rings, and carry-over across both an extension and a checkpoint restore.
I ran my own verification against the PR head. Nothing below is a request for changes.
Build and tests. Sources/SwiftLM/Server.swift at 4fe232f compiles against main (8469a17) with the submodule pins in the PR, and swift test --filter SwiftLMTests gives 276 tests / 0 failures, matching your number. Your 13 new tests all pass.
End-to-end. I exercised the path on a real hybrid model (mlx-community/Qwen3.5-0.8B-4bit, greedy, debug build) rather than only the unit level. A 4-turn conversation (~1500 tokens per turn, 6273 total) with the first words of the last turn edited:
- cold: 0.68 s prefill (9,217 t/s)
- diverged: 0.34 s, log shows
HIT (hybrid): 3148/6273 tokens reused (checkpoint of 6249) - identical repeat: 0.02 s, exact-prefix hit at 6266/6273
So the restore fires and roughly halves prefill here.
Equivalence. I ran the diverged prompt twice: once on a server warmed with the unedited prompt (so it resumed from the checkpoint), and once on a freshly restarted server (full prefill, no cache hit in the log). Greedy output was byte-identical between the two. That is only text on one model and one prompt shape, and there are no logprobs, so it is the same class of evidence as your output check — but it does corroborate that the rewind is not visibly different here.
Two observations, offered as data rather than objections:
- Reuse is quite sensitive to conversation shape. In a shorter conversation (system prompt plus two user turns) the newest checkpoint landing before the divergence was only 716 tokens back, so reuse was 716/15266 (~5%), versus ~50% in the 4-turn case. Your ~70% figure looks reachable, but a reader could reasonably expect that as a default rather than as a function of turn spacing.
- I could not exercise the
--ctx-sizepath — theRotatingKVCachebranch ofcanRewindAttentionis the subtle part and the one I only read, not ran.
Thanks again — happy to see this land.
|
Thanks for the review and the extra CI coverage in #213. On your second observation (the What ran. Qwen3.5-35B-A3B 4-bit (hybrid MoE), 16 GB M2,
The attention layers were What it did not cover.
|
|
Thank you for running the AI-assisted analysis (Claude Code). |
Problem
Hybrid models (recurrent GatedDeltaNet/Mamba layers plus attention, e.g. Qwen3.5/3.6) can only reuse a cached prompt that is an exact prefix of the new one, because recurrent state can't be rewound. Any edit to earlier history (a rewritten or dropped message, a compacted tool result, a regenerated turn) makes the whole conversation re-prefill from zero, even when 80% of it is unchanged. At ~11k tokens that is ~165 s on a 16 GB M2.
Change
While prefilling a hybrid prompt, snapshot the recurrent state at chosen turn boundaries ("checkpoints"). A later prompt that diverges from a cached one resumes from the newest checkpoint at or before the divergence: recurrent layers are set from the checkpoint, attention layers are trimmed back to it, and only the rest is prefilled.
--ctx-size) only until it fills (offset < maxSize, theRotatingKVCache.isTrimmable(after:)predicate, read from the saved metaState before the live cache is touched). Otherwise the exact-prefix path is used as before.Prefill is split into one
TokenIteratorper segment between checkpoints (the same call the single prefill already makes). Different chunk boundaries change GDN numerics slightly versus a single pass; the state is fp32 end to end, and cold prefill time did not change measurably (below).Measured
16 GB M2 Air, Qwen3.5-35B-A3B MoE 4-bit,
--no-vision --stream-experts --ssd-prefetch --prefill-size 2048. A system prompt plus three ~3.8k-token user messages with short answers between (10.9k tokens), then the same history with the third user message edited, then the same with the system prompt edited (no checkpoint can help). Baseline ismainwithout this change, same prompts.mainPeak swap stayed at 7–11 GB in both.
With
--ctx-size 32768(attention layers areRotatingKVCache, not yet wrapped at this length) the same three requests take 142.8 s, 47.7 s (7,669 of 10,834 tokens reused from a checkpoint) and 154.5 s.Output check. Resuming from a checkpoint should give the same answer as prefilling the whole prompt. With the edited 10.8k-token conversation above, greedy (
temperature: 0) decoding gave byte-identical 80-token replies for (a) a checkpoint resume (7,645 tokens reused, 58 s) and (b) two separate fresh-server full prefills (157 s and 163 s). This is one prompt and text comparison only (the server has no logprobs), so it shows the segmented prefill is not visibly different here, not that it is bit-exact.Tests
HybridCacheCheckpointTests(13 tests): checkpoint positions for cold and resumed prefills, thinning and de-duplication, resume from the newest checkpoint before the edit, a checkpoint after the divergence is not used, exact prefix still preferred, the entry with the longest resume point wins across entries, a wrapped ring refuses and an unwrapped ring accepts a checkpoint restore, and checkpoints carrying across an extension and across a checkpoint restore.swift test --filter SwiftLMTests: 276 tests, 0 failures.Builds on #210 (multi-entry cache, merged); useful with
--prompt-cache-entries 1too.🤖 Generated with Claude Code