Repository navigation
feat: keep several prompt cache entries (--prompt-cache-entries) and evict under memory pressure - #210
Merged
Conversation
…evict them under memory pressure PromptCache held a single entry, so a side request between two turns of the same conversation evicted that conversation's prefix. It is now an LRU of up to --prompt-cache-entries (default 1, i.e. unchanged). Saving an extension of an entry replaces it, restore picks the longest usable match, and a memory pressure source drops entries (all but the newest on warning, all on critical). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
CodeAndCanvas728
force-pushed
the
pr/prompt-cache-lru
branch
from
October 5, 2026 14:59
8703660 to
faa6672
Compare
…ry on eviction, document and test the flag Follow-up to the multi-entry prompt cache (--prompt-cache-entries). Stale entries. save() only replaced an entry the new prompt extends exactly, but the prompt saved for turn N is often not an exact prefix of turn N+1: templates that re-render history drop or rewrite the generation prompt it ended with, and edits or regenerations change the last message. The stale entry then held a slot until it aged out, pushing another conversation out of the cache. save() now also drops an entry that diverges from the new prompt in at most PromptCache.supersedeSlack (16) trailing tokens and shares more than it loses. The rule is stateless, so it needs no conversation identity. For a cache that can be rewound to any shorter prefix replacing is bounded: any prompt that could have used the old entry loses at most 16 tokens of reuse. Entries that merely start alike (a shared BOS or system prompt followed by more than 16 tokens of their own) are kept apart. Two kinds of state cannot be rewound that way and keep the strict exact-prefix rule: recurrent (hybrid) snapshots, which are only valid at exactly tokens.count, and sliding-window ring buffers that have already dropped tokens (offset beyond the window), which usableMatch rejects for any rewind past one token. Near-prefix replacement there would turn the old entry's full hit into a miss (two conversations sharing a long system prompt, each with a short first message). save() asks the same question usableMatch does, through one shared helper, so the two cannot drift apart. A ring that has not wrapped yet is rewound like any attention cache and gets the near-prefix rule. An empty prompt is no longer saved: nothing can restore it and it only took a slot. Memory pressure. evict(keepMostRecent:) now returns how many entries it dropped. The handler logs only when something was dropped and calls Memory.clearCache() then, from the Task that awaits the eviction rather than from the GCD handler thread, so the freed KV buffers actually go back to the system instead of staying in the MLX allocator pool. The comment and the flag help now state the real behavior: a warning keeps the most recent entry, a critical event drops all of them, so with the default of one entry only a critical event empties the cache. Flag. Document --prompt-cache-entries in the README option table and print the effective value in the startup Config line. A value below 1 prints a warning at startup and is raised to 1 (as --gpu-layers does for an invalid value), so the clamp is visible before any model is loaded. Tests. Cover the supersede rule (within the slack, one token beyond it, tiny entries, shared system prompt, ping-pong between two conversations, several dominated entries at once, exact duplicates, empty prompts, hybrid snapshots), wrapped and unwrapped ring buffers (shared system prompt with short tails, mixed full-attention and sliding-window layers, rewind boundaries), the longest-usable-match pick regardless of recency for restore() and restoreExactPrefix(), rejection of an unusable candidate falling through to another entry, tail divergence, hybrid recency refresh, evict() return values, the capacity floor and the flag parse. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Member
|
Thanks for this. I reviewed it and pushed one follow-up commit (4588b8b) on top of your branch rather than asking for another round trip. AI-assisted (Claude Code); tests run locally on origin/main + this branch (SwiftLMTests 262/0).
Not changed: no byte budget, a short prompt within 16 tokens of a longer one replaces it, and no end-to-end run on a sliding-window model. The PR description is out of date ("unchanged", "7 tests") and could use a refresh before merge. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
PromptCacheholds a single entry. Any request that doesn't share the cached prefix replaces it, so a client that interleaves conversations loses its cache on every switch. Typical cases: an agent that fires a side request (title generation, a summariser, a sub-agent with a different prompt) between two turns of the main conversation, or several chat sessions against one server. Each switch costs a full re-prefill of the long conversation.Change
PromptCachebecomes an LRU of up to--prompt-cache-entries Nentries (default 1, so default behaviour and memory use are unchanged).savereplaces any entry that is a prefix of the new one, so a single growing conversation still occupies one slot.restore(trimmable models) picks the entry with the longest usable match; the per-entry safety checks (excess vs. cached length, wrapped ring buffers) are unchanged, just factored intousableMatch.restoreExactPrefix(hybrid recurrent+attention models) picks the longest exact-prefix entry.DispatchSourcememory-pressure source evicts entries: all but the most recent on.warning, all on.critical. Cached KV is the largest evictable allocation, so it is returned before the OS has to compress or kill.PromptCache.evict(keepMostRecent:)is also the missing way to drop entries while idle.No byte budget in this PR; that can follow if pressure-based eviction proves too late in practice.
Tests
PromptCacheLRUTests(7 tests): unrelated prompts both hit with 2 entries and not with 1; a hit refreshes recency and the least recent is dropped; an extension replaces its prefix entry; longest match wins; hybrid exact-prefix picks the longest; bothevictmodes. Existing cache tests pass unchanged.swift test --filter SwiftLMTests: 230 pass.Measurements
Qwen3.5 35B-A3B (
qwen3_5_moe, hybrid GatedDelta + attention, 4-bit), loaded text-only (--no-vision) with--stream-experts --ssd-prefetch --prefill-size 2048, on a 16 GB M2. Sequence: turn 1 (12.7k tokens), an unrelated 2.7k-token request, then turn 2 (14.7k tokens, extending turn 1). Greedy,max_tokens16.--prompt-cache-entries 4Memory demand stayed about 8.4–8.7 GB in both runs. The pressure handler fired once (warning level) during the 4-entry run under swap on this machine and kept the newest entry.
Note
On hybrid models,
--ctx-sizecurrently disables the hybrid cache (attention layers becomeRotatingKVCache, whichhybridCacheBoundaryrejects); that is fixed separately in #207. The runs above did not use--ctx-size.🤖 Generated with Claude Code