Semantic cache for Russian-language LLM chats. A drop-in proxy in front of any OpenAI-compatible endpoint that catches paraphrased, typo-laden, and wrong-layout repeats of earlier questions and serves the cached answer — cutting cost and latency to near-zero on the hit path.
скок стоит достаовка в спб→ HIT →Сколько стоит доставка в Санкт-Петербург?ghbdtn, rfr dthyenm njdfh→ HIT →Как вернуть товар?доставка до питера сколько→ HIT →Сколько стоит доставка в Санкт-Петербург?
The cache is built around three ideas that matter specifically for Russian customer-facing bots:
- A Russian normalizer that fixes keyboard-layout mistakes
(
ghbdtn→привет), latin transliteration (privet→привет), small typos via trigram matching, and optional lemmatization viapymorphy3. - A pluggable embedder. The default is a dependency-free hashing embedder so you can try rucache with zero model downloads. For real deployments, plug in any multilingual sentence-transformers model.
- A tiny SQLite + numpy store that rebuilds the matrix in memory on startup. Fast enough for caches up to ~100k entries, boring enough to audit in one sitting.
pip install rucache # core only
pip install rucache[embeddings] # + sentence-transformers
pip install rucache[morph] # + pymorphy3 lemmatizer
pip install rucache[openai] # + openai for embeddings/proxy
pip install rucache[all] # everythingfrom rucache import SemanticCache, HashEmbedder, RussianNormalizer
cache = SemanticCache(
db_path="cache.sqlite",
embedder=HashEmbedder(), # swap for SentenceTransformerEmbedder
normalizer=RussianNormalizer(),
threshold=0.30, # see "Thresholds" below
)
cache.put(
"Сколько стоит доставка в Санкт-Петербург?",
"Доставка в Санкт-Петербург — 350 ₽ при заказе от 2000 ₽.",
)
hit = cache.lookup("скок стоит достаовка в спб")
if hit:
print(hit.entry.answer, f"(sim={hit.similarity:.2f})")Run the full demo:
python examples/quickstart.pyStart the proxy:
export OPENAI_API_KEY=sk-...
rucache serve --port 8000Then point any OpenAI-compatible client at it:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="sk-anything")
r = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Сколько стоит доставка в СПб?"}],
)
print(r.choices[0].message.content)
# Every subsequent semantically similar question returns from the cache.Cache hits come back shaped exactly like a normal chat completion, plus a
rucache metadata block that tells you the similarity and matched question.
rucache normalize "скок стоит достаовка в спб"
# → скок стоить доставка в спб
rucache put "Сколько стоит доставка в СПб?" "350 ₽, 1–2 дня"
rucache lookup "доставка в питер сколько"
# HIT similarity=0.71 answer: 350 ₽, 1–2 дня
rucache bench --threshold 0.30
rucache statsReproducible on any machine with just numpy:
python -m rucache.cli bench --threshold 0.30
# or
PYTHONPATH=src python3 -m rucache.cli bench --threshold 0.30Dataset: 10 canonical support questions × 10 Russian paraphrases each (including typos, abbreviations, latin transliteration, and wrong-layout tokens). 20 adversarial negative probes.
| threshold | hit rate | false-positive rate | p95 lookup |
|---|---|---|---|
| 0.25 | 51% | 0% | 0.31 ms |
| 0.30 | 45% | 0% | 0.29 ms |
| 0.35 | 35% | 0% | 0.29 ms |
| 0.40 | 33% | 0% | 0.29 ms |
The hash embedder has no concept of synonymy, so it catches paraphrases by lexical overlap only. Zero false positives across the sweep — the penalty for the low ceiling is paid in missed hits, not wrong answers.
Swap the embedder:
from rucache import SemanticCache, SentenceTransformerEmbedder
cache = SemanticCache(
embedder=SentenceTransformerEmbedder(
"sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2"
),
threshold=0.85,
)On the same dataset, hit rates typically jump to 85–95% at thresholds around 0.80–0.88 while keeping false positives near zero. Numbers vary by model and domain — please run the benchmark on your data before trusting anything in this README.
Using the built-in estimator (rucache.bench.token_savings_estimate):
10,000 chats/month, 500 input + 200 output tokens average,
$0.15/$0.60 per 1k (gpt-4o-mini pricing), 40% hit rate
→ 4,000 cached responses
→ 2,000,000 input + 800,000 output tokens saved
→ ≈ $780/month saved
Think of threshold as "how similar must two questions be before we
consider them the same question?"
HashEmbedder: use 0.25–0.35. Higher is safer but drops hit rate fast because this embedder's similarity distribution is compressed.SentenceTransformerEmbedder: use 0.80–0.90. Sentence embeddings have a sharper similarity distribution and live in a different regime.OpenAIEmbedder(text-embedding-3-small): use 0.82–0.90.
When in doubt, run rucache bench with a threshold sweep on your own data
and pick the knee of the curve.
Semantic caches make mistakes that raw LLMs don't. Know them:
- Negations.
"можно ли отменить заказ?"and"можно ли НЕ отменять заказ?"can collide on a weak embedder. Use a stronger model and/or add a negation-aware rule in your normalizer. - Numbers and dates.
"доставка завтра"and"доставка через неделю"can be embedded close together. If your domain is number-sensitive, either raise the threshold or scope the cache (see below). - Personal data. The proxy scopes cache entries by system prompt and model, but not by user. If one user's answer contains their order number, don't let the cache serve it to others. Scope manually via metadata and a custom key, or disable the cache for authenticated endpoints.
- Stale answers. A cache is a liability if your underlying knowledge
changes. Use the
created_atfield in metadata to expire entries — rucache stores it for you but does not auto-evict (yet). - Tone-sensitive responses. Sentiment, emotional support, or legal advice questions should probably not be cached at all.
State lives in SQLite (cache.sqlite): durable, auditable, and diffable in
a text editor if something goes wrong. The diagram can be regenerated with
python benchmarks/make_architecture_diagram.py.
rucache/
├── src/rucache/
│ ├── normalizer.py # layout / translit / typo / lemma pipeline
│ ├── embedder.py # HashEmbedder, SentenceTransformerEmbedder, OpenAIEmbedder
│ ├── cache.py # SemanticCache core
│ ├── proxy.py # FastAPI OpenAI-compatible proxy
│ ├── cli.py # rucache CLI (typer)
│ └── bench.py # benchmark harness
├── benchmarks/data/support_ru.jsonl # 10 groups × 10 paraphrases
├── tests/ # pytest suite
├── examples/quickstart.py
└── run_tests.py # minimal pytest-free runner
- TTL / LRU eviction policies on the SQLite store
- Multi-turn dialog caching (not just the last user message)
- Streaming cache hits so
stream=trueclients also benefit - Built-in Prometheus metrics
- Benchmark on MERA / Russian SuperGLUE subsets
- HNSW index when n > 50k entries
- Optional LLM-judge-based false-positive filter
PRs welcome, especially:
- New entries for the Russian typo vocabulary in
normalizer.py - More benchmark groups (different domains: food delivery, banking, etc.)
- Stronger default embedder recommendations
- Integrations with LangChain / LlamaIndex
Run the test suite locally with:
python run_tests.py # no pytest needed
pytest # if you have itMIT — see LICENSE.
