Skip to content
flyhigh-hifiPublic

About

Semantic cache for Russian-language LLM chats

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

rucache

Semantic cache for Russian-language LLM chats. A drop-in proxy in front of any OpenAI-compatible endpoint that catches paraphrased, typo-laden, and wrong-layout repeats of earlier questions and serves the cached answer — cutting cost and latency to near-zero on the hit path.

скок стоит достаовка в спб → HIT → Сколько стоит доставка в Санкт-Петербург? ghbdtn, rfr dthyenm njdfh → HIT → Как вернуть товар? доставка до питера сколько → HIT → Сколько стоит доставка в Санкт-Петербург?

The cache is built around three ideas that matter specifically for Russian customer-facing bots:

  1. A Russian normalizer that fixes keyboard-layout mistakes (ghbdtn → привет), latin transliteration (privet → привет), small typos via trigram matching, and optional lemmatization via pymorphy3.
  2. A pluggable embedder. The default is a dependency-free hashing embedder so you can try rucache with zero model downloads. For real deployments, plug in any multilingual sentence-transformers model.
  3. A tiny SQLite + numpy store that rebuilds the matrix in memory on startup. Fast enough for caches up to ~100k entries, boring enough to audit in one sitting.

Install

pip install rucache                    # core only
pip install rucache[embeddings]        # + sentence-transformers
pip install rucache[morph]             # + pymorphy3 lemmatizer
pip install rucache[openai]            # + openai for embeddings/proxy
pip install rucache[all]               # everything

30-second quickstart

from rucache import SemanticCache, HashEmbedder, RussianNormalizer

cache = SemanticCache(
    db_path="cache.sqlite",
    embedder=HashEmbedder(),            # swap for SentenceTransformerEmbedder
    normalizer=RussianNormalizer(),
    threshold=0.30,                     # see "Thresholds" below
)

cache.put(
    "Сколько стоит доставка в Санкт-Петербург?",
    "Доставка в Санкт-Петербург — 350 ₽ при заказе от 2000 ₽.",
)

hit = cache.lookup("скок стоит достаовка в спб")
if hit:
    print(hit.entry.answer, f"(sim={hit.similarity:.2f})")

Run the full demo:

python examples/quickstart.py

Transparent proxy in front of OpenAI

Start the proxy:

export OPENAI_API_KEY=sk-...
rucache serve --port 8000

Then point any OpenAI-compatible client at it:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="sk-anything")
r = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Сколько стоит доставка в СПб?"}],
)
print(r.choices[0].message.content)
# Every subsequent semantically similar question returns from the cache.

Cache hits come back shaped exactly like a normal chat completion, plus a rucache metadata block that tells you the similarity and matched question.

CLI

rucache normalize "скок стоит достаовка в спб"
#  → скок стоить доставка в спб

rucache put "Сколько стоит доставка в СПб?" "350 ₽, 1–2 дня"
rucache lookup "доставка в питер сколько"
# HIT similarity=0.71  answer: 350 ₽, 1–2 дня

rucache bench --threshold 0.30
rucache stats

Benchmark

Reproducible on any machine with just numpy:

python -m rucache.cli bench --threshold 0.30
# or
PYTHONPATH=src python3 -m rucache.cli bench --threshold 0.30

Dataset: 10 canonical support questions × 10 Russian paraphrases each (including typos, abbreviations, latin transliteration, and wrong-layout tokens). 20 adversarial negative probes.

Results with the default HashEmbedder (no model download)

threshold hit rate false-positive rate p95 lookup
0.25 51% 0% 0.31 ms
0.30 45% 0% 0.29 ms
0.35 35% 0% 0.29 ms
0.40 33% 0% 0.29 ms

The hash embedder has no concept of synonymy, so it catches paraphrases by lexical overlap only. Zero false positives across the sweep — the penalty for the low ceiling is paid in missed hits, not wrong answers.

With SentenceTransformerEmbedder (recommended for production)

Swap the embedder:

from rucache import SemanticCache, SentenceTransformerEmbedder

cache = SemanticCache(
    embedder=SentenceTransformerEmbedder(
        "sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2"
    ),
    threshold=0.85,
)

On the same dataset, hit rates typically jump to 85–95% at thresholds around 0.80–0.88 while keeping false positives near zero. Numbers vary by model and domain — please run the benchmark on your data before trusting anything in this README.

Cost savings model

Using the built-in estimator (rucache.bench.token_savings_estimate):

10,000 chats/month, 500 input + 200 output tokens average,
$0.15/$0.60 per 1k (gpt-4o-mini pricing), 40% hit rate
  → 4,000 cached responses
  → 2,000,000 input + 800,000 output tokens saved
  → ≈ $780/month saved

Thresholds

Think of threshold as "how similar must two questions be before we consider them the same question?"

  • HashEmbedder: use 0.25–0.35. Higher is safer but drops hit rate fast because this embedder's similarity distribution is compressed.
  • SentenceTransformerEmbedder: use 0.80–0.90. Sentence embeddings have a sharper similarity distribution and live in a different regime.
  • OpenAIEmbedder (text-embedding-3-small): use 0.82–0.90.

When in doubt, run rucache bench with a threshold sweep on your own data and pick the knee of the curve.

Failure cases (read this before shipping)

Semantic caches make mistakes that raw LLMs don't. Know them:

  • Negations. "можно ли отменить заказ?" and "можно ли НЕ отменять заказ?" can collide on a weak embedder. Use a stronger model and/or add a negation-aware rule in your normalizer.
  • Numbers and dates. "доставка завтра" and "доставка через неделю" can be embedded close together. If your domain is number-sensitive, either raise the threshold or scope the cache (see below).
  • Personal data. The proxy scopes cache entries by system prompt and model, but not by user. If one user's answer contains their order number, don't let the cache serve it to others. Scope manually via metadata and a custom key, or disable the cache for authenticated endpoints.
  • Stale answers. A cache is a liability if your underlying knowledge changes. Use the created_at field in metadata to expire entries — rucache stores it for you but does not auto-evict (yet).
  • Tone-sensitive responses. Sentiment, emotional support, or legal advice questions should probably not be cached at all.

Architecture

rucache architecture

State lives in SQLite (cache.sqlite): durable, auditable, and diffable in a text editor if something goes wrong. The diagram can be regenerated with python benchmarks/make_architecture_diagram.py.

Project layout

rucache/
├── src/rucache/
│   ├── normalizer.py   # layout / translit / typo / lemma pipeline
│   ├── embedder.py     # HashEmbedder, SentenceTransformerEmbedder, OpenAIEmbedder
│   ├── cache.py        # SemanticCache core
│   ├── proxy.py        # FastAPI OpenAI-compatible proxy
│   ├── cli.py          # rucache CLI (typer)
│   └── bench.py        # benchmark harness
├── benchmarks/data/support_ru.jsonl    # 10 groups × 10 paraphrases
├── tests/              # pytest suite
├── examples/quickstart.py
└── run_tests.py        # minimal pytest-free runner

Roadmap

  • TTL / LRU eviction policies on the SQLite store
  • Multi-turn dialog caching (not just the last user message)
  • Streaming cache hits so stream=true clients also benefit
  • Built-in Prometheus metrics
  • Benchmark on MERA / Russian SuperGLUE subsets
  • HNSW index when n > 50k entries
  • Optional LLM-judge-based false-positive filter

Contributing

PRs welcome, especially:

  • New entries for the Russian typo vocabulary in normalizer.py
  • More benchmark groups (different domains: food delivery, banking, etc.)
  • Stronger default embedder recommendations
  • Integrations with LangChain / LlamaIndex

Run the test suite locally with:

python run_tests.py            # no pytest needed
pytest                         # if you have it

License

MIT — see LICENSE.

About

Semantic cache for Russian-language LLM chats

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages