Warning Secure RAG is an experimental research alpha (
0.2.0a1). It is intended for privacy-aware RAG experimentation and benchmarking, not production or clinical use.
Secure RAG is a privacy-aware Retrieval-Augmented Generation framework for local document intelligence.
Its key research contribution is Pre-Embedding Privacy Enforcement, where sensitive information is detected and masked before chunking, embedding, and vector indexing, ensuring private data never enters the retrieval pipeline in raw form.
This allows the framework to prevent sensitive data from entering the vector store in raw form. The current system demonstrates the privacy-utility tradeoff: pre-embedding masking improves privacy but can reduce retrieval quality for identity-based queries.
- Mandatory pre-embedding document masking
- Support for
.txtand.pdfinputs - FAISS-backed semantic retrieval
- Streaming answer generation
- Lazy loading for embedding and API clients
- Packaged Python API and CLI entry point
- Research-oriented architecture for privacy-aware RAG experiments
Secure RAG follows a privacy-first retrieval pipeline:
- Load a local document from
.txtor.pdf - Parse and validate input
- Apply privacy masking before chunking
- Split the document into overlapping chunks
- Convert chunks into embeddings
- Index embeddings in FAISS
- Retrieve relevant chunks using the raw user query
- Stream the grounded response back to the caller
This design ensures sensitive entities can be protected before entering the vector store, which is the framework's core research contribution.
Secure RAG enforces privacy by masking sensitive entities before chunking, embedding, and vector indexing. This ensures:
- Raw sensitive data never enters the vector store
- The user query is never modified
- Retrieval operates on masked content only
The privacy-utility tradeoff is a core research result: pre-embedding masking improves privacy but can reduce retrieval quality for identity-based queries.
- This is an experimental research framework, not a production-ready RAG package.
- Identity-based retrieval can fail because raw names in queries may not align with masked indexed content.
- Some LLM backends may echo prompt structure or use outside knowledge despite grounding instructions.
- Output quality depends significantly on the configured model provider.
- Local setup may require the spaCy model
en_core_web_sm.
pip install -i https://test.pypi.org/simple/ secure-ragpip install -e ".[dev]"Install the spaCy model:
pip install https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.8.0/en_core_web_sm-3.8.0-py3-none-any.whlDepending on the backend you use, you may also need:
HF_API_KEY/HF_TOKENLLM_PROVIDER=ollamaandOLLAMA_MODEL=llama3.2
Secure RAG includes a ready-to-use Docker Compose setup.
Project structure:
data/
└── sample_patient_data.txt
The data/ directory is mounted into the container at /data. Any supported document placed in this directory is accessible to the CLI.
Quick start:
# Build the image and start an interactive CLI session
docker compose run --rm secure-ragThe first run builds the image, which installs dependencies and bakes in the spaCy model. Subsequent runs are instant.
docker compose run --rm secure-ragis the recommended way to interact with the CLI. It attaches directly to your terminal, handles stdin and signals correctly, and removes the container after exit.
docker compose upis useful when running the full Compose stack (e.g. with the Ollama service via--profile ollama), but it is not the preferred way for interactive CLI usage.
Changing the input document:
Open .env and update RAG_INPUT_FILE to point to a different file inside data/:
RAG_INPUT_FILE=my_own_document.txt
Place the file in data/ and run docker compose up again.
Using a local Ollama instance:
docker compose --profile ollama upThis starts both Secure RAG and an Ollama container. Set LLM_PROVIDER=ollama in .env to use it.
Notes:
- The spaCy model
en_core_web_smis baked into the image — no manual download needed. - The embedding model (
all-MiniLM-L6-v2) downloads on first use and is cached inside the container. - Environment variables are loaded from
.envautomatically (API keys, provider, model selection).
from secure_rag import build_rag, rag_answer
vector_store, chunks = build_rag("test_data.txt")
answer = "".join(rag_answer("what treatment is given for chest pain?", vector_store, chunks))
print(answer)secure-rag test_data.txtSecure RAG includes a benchmark script for comparing privacy strategies:
python benchmarks/privacy_eval.pyThe benchmark compares three evaluation configurations:
- Baseline A: Raw Retrieval-Augmented Generation (no masking)
- Baseline B: Post-Retrieval Privacy Masking (masking after retrieval)
- Secure RAG: Pre-Embedding Privacy Enforcement (canonical runtime)
Current evaluation metrics:
- Document leakage
- Retrieval leakage
- Masking recall
- PHI in answers
Note: results depend on the configured model backend and local environment. LLM-based evaluation requires API or local model setup.
Secure RAG is a research framework for studying privacy-preserving retrieval-augmented generation. It is not intended for medical decision-making, clinical deployment, or production handling of regulated sensitive data.