Two supported topologies: Docker Compose on a single host (the local profile), and free
hosted tiers with the cloud profile (hosted Qdrant/Jina/Groq/etc.). For cloud, the API
can run on Vercel (Python FastAPI function) or Render (Docker); frontends stay on
Vercel. This page covers hosted deploy; see local-setup.md for Compose.
Frontends (separate Vercel projects; Root Directory = app folder):
website/- public landing (hero ask, chat answers, product narrative).web/- operator console (latency, strategy compare, CRAG/guardrail traces, bench).
Neither Vercel nor Render free has an RQ worker or durable disk for ingestion. Run ingest and calibrate from your machine against hosted Qdrant, Postgres, and embedding:
cp .env.cloud.example .env # fill in every credential first
uv run python scripts/ingest-self.py --site-url https://your-site.example
# ...or the MSMARCO-XI corpus instead:
# uv run python scripts/ingest-msmarco.py --rows-per-language 250 --max-chunks 90000
uv run python -m fastrag.calibrate \
--golden eval/calibration.jsonl --cache-pairs eval/cache_pairs.jsonlThe index lives in Qdrant Cloud + Neon. The API only needs config/calibration.json at
startup (it refuses to boot without it). Re-ingest offline when the corpus changes.
Calibration records the content_version it was fitted against, and startup rejects a
mismatch, so re-ingesting without recalibrating fails loudly instead of serving the new
corpus through the old thresholds. When re-ingesting against a deployment that is already
serving, use --no-activate and flip the alias after calibrating; see
self-corpus.md.
Use a separate Vercel project with Root Directory = repository root (not web/ /
website/). Vercel detects FastAPI from pyproject.toml; the entrypoint is declared there:
[tool.vercel]
entrypoint = "src.fastrag.api:app"Root vercel.json sets maxDuration to 300s on src/fastrag/api.py so
voice + SSE streams are not cut off early, and excludes the Next.js trees from the Python
bundle.
- Import the Git repo as a new Vercel project (root directory
.). - Copy every variable from
.env.cloud.example/ your local.envinto Project → Environment Variables. Minimum for a working boot:FASTRAG_PROFILE=cloudFASTRAG_ENVIRONMENT=production(ordevelopment)FASTRAG_SPARSE_RETRIEVAL_ENABLED=falseFASTRAG_QUERY_API_KEY,FASTRAG_ADMIN_API_KEYFASTRAG_QDRANT_URL,FASTRAG_QDRANT_API_KEYFASTRAG_JINA_API_KEYFASTRAG_LLM_BASE_URL,FASTRAG_LLM_API_KEY,FASTRAG_LLM_MODELFASTRAG_DATABASE_URL,FASTRAG_REDIS_URLFASTRAG_CALIBRATION_JSON- paste the full JSON from localconfig/calibration.json(gitignored; without this, startup fails)FASTRAG_ALLOW_QUERY_OVERRIDES=false(default in.env.cloud.example; settrueonly on trusted staging for/query-style experiments)
- Redeploy. Open
/health/ready- on failure it returns{"status":"not_ready","error":"..."}instead of a blank 500. Fix whatevererrornames, then confirm/buildwith the admin key.
Point frontend projects at this URL via FASTRAG_API_URL (and FASTRAG_QUERY_TOKEN =
FASTRAG_QUERY_API_KEY).
Cold starts on Fluid compute rebuild the pipeline in lifespan; first request after idle is
slower. Python excludeFiles in root vercel.json keeps web/ / website/ out of the
function bundle; they must still exist in the Git upload for the frontend projects.
render.yaml is a Blueprint: free Docker web service, health check
/health/ready, binds 0.0.0.0:$PORT. FASTRAG_QUERY_API_KEY / FASTRAG_ADMIN_API_KEY
use generateValue; credentials marked sync: false are prompted on first deploy.
config/calibration.json is COPY’d only if present in the build context - it is
gitignored, so production must set FASTRAG_CALIBRATION_JSON (declared in
render.yaml). The image entrypoint writes that env var to
/app/config/calibration.json before uvicorn starts.
Blueprint defaults: FASTRAG_PROFILE=cloud, FASTRAG_SPARSE_RETRIEVAL_ENABLED=false.
Free instances spin down after ~15 minutes idle.
The pair described in providers.md
spends no Jina balance, but it peaks at 3.7 GB resident and the weights total about 3.4 GB,
so it cannot run as a Vercel function or on Render's free instance. Use a Docker host with at
least 4 GB of memory (Render's 4 GB plan, or your own box). Reranking dominates latency:
5.4 s for 20 candidates on 16 CPU cores with FASTRAG_RERANKER_BATCH_SIZE=1, and it scales
with cores, so a two-core instance would approach the 60 s request deadline. Prefer a host
with a GPU (FASTRAG_RERANKER_EXECUTION_PROVIDERS=CUDAExecutionProvider, 0.47 s on an
RTX 3050 Ti, which needs the onnxruntime-gpu wheel in place of onnxruntime), or lower
FASTRAG_RETRIEVAL_CANDIDATE_K, which trades recall for time.
The image does not bake models in. Set these alongside the variables in that section, and the entrypoint fetches the pinned revisions at boot and verifies their checksums before uvicorn starts, refusing to serve on a mismatch:
FASTRAG_ENVIRONMENT=production
FASTRAG_DOWNLOAD_MODELS=true
FASTRAG_DENSE_MODEL_PATH=/models/dense
FASTRAG_RERANKER_MODEL_PATH=/models/reranker
FASTRAG_CHUNK_SIZE=150
FASTRAG_RERANKER_BATCH_SIZE=1
FASTRAG_GUARDRAIL_OFFTOPIC_ENABLED=falseWithout a persistent disk mounted at /models, every boot downloads the 3.4 GB again.
Ingest and calibrate with the identical model variables: the embedding fingerprint covers
model, revision, checksum and prefixes, and startup rejects an index built with any other.
FASTRAG_JINA_API_KEY is not needed in this configuration.
Use separate Vercel projects from the FastAPI API project (do not set Root Directory to
. for these).
| Project | Root Directory | Role |
|---|---|---|
| Landing | website/ |
Public marketing site + hero ask / chat + /query pipeline trace |
| Console (optional) | web/ |
Operator latency / CRAG / bench UI |
Do not put web/ or website/ in the root .vercelignore - that file
applies to every project in the monorepo and would strip the Next.js app before install.
Each has its own vercel.json (maxDuration 60s on the RAG proxy for SSE). In each project set:
FASTRAG_API_URL- your FastAPI Vercel URL (e.g.https://fast-rag-….vercel.app).FASTRAG_QUERY_TOKEN- same value asFASTRAG_QUERY_API_KEYon the API.
Neither is prefixed NEXT_PUBLIC_. Proxy route app/api/rag/[...path]/route.ts keeps the
token server-side, so the landing origin does not need to be in FASTRAG_CORS_ORIGINS
unless you bypass the proxy and call the API from the browser.
The landing project also reads two optional public settings, inlined at build time (redeploy after changing them). Leave either unset to hide what depends on it:
NEXT_PUBLIC_GITHUB_REPO_URL- the repository (e.g.https://github.com/owner/FastRAG): the header star button, footer and call-to-action links, and docs links to repository files.NEXT_PUBLIC_GITHUB_REPO_BRANCHsets the branch those file links use (defaultmaster).NEXT_PUBLIC_DEVELOPER- the developer credited in the Developer section and footer, as one line of JSON:{"name":"…","role":"…","bio":"…","photo":"/developers/x.png","github":"…", "linkedin":"…","website":"…","email":"…"}. Onlynameis required.
Query overrides: .env.cloud.example sets FASTRAG_ALLOW_QUERY_OVERRIDES=false.
Keep that on production API deploys. See query-trace.md.
docker compose up no longer starts Langfuse. Tracing points at Langfuse Cloud's free tier
by default, which needs no ClickHouse, no MinIO, and no second Redis. To run it yourself:
docker compose --profile langfuse up -dand set the commented block at the bottom of .env.local.example, plus
LANGFUSE_BASE_URL=http://langfuse-web:3000. Nothing else changes; the application code
is identical either way.
curl -fsS https://your-api.vercel.app/health/ready
curl -fsS -H "Authorization: Bearer $KEY" https://your-api.vercel.app/build
# or https://your-api.onrender.com/.../build reports the release, profile, and the three active providers, which is the quickest
way to confirm the deployed service is wired the way you think it is. Then drive a real voice
query through website/ or web/ and check that transcript, citations, guardrail decision,
CRAG trace, and Langfuse spans all appear. On website/, run a hero query and open /query
to confirm the trace payload and Experiment re-run path.
No HA for Compose; cold starts on free hosts; dense-only when sparse is off; LLM free-tier rate limits. Demo and evaluation, not production HA. Same application code on local and cloud.