Run local LLMs with Ollama, a clean web UI (Open WebUI), and secure remote access over a Tailscale VPN β your own private ChatGPT on your hardware.
Keep your prompts and data private and free: serve open models locally, reach them from anywhere via VPN (no ports exposed to the internet), and avoid per-token costs for everyday use.
- π§ Ollama β local model serving. Ships with a small, CPU-friendly default
set:
llama3.2:3bandqwen2.5:3bfor chat, plusnomic-embed-textfor Open WebUI's RAG/embeddings. Fully configurable viaOLLAMA_MODELS. - π¬ Open WebUI β chat interface, model switching, conversation history, RAG.
- π Tailscale β private mesh VPN; access your stack from laptop/phone with zero open ports (Tailscale Serve publishes Open WebUI over tailnet HTTPS).
- π³ Docker Compose β one command to bring it all up.
- π Caddy (optional) β reverse proxy with automatic public TLS for real
domains, behind the
tlsCompose profile.
All images are pinned for reproducibility.
flowchart LR
Dev[Laptop / Phone] -- Tailscale VPN --> TS[tailscale container]
subgraph Host[Docker host]
TS -- "Serve (HTTPS)" --> UI[Open WebUI :8080]
UI --> O[Ollama :11434]
end
Dev -. "optional public HTTPS" .-> Caddy[Caddy :443]
Caddy --> UI
No host ports are published by default. The only way in is through your tailnet (or, optionally, the Caddy reverse proxy if you explicitly enable it).
git clone https://github.com/AntoniRomera/self-hosted-ai.git
cd self-hosted-ai
cp .env.example .env
# Edit .env and set at minimum:
# TS_AUTHKEY - from https://login.tailscale.com/admin/settings/keys
# WEBUI_SECRET_KEY - generate: openssl rand -hex 32
$EDITOR .env
# Bring it all up and preload the configured models:
make upThen open https://<TS_HOSTNAME>.<your-tailnet>.ts.net from any device on your
tailnet. (TS_HOSTNAME defaults to self-hosted-ai.)
Local-only access without Tailscale? Copy
docker-compose.override.yml.exampletodocker-compose.override.yml, thenmake up. Open WebUI will be reachable athttp://127.0.0.1:3000.
Generate an auth key in the Tailscale admin console. For an unattended home
server, a reusable + ephemeral + tagged key works well. Put it in .env:
TS_AUTHKEY="REPLACE-WITH-YOUR-TAILSCALE-AUTHKEY"The key is read only from the environment β it is never committed (.env is
gitignored). The container runs in userspace networking mode, so it needs no
special host privileges.
| Variable | Description | Default |
|---|---|---|
OLLAMA_MODELS |
Comma-separated models to preload | llama3.2:3b,qwen2.5:3b,nomic-embed-text |
OLLAMA_KEEP_ALIVE |
How long a model stays in memory after last use | 5m |
WEBUI_PORT |
Host port (only when local-only override is enabled) | 3000 |
WEBUI_SECRET_KEY |
Required. Signs session cookies (openssl rand -hex 32) |
none |
WEBUI_AUTH |
Set false to disable login (trusted single user only) |
true |
TS_AUTHKEY |
Required secret. Tailscale auth key | none |
TS_HOSTNAME |
Node name on the tailnet (MagicDNS) | self-hosted-ai |
TS_EXTRA_ARGS |
Extra args for tailscale up |
empty |
COMPOSE_PROFILES |
Set to tls to enable the Caddy reverse proxy |
empty |
CADDY_DOMAIN |
Public domain for Caddy TLS (required with tls profile) |
ai.example.com |
*_IMAGE |
Pinned image tags (OLLAMA_IMAGE, OPENWEBUI_IMAGE, β¦) |
see .env.example |
Full deep-dive: docs/CONFIGURATION.md.
LLMs are memory-bound. Pick model sizes to fit your RAM (CPU) or VRAM (GPU):
| Hardware | Comfortable model sizes | Examples |
|---|---|---|
| CPU only, 8 GB RAM | ~3B (quantized) | llama3.2:3b, qwen2.5:3b |
| CPU only, 16 GB RAM | up to ~7-8B (slow) | llama3.1:8b, qwen2.5:7b |
| GPU, 8-12 GB VRAM | 7-8B comfortably | llama3.1:8b, mistral:7b |
| GPU, 24 GB VRAM | 14-32B (quantized) | qwen2.5:14b, qwen2.5:32b-instruct-q4 |
The defaults target CPU-only 8 GB+ hosts so the stack just works out of the box.
Enable an NVIDIA GPU: uncomment the deploy.resources (or gpus: all) block
under the ollama service in docker-compose.yml. Requires the
NVIDIA Container Toolkit.
Details in docs/CONFIGURATION.md.
Two supported paths:
- Tailscale Serve (default, recommended). Zero-config HTTPS using tailnet
certificates. Open WebUI is published only on your tailnet β nothing is
exposed publicly. Config lives in
tailscale/serve.json. - Public Caddy + real domain. For exposing the UI on the public internet.
Set
COMPOSE_PROFILES=tlsandCADDY_DOMAIN=your.domainin.env, point the domain's DNS at the host, and Caddy provisions a Let's Encrypt certificate automatically. SeeCaddyfile.
Chat history and settings live in the openwebui-data volume (a SQLite DB at
/app/backend/data).
# Back up Open WebUI data to ./backups/openwebui-<timestamp>.tar.gz
make backup
# Also back up downloaded models (large):
make backup-all
# Restore a backup (stop the stack first):
make down
make restore FILE=backups/openwebui-20260101-120000.tar.gz
make upBackups use an ephemeral alpine container to tar the named volume, so they work
while the stack is running and need no host tooling.
| Target | What it does |
|---|---|
make up |
Start the stack and preload models |
make down |
Stop and remove containers |
make logs |
Follow logs |
make ps |
Show service status |
make pull-models |
Pull models from OLLAMA_MODELS |
make backup |
Back up Open WebUI data |
make restore |
Restore a backup (FILE=...) |
make validate |
docker compose config sanity check |
make lint |
Run yamllint + shellcheck |
make test |
Run the compose smoke tests |
make clean |
Remove containers and volumes (data loss) |
This is an infra project, so "tests" mean validating the stack definition
rather than running it. tests/smoke-compose.sh
renders the Compose config with dummy secrets and asserts the invariants that
matter here:
- every service is defined and wired correctly (
open-webui->ollama); - no host ports are published in the default config;
- the optional override binds only to
127.0.0.1; - Caddy appears only under the
tlsprofile; - all images are pinned (no
:latest); - required secrets (
WEBUI_SECRET_KEY,TS_AUTHKEY) are enforced; tailscale/serve.jsonis valid JSON proxying toopen-webui:8080;.env.exampleholds placeholders, never real keys.
make test # runs tests/smoke-compose.shThe same checks (plus yamllint, shellcheck, and caddy validate) run in CI on
every push and pull request via
.github/workflows/ci.yml.
- Can't reach the UI on the tailnet. Check
make logsfor thetailscaleservice. Confirm the node appears in your Tailscale admin console and that MagicDNS + HTTPS certificates are enabled for your tailnet. WEBUI_SECRET_KEY is required. You started Compose without setting it. Runmake up, which validates.env, or set the value beforedocker compose up.- Model pull is slow / fails. Large models take time. Re-run
make pull-models; it is idempotent and resumes. - Out of memory / killed. The model is too big for your RAM/VRAM. Pick a smaller size from the hardware table above.
- Caddy can't get a cert. Ensure
CADDY_DOMAINresolves to this host and ports 80/443 are reachable from the internet.
Part of a small self-hosting / infra portfolio:
terraform-aws-modulesβ reusable AWS infrastructure modules.ai-agents-frameworkβ agent orchestration you can point at this stack's Ollama endpoint.mcp-erp-serverβ MCP server example.ai-job-aggregatorβ data pipeline example.openclaw-integrationβ integration glue.
MIT Β© Antoni Romera Luis