Skip to content
kenjiroePublic

About

Security-first async agent harness for MCP

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

46 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ozr

CI

Security-first, local-first async agent harness for the Model Context Protocol (MCP).

ozr orchestrates LLM planning, MCP tool execution, human approval, and audit logging in a pure async Rust runtime. Use it from the CLI, HTTP API, OpenAI-compatible clients, or the Tauri desktop GUI.

Why ozr?

  • Security-first guardrails — unknown tools escalate to high-risk Shell; medium/high actions block on human approval.
  • Non-blocking runtime — agent runs spawn on Tokio; API sessions poll independently (approval + concurrent /v1/run supported).
  • Provider-agnostic — mock, OpenAI-compatible, Anthropic, Gemini, and Ollama LLM backends; MCP via mock or stdio. Optional isolated execution via sandboxd (local setup guide).
  • Audit-ready — run logs, session checkpoints, replay reports under .ozr/.

See docs/architecture.md for the system blueprint.

Quick start

Build

cargo build --release

CLI (mock backend, no API keys)

./target/release/ozr run "read docs"
./target/release/ozr run "run mystery shell task"   # triggers approval gate

Initialize local config on first run:

./target/release/ozr init    # creates .ozr/config.env from defaults

HTTP API

./target/release/ozr serve
curl http://127.0.0.1:8080/health

curl -s http://127.0.0.1:8080/v1/run \
  -H 'content-type: application/json' \
  -d '{"prompt":"read docs"}'

OpenAI-compatible shim (JSON or SSE streaming):

curl -s http://127.0.0.1:8080/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"model":"ozr","messages":[{"role":"user","content":"read docs"}]}'

curl -N http://127.0.0.1:8080/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"stream":true,"messages":[{"role":"user","content":"read docs"}]}'

Endpoints: POST /v1/run · GET /v1/session/{id} · POST /v1/session/{id}/approve · POST /v1/chat/completions

Desktop GUI (beta)

cargo build
cd ui && npm install && npm run tauri dev

Default: GUI spawns ozr serve on http://127.0.0.1:18787. See ui/README.md.

Docker stack: start ./scripts/docker-up-stack.sh, then point the GUI at the container API:

OZR_GUI_API_BASE=http://127.0.0.1:8080 npm run tauri dev

Docker

Standard ports: ozr API 8080 · sandboxd 9090 · Qdrant 6333 · GUI dev 18787

Full stack — one command (ozr API + sandboxd + Qdrant, production policy, sandbox auto-wired):

chmod +x scripts/docker-up-stack.sh
./scripts/docker-up-stack.sh
curl http://127.0.0.1:8080/health

This validates compose, boots the stack, creates a sandboxd sandbox, and writes OZR_SANDBOXD_SANDBOX_ID into the ozr container plus host .ozr/config.env. High-risk shell tasks still require human approval in the GUI or API — that is expected.

GUI against the container API:

cd ui && OZR_GUI_API_BASE=http://127.0.0.1:8080 npm run tauri dev

Run a prompt such as run mystery shell task, approve when prompted, and confirm the run completes (no no such sandbox).

Infra only (host-native cargo / Tauri GUI — sandboxd + Qdrant in Docker):

chmod +x scripts/docker-up-infra.sh
./scripts/docker-up-infra.sh
./scripts/wire-sandboxd.sh
cd ui && npm run tauri dev

API only (mock backends, no sandboxd):

docker compose up -d --build
curl http://127.0.0.1:8080/health

Config template: .env.stack.example · Details: docs/sandboxd-local.md

Stack smoke tests

From repo root (no Docker required for Rust/UI unit tests):

./scripts/validate-stack-compose.sh
cargo test --lib -- --test-threads=1
cargo test --test api -- --test-threads=1
cargo test --test memory_recall_eval -- --test-threads=1
cd ui && npm test && npm run build

With the full stack running (./scripts/docker-up-stack.sh):

curl -sf http://127.0.0.1:8080/health && echo OK
OZR_API_BASE=http://127.0.0.1:8080 ./scripts/test-sandboxd-shell-api.sh

The live script posts run mystery shell task, auto-approves when pending_approval, and passes when the session completes with sandboxd_task= in the result.

Mock backend demo prompts (default OZR_LLM_BACKEND=mock, OZR_MCP_BACKEND=mock; catalog: read_file, list_tools):

Prompt Mock tool Approval What to expect
read docs read_file auto Summary with mock read result
list available tools list_tools auto Tool names from mock catalog
run mystery shell task run_shell required Routes to sandboxd when executor enabled

With sandboxd, an empty workspace may still complete successfully: the in-sandbox agent can reply that it has no shell tool, build_ok=false (no package.json), and files_changed=[] — that is normal. Failure signals include no such sandbox, session failed, or missing sandboxd_task= in the API result.

Configuration

All settings use OZR_* environment variables or .ozr/config.env.

cp .env.example .ozr/config.env

Key flags:

Variable Default Purpose
OZR_LLM_BACKEND mock LLM provider
OZR_MCP_BACKEND mock MCP transport
OZR_APPROVAL_MODE prompt CLI approval (auto / deny / prompt)
OZR_POLICY_PACK balanced production requires sandboxd for Shell/Write/Network
OZR_FEATURE_SANDBOXD_EXECUTOR false Route risky actions to sandboxd
OZR_API_BIND 127.0.0.1:8080 HTTP API listen address

Run ozr config to print the effective configuration (secrets shown as set/unset only).

Repository layout

src/          Rust core — agent loop, policy, API, CLI
ui/           Tauri + React desktop GUI (alpha)
tests/        Integration and E2E tests
docs/         Architecture and deployment guides
scripts/      Docker stack, sandboxd wiring, secret audit
ozr.md        Product blueprint
INTEGRATION_SPEC.md

Development

cargo test
cargo fmt --all -- --check
cargo clippy --all-targets -- -D warnings
./scripts/audit-secrets.sh

See CONTRIBUTING.md · SECURITY.md · CHANGELOG.md

Status

v0.1.0-beta.2 — one-click Docker stack, sandbox readiness, GUI external API settings + polish, memory recall eval CI, Vitest coverage. Prior: async core, Axum API, live LLM streaming, production hardening, sandboxd executor, GUI beta. See CHANGELOG.md.

License

MIT — see LICENSE.

About

Security-first async agent harness for MCP

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages