This document explains crucial components, why they exist, and the main tradeoffs behind each one.
Source: src/presentation/ui_bridge.py
Purpose:
- Hold small shared runtime signals for the CLI and graph streaming loop.
- Keep UI state outside core graph state.
Key fields:
current_nodefor live progress display.trial_numfor retry progress.disambiguatorcallback for interactive title resolution.page_cachesfor task-scoped cache handoff.
Tradeoff:
- This is global process state, which is simple and fast for a CLI process, but should be replaced by stronger isolation if moving to multi-tenant server runtime.
Source: src/utils/page_cache.py
Purpose:
- Overlap fetch latency with ongoing graph execution.
- Deduplicate page fetches in a single run.
Design:
- Cache stores asyncio tasks keyed by normalized IDs.
- Prefetch can be triggered after resource resolution and execution.
- Supports refresh for mutated pages.
Tradeoff:
- Not serializable by design, so it stays out of graph state and is injected via runnable config.
Source: src/utils/telemetry.py
Purpose:
- Track which Notion pages were read or mutated by generated code.
Design:
- Injects a request interceptor into generated code before execution.
- Extracts IDs from page/block endpoint traffic.
- Injects
RESOURCE_MAPdirectly into code scope. - Supports fallback file-based affected-ID retrieval.
Tradeoff:
- Monkey-patching is pragmatic but requires careful re-entrancy handling in reused interpreters.
Sources:
Purpose:
- Separate stable secrets from ephemeral evaluation IDs.
Design:
.envfor stable credentials/config..env.sandboxfor generated test IDs and short-lived sandbox artifacts.- Runtime can load both with sandbox values overriding overlaps.
Tradeoff:
- Slightly more setup complexity in exchange for safer evaluation operations and fewer accidental edits of base secrets.
Sources:
Purpose:
- Stop unsafe/out-of-scope prompts early.
- Extract required resource titles before code generation.
Design:
- General guard returns strict JSON: reasoning, scope flag, required resources.
- Security guard uses Llama Guard result.
- Pipeline joins both checks before continuing.
Tradeoff:
- Additional model calls add latency but significantly improve reliability and safety.
Source: src/nodes.py
Purpose:
- Resolve user-mentioned page titles into concrete Notion IDs.
Design:
- Searches Notion by title for each required resource.
- Uses interactive disambiguation callback when multiple matches exist.
- Fails explicitly when no disambiguator is available in non-interactive runs.
Tradeoff:
- Extra pre-execution API calls reduce downstream code hallucination and bad ID usage.
Sources:
Purpose:
- Execute generated code in isolated runtime by default.
Design:
- Prepare/connect sandbox, execute with timeout, then cleanup.
- Egress policy allows only
api.notion.com. - Sends only
NOTION_*env values to sandbox.
Tradeoff:
- Slight startup overhead for sandbox setup, but improved isolation and safer execution.
Source: src/nodes.py
Purpose:
- Prevent accidental secret leakage in outputs.
Design:
- Post-execution scan checks output text for configured sensitive token values.
- If leaked, output is replaced and terminal status is set to
security_blocked.
Tradeoff:
- String matching is conservative and simple; false positives are possible but acceptable for safety.
Source: src/utils/openai_utils.py
Purpose:
- Keep async LLM client lifecycle bounded to each run.
Design:
- Context-managed session with contextvar storage.
- Reset and close at end of lifecycle.
Tradeoff:
- Slight setup/teardown overhead per run, but avoids stale client/resource leaks across CLI turns.
Source: src/models/hardcoded_contexts.py
Purpose:
- Provide deterministic retrieval context variants without requiring live vector retrieval.
Design:
- Loads context files from
data/contextplus in-code baselines. - Builds named combinations (for example schema report + Notion API summary).
- Supports
dynamicmode as a separate route.
Tradeoff:
- Strong determinism and simpler ops, but less adaptive than live retrieval.
Sources:
Purpose:
- Convert evaluation outputs into structured diagnostics and summaries.
Design:
- Can run automatically after evaluation orchestration.
- Supports section toggles for code, execution, statements, and retrieval context.
Tradeoff:
- Additional post-processing time, but much faster iteration on failure patterns.
Source: src/models/schema.py
Frequently surfaced fields:
execution_outputmessage_to_userfeedbackverdictaffected_notion_idsrelevant_page_ids