Ultrafuzz is a TypeScript workspace managed with pnpm.
The supported host runtime is Node.js 22.19 or newer. This satisfies the
pinned pnpm 11 toolchain, the pinned Smithers release's Node 22 declaration, and
the exact Kimi Code 0.29.1 CLI used by the adapter contract tests. Smithers'
executable and those contract tests run with Bun 1.3+. The Modal image uses
Node.js 22.23.
Generated-adapter contracts are registered under the Bun adapter contract:
test-name prefix and declare a 30-second cold-cache timeout (real Kimi CLI
probes declare 60 seconds). Bun's node:test shim ignores the { skip: ... }
option even though it honors timeout; use the shared testWhen selector,
which calls test.skip, for any conditional test in the runtime suite. A
supporting test rejects option-based skips in that Bun-loaded file.
Workflow execution is supported on Linux with procfs mounted at /proc.
Sealed workflow controls are opened through directory descriptors, and the
Smithers child receives the held generation as
/proc/<controller-pid>/fd/<descriptor>. A self-process /dev/fd alias does
not provide that cross-process execution path. Windows and macOS execution are
not currently supported by this hardened boundary.
pnpm -w format:check
pnpm -w lint
pnpm -w typecheck
pnpm -w build
pnpm -w test
pnpm -w validate:release
pnpm -w docs:checkThe root CI script runs format check, lint, build, and release validation:
pnpm -w run ciEvery CI run, pull requests included, runs the build gates (CI policy checks,
formatting, lint, dead-code checks, the workspace build, strict lint of changed
lines, bundle budgets, and dependency policy) and all nine release validation
lanes: package gates, one runtime supporting lane that includes the Bun adapter
contracts, four runtime integration shards, the CLI suite, the cli-e2e lane,
which runs one campaign end to end (see below), and benchmark history with the
workspace typecheck. The lanes do not wait for the build gates and run at most
eight at a time. They run their tests under eatmydata, which turns fsync
into a no-op in the test processes but not in the Smithers engine processes
those tests launch. Feature branches are validated only by the pull-request
event, avoiding a duplicate push run. A newer push to a pull request cancels
that pull request's older run; a push to unstable never cancels another run. On
unstable, the lane results are merged into one JSON report in stable gate order.
Focused package iteration uses pnpm filters:
pnpm --filter @ultrafuzz/config test
pnpm --filter @ultrafuzz/topology test
pnpm --filter @ultrafuzz/prompts test
pnpm --filter @ultrafuzz/references test
pnpm --filter @ultrafuzz/artifacts test
pnpm --filter @ultrafuzz/runtime test
pnpm --filter @ultrafuzz/cli test
pnpm --filter @ultrafuzz/modal testpnpm --filter @ultrafuzz/cli test:e2e runs the end-to-end campaign test in
packages/cli/test/e2e/. It drives init, run, resume, status, stats,
report, and events as separate CLI processes, and the generated workflow
runs on the pinned Smithers engine under Bun, with a stub codex executable in
place of the model. It SIGKILLs the detached controller while one node is
running, resumes the run, and checks that it succeeds with a verified report,
that no finished task started again, and that status and stats count the
same agent attempts, including the one the kill interrupted. It needs Linux,
Bun, Git, and access to the npm registry, because run and resume install the
pinned engine from npm as they do for any campaign. On SIGINT or SIGTERM, the
test kills the detached campaign and deletes its fixture, about 1 GB, before it
exits.
Package-local typecheck and test scripts may build direct workspace
dependencies first because package exports point at dist/**.
pnpm -w docs:checkThe docs check verifies required documentation entrypoints and rejects drift in the generated audit-profile and prompt-catalog references. It does not build a static site.
Modal unit and contract tests are deterministic and local:
pnpm --filter @ultrafuzz/modal test
pnpm --filter @ultrafuzz/modal typecheck
pnpm --filter @ultrafuzz/modal buildThese tests also validate the exact three-target smoke and four-provider full EVMBench configuration, the fixed Sol judge, bounded row and control deadlines, immutable image naming, and hash-manifested public bundles. A workflow policy test keeps Actions limited to ordinary CI without provider secrets or paid launch commands. The tests make no cloud or model calls; paid benchmarks and history publication require explicit maintainer runs outside Actions.
The real-cloud smoke is deliberately separate from every normal test and CI script. It must be selected explicitly, once per provider:
pnpm --filter @ultrafuzz/modal smoke -- --provider openai
pnpm --filter @ultrafuzz/modal smoke -- --provider anthropic
pnpm --filter @ultrafuzz/modal smoke -- --provider deepseek
pnpm --filter @ultrafuzz/modal smoke -- --provider kimiDo not add the smoke to test, gate it on an environment variable, or replace
its generic fixture with real project material. A passing run exercises the
published production image, provider-isolated subscription auth, non-root
execution, writable durable storage, mid-run termination, same-volume resume,
completed-work reuse, and single launch ownership. Output is restricted to
aggregate checks; cloud IDs, credentials, provider output, prompts, findings,
source contents, and artifacts are not test output.
Build output lives under package dist/ directories and is not a product
surface. User-owned Ultrafuzz state remains root ultrafuzz.toml and
.ultrafuzz/**.