All notable changes to AgentDiff (agent-trajectory-diff) are documented here.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
- Website extracted to
kerrshift/agentdiff-website(full git history preserved): this repository is now product-only Python. The site owns its own CI/CD — GitHub Actions builds and deploys agentdiff.app to Cloudflare Pages on every merge there.Makefilewebsite-*targets and the product/website PR-split rules are retired; the token-serviceworker/remains here as product infrastructure.
0.5.1 - 2026-09-17
AgentDiff v0.5.1 — a patch for a false positive found by the project's own live gate: consecutive repetition of a step pattern is now a loop only when the repeated steps are stagnant. Iterating one tool across items no longer blocks the build.
- Sequence-loop detector no longer blocks legitimate iteration: repetition
of a step-name pattern was treated as a loop even when the repeated steps had
different inputs and outputs, so a tool-calling agent that iterates one tool
across items (
decide -> call -> decide -> call -> decide -> call, e.g. one call per state) was hard-blocked withBlocked by: loopswhile its trajectory divergence was0.0. Repetition now counts as a loop only when the repeated steps are stagnant (identical input payloads and outputs), matching the identical-call invariant's standard. Runaway behaviour with drifting arguments remains covered bymax_tool_repeatsand the resource bands. Found by this project's own live gate running against the public demo repository.
0.5.0 - 2026-09-02
AgentDiff v0.5.0 — the Circuit Breaker release: statistical baselines, an
honest gate, and a self-serve approve bot. Highlights: N-run baseline
envelopes with variance bands (flaky agents stop flipping the gate), decoupled
failure modes (loops always block; path drift and cost are human judgments),
/agentdiff approve re-baselines from the PR thread as agentdiff-ci[bot]
via the hosted identity service, agentdiff init writes a production-ready
gate in seconds, and every report's provenance now describes the gate that
actually judged it.
- Hosted identity service (
worker/): a stateless Cloudflare Worker (free tier, ~100 lines, nothing stored) that mints 1-hour GitHub App installation tokens for AgentDiff workflows. Once the project App is installed on a repo, the approve bot comments as agentdiff[bot] with zero per-repo configuration — no secrets, no variables. Security model: the caller must present a Actions token that already has access to the repo it claims, and the App must be installed there; tokens expire in ≤1 hour regardless. Deploy checklist inworker/README.md. - Three-tier bot identity in generated workflows: self-managed App secrets →
hosted token service →
github-actions[bot]— each tier fully automatic, degrading gracefully. /agentdiff approve— the interactive PR bot (Pillar 3 of the 0.5.0 release, pulled forward): reviewers comment/agentdiff approveon a flagged PR and the candidate run becomes (or joins) the golden baseline — no local checkout, no manual JSON. Policy D3: loop violations are never blessable; path drift and cost spikes are human judgments.agentdiff approve BASELINE CANDIDATE [--pr N]powers it.- Branded bot identity: the approve workflow authenticates as a GitHub
App (
agentdiff[bot]with AgentDiff's logo — the repo ships the avatar PNG) viaactions/create-github-app-token, falling back togithub-actions[bot]when the App isn't configured. App-token pushes also re-trigger the check (GITHUB_TOKEN pushes never do), so the check genuinely flips to PASSED after approval. - Zero-setup approve bot: the default path needs no configuration at
all — the bot runs on the workflow's own
GITHUB_TOKENand, after re-baselining, posts a genuine greenAgentDiff Checkresult on the new head via the Checks API. The GitHub App is purely optional branding (identity + auto re-trigger), and a one-click website installer (App Manifest flow, client-side, $0 hosting) lands with the site update. agentdiff init --with-approve: additionally writes the approve-bot workflow — command filter (/agentdiff approve, bots excluded), write-permission gate on the commenter, concurrency guard, and candidate-trace artifact handoff from the check workflow (which now uploads it).- Error Recovery Cascade is now a default hard gate (PRD failure class):
a candidate that spends ≥ 3× the baseline's post-error recovery steps
blocks CI (exit 1). Explicit
--max-recovery-ratiostill wins; pass a large value (or[cli] max_recovery_ratio) to tune it. agentdiff init— zero-config CLI wizard (Pillar 4 of the 0.5.0 "Circuit Breaker" release): auto-detects the agent framework (LangGraph, CrewAI, OpenAI Agents SDK, OpenTelemetry/OpenInference, or generic) and writes a production-readyagentdiff.toml(v0.5 spec, statistical scenario) plus a GitHub Actions gate workflow in seconds. Idempotent — refuses to clobber without--force;--adapteroverrides detection;--scenario/--runsparameterize the generated config.- Human-first PR verdict (positioning fix, product side): the PR comment now leads with what reviewers care about — "No infinite loops. No cost spikes. Trajectory within budget — safe to merge." or "Blocked by: tool_loop." — with the metric table demoted to a collapsible Gate details section.
- Statistical envelopes — N-run baselines with variance bands (Pillar 1 of
the 0.5.0 "Circuit Breaker" release): baselines can now capture N ≥ 2 runs
(
agentdiff record ... --runs 5) into a versionedagentdiff_baseline_envelopeartifact (schema/baseline_envelope.schema.json, schema 2.0.0 — additive;agent_trace.schema.jsonuntouched). A candidate passes when some recorded run explains it (min-TDI-of-N) and its resource profile sits within mean ± k·sigma bands. Existing v1 baselines keep working:load_baseline()wraps a bare trace as an envelope with N=1, modestrict. - Topological equivalence (commutative subgraphs): executing
[A -> B]or[B -> A]is zero-penalty when the swapped steps are data-independent (no parent-graph path in either direction) and did identical work; steps with a dependency path, changed arguments, or changed outcomes still count as real divergence. Newmatched_commutativediff status renders as "reordered" across terminal/JSON/markdown/PR. [scenario.<name>]config sections (PRD v0.5 spec):mode,sample_runs,max_cost_increase_pct,[scenario.x.hard_invariants](fail_on_identical_loops,max_tool_repeats), and[scenario.x.tolerances](step_count_std_dev,divergence_ceiling). A repo with a single scenario needs no--scenarioflag.- Branded PR reports from generated gate workflows (
agentdiff init): the gate's PR-report step mints anagentdiff[bot]token from the hosted identity service (token.agentdiff.app) when the AgentDiff App is installed on the repo, and silently falls back to the workflow's ownGITHUB_TOKENotherwise. Verified live onagentdiff-demo. - CLI statistical mode: envelope baselines gate via variance bands
(step-count band, envelope-relative cost ceiling = max of the relative cap
and the k·sigma band, divergence ceiling); strict single-run baselines are
unchanged.
--update-baselineon an envelope rotates a rolling window ofsample_runsand recomputes bands. - Envelope benchmarks (N=3/N=5 at 100/500/1000 steps): compare cost scales linearly with N (N=5 @ 1000 steps ≈ 7.4s worst case; realistic KB-scale traces are milliseconds).
- Decoupled failure modes — hard gates vs soft warnings (Pillar 2 of the
0.5.0 "Circuit Breaker" release): gate evaluation is now severity-aware.
HARD violations block CI (exit 1); SOFT warnings render in every report
format but never flip the exit code. New
GateResult/GateFindingtypes;evaluate_gate()is the shared severity-aware gate, whileevaluate_report()keeps its historical string contract. - Cyclical tool loop invariant (
[invariants] fail_on_identical_loops, default on): the same endpoint called ≥ 2 times with identical inputs and stagnant output state — even non-consecutively (A -> B -> A) — is a hard block.--allow-identical-loopsdisables it for a run. - Tool repeat cap invariant (
[invariants] max_tool_repeats, opt-in): hard cap on calls to any single endpoint;--max-tool-repeatsflag. - Path drift soft warning: a run that passes every budget but took an
alternate route renders a non-blocking "alternate valid route" note in
terminal, JSON (
warnings), markdown, and PR comments ([!NOTE]block). - Hard violations now render in terminal, JSON (
violations), markdown, and PR comments — a FAILED report shows exactly which rules fired. - Gate provenance (G7) and threshold-change flagging (G6) now cover the new
[invariants]knobs; config sections ignore unknown keys, so forward-compatibleagentdiff.tomlfiles load cleanly.
- Repository moved to the
kerrshiftorganization: canonical URLs now point togithub.com/kerrshift/agentdiff— package metadata (pyproject.tomlHomepage/Repository), README badge + clone link, the baseline-envelope schema$id, and the docs link inagentdiff initgenerated workflows. Reusable-action references are nowkerrshift/agentdiff/.github/actions/agentdiff-check; oldlostmartian/agentdiffURLs continue to work via GitHub redirects.
agentdiff initgenerated workflows used an invalid CLI form: the gate and PR-report steps invokedagentdiff <candidate> --baseline <baseline>, which the E1 diff-default patch parses asdiffwith a missing positional (Missing argument 'candidate_path') —--baselineis the baseline-store option, not the first positional argument. Generated workflows now invokeagentdiff diff <baseline> <candidate> --scenario <name>explicitly, with a regression test pinning the valid form.agentdiff init --with-approveruntime fixes (found by running the generated bot end-to-end on a live repo): pre-checkoutghcalls (pr view,run list,pr comment) now pass-R ${{ github.repository }}so they resolve outside a git checkout; the trigger filter uses nativeifexpressions (no${{ }}wrapper, which GitHub re-evaluates and mis-handles) plus a[bot]-login-suffix guard so the bot's own approval comment can never re-trigger itself.- Gate provenance misdescribed statistical runs (G7): in envelope mode the
Gate: …line showed the legacy single-run knobs (max_divergence=0.3,max_cost_delta=10.0%) whilecompare_envelopeactually judged with the scenario tolerances. The line now rendersGate [statistical envelope: <scenario>, N=<n>]: divergence_ceiling=…, max_cost_increase_pct=…, step_count_std_dev=…for envelope runs; strict/v1 baselines keep the legacy form.
AgentDiff v0.4.0 — trace capture, Goodhart-resistant gating, and a full docs
site redesign. Highlights: capture traces directly from any callable with
agentdiff record, every report self-describes its gate provenance, threshold
changes are flagged in the PR comment, stale baselines warn on --explain, and
the website gains a technical /features deep dive, /adapters, /action,
/compare, and /quickstart pages.
recordsubcommand (#23): capture execution traces from any Python callable (agentdiff record my_agent:run --input '{...}' --out traces/run.json), making baseline capture a one-liner instead of hand-instrumented serialization.- Gate provenance in every report (#25): terminal, JSON, and PR report formats self-describe the gate rules and their source, so a green check is always auditable against the thresholds that produced it.
- Threshold-change flagging (#24): when gate values in
agentdiff.tomldiffer from the baseline commit, the PR comment says so — above the diff — closing the Goodhart loop where thresholds drift to keep CI green. - Stale-baseline warning (#26):
--explainsurfaces the baseline's age past a configurable threshold, so re-recording is a deliberate act rather than silent decay. - Shell completion install instructions for bash/zsh/fish (#27).
- Website redesign (#28): new editorial
/featurespage with engine and metric diagrams, plus dedicated/adapters,/action,/compare, and/quickstartpages; landing sections rebuilt under a(site)route group.
0.3.0 - 2026-08-23
AgentDiff v0.3.0 — framework-native ingestion, suite-level gating, and recovery-effort metrics. Highlights: diff LangGraph and CrewAI artifacts natively (no OTel required), gate whole scenario suites in one call, benchmark two agents head-to-head, and gate on how expensive recovery from errors is.
- Recovery Step Ratio (RSR) — new metric quantifying how expensive it is
to get back on track after errors: the successful steps spent after each
ERROR/RETRY/ABANDONED cluster until the trajectory re-aligns with the
baseline path.
DiffReportgainsbaseline_recovery_steps,candidate_recovery_steps, andrecovery_step_ratio(additive, defaulted); surfaced insummary(), terminal, PR markdown (row appears when a threshold is set), and--explainfindings. Gateable via the opt-inassert_no_regressions(..., max_recovery_step_ratio=...)and CLI/config--max-recovery-ratio/[cli] max_recovery_ratio. - LangGraph adapter (
agentdiff.adapters.LangGraphAdapter, name"langgraph"): direct ingestion of native LangGraph artifacts — state snapshots, checkpoint dumps (channel_values), and message lists — in all three serialization shapes (message_to_dictdumps, LangChain constructor dumps, plain OpenAI-style role dicts). Tool decisions map to routing steps, tool results to tool_call steps (status honored), and the final AI answer to a response step; token usage read fromusage_metadata/response_metadata.token_usage. Participates inautodetection. - CrewAI adapter (
agentdiff.adapters.CrewAIAdapter, name"crewai"): direct ingestion of CrewAI kickoff output (CrewOutput.model_dump()) — per-task message logs map to fine-grained routing/tool/response steps prefixed by agent role, aggregatetoken_usagepopulates trace totals, and simplified exports without logs degrade to one step per task. Participates inautodetection. - Shared role-message engine (
adapters/_messages.py): one interpretation of OpenAI-style role conversations across all serialization shapes, powering the LangGraph and CrewAI adapters identically. - Adapter registry + plugin discovery (
agentdiff.adapters.register_adapter/available_adapters): custom adapters can be registered at runtime or via the standardagentdiff.adaptersentry-point group, and resolved by name throughload_trace(..., adapter_name=...)/[adapter] name = "...". Custom adapters may implement adetect(data) -> boolclassmethod to joinautodetection (built-ins always take priority). - Scenario runner (
agentdiff.run_scenarios): run multi-scenario regression suites programmatically. EachScenariopairs a baseline and a candidate (trace objects or file paths) with its ownGateThresholds; the resultingSuiteReportaggregates pass/fail/error per scenario with a CI-friendlysummary()— one broken trace or failing scenario never aborts the rest of the suite. - Parallel scenario execution:
run_scenarios(..., workers=N)executes scenarios on a thread pool (opt-in; sequential by default). Results are always returned in input order, and per-scenario errors stay contained. - A/B benchmark mode (
agentdiff.run_benchmark): head-to-head comparison of two agents on the same task across frameworks. Each case is scored on four deterministic efficiency dimensions (steps, wasted effort, tokens, latency); majority wins with explicit ties, plus a per-case side-by-side table (summary()/to_markdown()) and overall tally. Opt-in parallel execution viaworkers=N. Verdicts are structural-efficiency only — semantic quality remains a non-goal. - Performance benchmark suite (
make bench,benchmarks/): pytest-benchmark coverage for alignment, end-to-end compare, loop detection, recovery metrics, adapter parsing, and report serialization at 100/500/1000-step scales. Baseline finding: LCS alignment dominates runtime (~1.6s at 1000 steps); ingestion and metric computation are negligible. Excluded from the regular test suite; results are local (.benchmarks/gitignored). - Gate semantics are centralized in a pure
evaluate_report()shared byassert_no_regressionsand the scenario runner (single source of truth for current and future gates).
- Added
live_langgraph.pyandlive_crewai.py: real framework executions instrumented with OpenInference, diffed through theopeninferenceadapter — divergence, loop flags, and gate blocking on genuine graph/crew output. - Added
ingestion_langgraph.pyandingestion_crewai.py: offline direct- ingestion recipes against real captured fixtures (no OTel, no API keys), each diffing a regression variant of the fixture.
0.2.2 - 2026-08-19
agentdiff-checkGitHub Action gainsprandgithub-tokeninputs: whenpris set, the action runsagentdiff --pr <N>and posts the PR-ready report (status, gate thresholds, root-cause culprit, collapsed divergence tree, loops) as a comment on the pull request. The comment is posted even when the gate blocks (exit 1).
0.2.1 - 2026-08-19
Patch release. Every change below was found by live-testing the adapters against real SDK/API output and fixed so the adapters ingest real traces natively (no caller-side reshaping).
- openai_agents: skip
task/turnbookkeeping spans as the real SDK emits them (typecustom+ nametask/turn, not a dedicated span type). - langfuse: map the v4 SDK's
AGENT/CHAIN/TOOL/RETRIEVERobservation types (previously they fell through tothought); sort observations chronologically by start time; accept snake_case SDK keys (start_time,parent_observation_id,status_message,latency). - openinference: accept real OTel values natively — integer span/trace ids are stringified, and nanosecond timestamps are auto-detected (the legacy shape used seconds).
- testing:
assert_no_regressionsnow marks the reportpassed=Falseon failure, so a printedsummary()no longer misleadingly shows "PASSED" after a blocked gate.
- Added live recipes that ingest real traces through each adapter:
live_openai_agents.py(OpenAI Agents SDK),live_openinference.py(OpenInference/OTel),live_langfuse.py(Langfuse Cloud, v4 SDK).
- README/cookbook guides updated for the live adapter recipes.
0.2.0 - 2026-08-19
AgentDiff v0.2.0 — the first PyPI release of a developer-first trajectory regression engine for multi-turn, tool-using AI agents.
- Trajectory diff engine: structurally compares two agent execution DAGs, aligning steps to compute the Trajectory Divergence Index (TDI), Wasted Effort Index (WEI), resource/token deltas, and tool-loop detection.
- Local-first CI/CD gate:
assert_no_regressions(...)forpytest, a CLI with exit codes, and a composite GitHub Action (agentdiff-check). - Telemetry adapters:
generic,openinference,langfuse,langsmith, andopenai_agents— withautodetection from raw JSON/OTel files. - CLI:
agentdiff baseline.json candidate.jsonwith--format(terminal/json/markdown/pr),--fail-on-regression,--max-divergence,--max-loops,--max-cost-delta,--baseline/-b,--update-baseline, baseline rotation (--baseline-rotation/--max-drift), and--prPR report generation. - Config-as-code:
agentdiff.toml(load_config/find_config_file), with CLI flags overriding config. - Reporters: terminal (rich), JSON, Markdown, PR reports, tree rendering, and human-readable change explanations / culprit location.
- Public Python API:
compare,assert_no_regressions,load_trace,parse_trace_data,load_config,decide_rotation,generate_pr_markdown, plus canonical models (AgentTrace,DiffReport,StepDiff, ...). - pytest plugin:
agentdiffmarker/entry-point for threshold gating.
- Python 3.10 ISO timestamp parsing: telemetry timestamps with a short
fractional second (e.g.
...00.5Z) now parse correctly across all adapters, instead of silently dropping latency to0.0.
- Full docs site (markdown) covering adapters, CLI, regression gates, CI/CD, configuration, and baseline workflows.
cookbooks/with ingestion examples for each telemetry adapter.
- Requires Python 3.10+.
- Licensed under MIT.