Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,15 @@
> Current package stable line: v6.3.0.
> v4.2.0 remains the previous plugin SDK stable release, v4.1.0 remains the previous enterprise governance stable release, and v2.4.1 remains the previous P95 stable release. Ecosystem maturity is tracked separately from P98 controlled maturity.

## Unreleased

- Added `failure-doctor heal`, a conservative Playwright locator repair loop that joins diagnosis, one exact source edit, a real rerun, and before/after evidence.
- Added automatic source backup and rollback when the verification command fails.
- Added explicit `verified`, `failed_rolled_back`, `failed_patch_retained`, and `manual_restore_required` states instead of treating an edited file as a successful repair.
- Blocked unsupported failure types, ambiguous matches, path escapes, symlink targets, and dirty Git targets unless the caller explicitly opts in.
- Refocused the default CLI and README first screen on diagnosis, verified repair, the offline benchmark, and the local console; specialist tracks remain available under `advanced`.
- Added unit and CLI coverage for successful reruns, failed reruns, automatic rollback, ambiguity blocking, and diagnosis eligibility.

## v6.3.0

- Added a simplified default CLI with `diagnose`, `bench`, `console`, `doctor`, and an explicit `advanced` entry for the full legacy command set.
Expand Down
93 changes: 38 additions & 55 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,35 +1,48 @@
# Agent Failure Doctor
# Agent Failure Doctor

[中文文档](README.zh-CN.md)

![CI](https://github.com/tobybgy-lsd/web-agent-runtime-bench/actions/workflows/benchmark.yml/badge.svg)
![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)
![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue.svg)

Local-first failure diagnosis lifecycle tool for AI browser automation,
Playwright, crawler, RPA, and business automation failures.
Diagnose Playwright failures locally, apply one scoped locator repair, and keep
the change only when the same real test command passes.

**Input:** trace.zip / error.log / console.txt / network.json / probe_report.json / screenshot metadata / user_description.txt / visual_run / OCR or document evidence.
The public core deliberately does two things:

**Output:** diagnosis, evidence, next action, repair suggestions, GitHub issue draft, Codex fix prompt.
1. Turn a trace, log, screenshot, or failed-run directory into an evidence report.
2. Repair an explicitly supplied stale locator, rerun the real test, and emit a
truthful `verified`, `failed_rolled_back`, or `manual_restore_required` result.

```powershell
git clone https://github.com/tobybgy-lsd/web-agent-runtime-bench.git
cd web-agent-runtime-bench
python -m pip install agent-failure-doctor
failure-doctor diagnose .\examples\failed_runs\proxy_network_error --out .\report
failure-doctor plan .\report --out .\fix_plan
failure-doctor bench
```

**Lifecycle commands:** `diagnose` / `plan` / `verify` / `run`.
**Classic lifecycle:** diagnose -> plan -> AI handoff / patch proposal -> verify -> sanitize/share.
**Share/adapt commands:** `sanitize` / `adapt`.
**Bootstrap command:** `failure-doctor agent-bootstrap`.
**Patch command:** `failure-doctor propose-patch`.
**Fleet command:** `failure-doctor batch`.
Compatibility track: Agent Failure Doctor v4.2.0 Plugin SDK & Adapter Ecosystem Pack.
## Verified locator repair

```powershell
failure-doctor heal .\test-results\failed-case `
--project . `
--target tests\checkout.spec.ts `
--old-locator "button.old-submit" `
--new-locator "button[data-testid=submit]" `
--out .\heal-report `
--test-command npx playwright test tests\checkout.spec.ts
```

The command writes the diagnosis, exact diff, source backup, sanitized command
logs, exit code, and verification report. It blocks ambiguous or unsupported
repairs. A non-zero rerun restores the original file by default. It never marks
a proposal as repaired merely because a file was edited. See
[Verified repair contract](docs/VERIFIED_HEAL.md).

Current milestone: Agent Failure Doctor v6.3 Lite UX & Bench Core Release.
Unreleased focus: verified Playwright locator repair with automatic rollback.
Current stable line: v6.3.0.
Previous stable line: Agent Failure Doctor v6.2 Local RPA Ops & Browser Backend Release.
Earlier stable line: Agent Failure Doctor v6.1 Composite Diagnosis Runtime Triage Release.
Expand All @@ -47,13 +60,14 @@ Install only the local diagnosis and benchmark core:
python -m pip install agent-failure-doctor
failure-doctor doctor
failure-doctor diagnose .\examples\failed_runs\proxy_network_error --out .\report
failure-doctor heal --help
failure-doctor bench
failure-doctor console
```

The default CLI promotes only five entries: `diagnose`, `bench`, `console`,
`doctor`, and `advanced`. The existing full command set remains available under
`failure-doctor advanced` and through its original command names.
The default CLI promotes only six entries: `diagnose`, `heal`, `bench`,
`console`, `doctor`, and `advanced`. The existing full command set remains
available under `failure-doctor advanced` and through its original command names.

Install optional capabilities only where they run:

Expand All @@ -64,49 +78,18 @@ python -m pip install "agent-failure-doctor[ocr-paddle]"
python -m pip install "agent-failure-doctor[enterprise]"
```

Focused console modes are available with `--mode lite`, `diagnose`, `bench`,
or `full`. The built-in `lite-core` benchmark is offline and uses packaged,
sanitized failure artifacts; it does not access real commerce or ERP systems.
Challenge/CAPTCHA cases are detection and manual-handoff tests, not bypasses.

Classic quickstart: `failure-doctor diagnose .\examples\failed_runs\proxy_network_error --out .\report`; `failure-doctor plan .\report --out .\fix_plan`; `failure-doctor propose-patch --repo . --report .\report --out .\patch_plan`; `failure-doctor agent-bootstrap --target all --project .`.

Earlier stable line: Agent Failure Doctor v4.2.0 Plugin SDK & Adapter Ecosystem Pack.
Previous P95 stable: v2.4.1.
The built-in `lite-core` benchmark is offline and uses packaged, sanitized
failure artifacts. Challenge/CAPTCHA cases are detection and manual-handoff
tests, not bypasses.

**Lifecycle commands:** `diagnose` / `plan` / `verify` / `run`.

**Classic lifecycle:** diagnose -> plan -> AI handoff / patch proposal -> verify -> sanitize/share.

**Key commands:** `failure-doctor propose-patch`; `failure-doctor batch`; `sanitize` / `adapt`.

Optional v6.0 output: sanitized public cases, benchmark reports, adapter reports, Android APK UI evidence reports, Android Pro hardening reports, Android Ops reports, Android authoring reports, Android pilot reports, Android deep diagnostic bundles, Android playbooks, Android device-lab reports, mobile-stability reports, deployment health reports, stability reports, plugin validation reports, evidence-bound reasoning, local web console, and CI/CD gate.

- Current milestone: Agent Failure Doctor v6.3 Lite UX & Bench Core Release
- Current stable line: v6.3.0
- Previous stable line: Agent Failure Doctor v6.2.0 Local RPA Ops & Browser Backend Release
- Previous stable line: Agent Failure Doctor v6.1.0 Composite Diagnosis Runtime Triage Release
- Previous stable line: Agent Failure Doctor v6.0.0 Mobile Automation Stable Standardization Release
- Previous stable line: Agent Failure Doctor v5.3.0 Android Real Device Farm & Business Workflow Operations Pack
- Previous stable line: Agent Failure Doctor v5.2.0 Android APK Production Hardening & Workflow Template Pack
- Previous stable line: Agent Failure Doctor v5.1.0 Android APK UI Automation Adapter Pack
- Previous stable line: Agent Failure Doctor v5.0.0 Stable API / Schema / Plugin ABI Standardization Release
- Earlier stable line: Agent Failure Doctor v4.3.0 Real User Case Program & Public Benchmark Pack
- Previous stable line: Agent Failure Doctor v4.2.0 Plugin SDK & Adapter Ecosystem Pack
- Earlier stable line: Agent Failure Doctor v4.1.0 Enterprise Governance & Role-Based Console Pack
- Earlier stable line: Agent Failure Doctor v4.0.0 Hybrid Evidence Reasoning Pack
- Earlier stable line: Agent Failure Doctor v3.9.0 Local Failure Knowledge Base Pack (v3.9 Local Failure Knowledge Base Pack)
- Previous P95 stable line: Agent Failure Doctor v2.4.1 P95 Alignment & Missing Tracks Pack

**Classic lifecycle:** diagnose -> plan -> AI handoff / patch proposal -> verify -> sanitize/share.
Compatibility track: Agent Failure Doctor v4.2.0 Plugin SDK & Adapter Ecosystem Pack.
Previous P95 stable: v2.4.1.

**Patch command:** `failure-doctor propose-patch`.

**Share/adapt commands:** `sanitize` / `adapt`.

**Fleet command:** `failure-doctor batch`.

**Core commands:** `diagnose` / `plan` / `verify` / `run`; `android`; `android-pro`; `android-ops`; `android-author`; `android-pilot`; `android-dx`; `android-playbook`; `android-real-pilot`; `android-lab`; `mobile-stability`; `case`; `issue-pack`; `benchmark`; `adapter`; `deploy`; `stability`; `plugin`; `reason` / `root-cause` / `causal-chain`; `agent-bootstrap`; `sanitize` / `adapt`; `ocr-evidence`; `visual-runtime`; `regulated-eval`; `full-chain-eval`; `console`; `ci`; `kb`; `failure-doctor propose-patch`; `failure-doctor batch`.
Specialist Android, OCR, enterprise, adapter, and LAN Worker capabilities are
optional compatibility tracks. They are intentionally absent from the default
help; use `failure-doctor advanced` when you actually operate those components.

### Unreleased LAN Worker v1

Expand Down
29 changes: 26 additions & 3 deletions README.zh-CN.md
Original file line number Diff line number Diff line change
@@ -1,20 +1,43 @@
# Agent Failure Doctor 中文文档
# Agent Failure Doctor 中文文档

[English documentation](README.md)

## 这个项目现在只突出两件事

1. 把 Playwright 的 trace、日志、截图或失败目录整理成有证据的诊断报告。
2. 对一个明确的失效定位器做最小修改,真实重跑同一条测试命令;只有重跑通过才保留修改。

```powershell
failure-doctor diagnose .\test-results\failed-case --out .\report

failure-doctor heal .\test-results\failed-case `
--project . `
--target tests\checkout.spec.ts `
--old-locator "button.old-submit" `
--new-locator "button[data-testid=submit]" `
--out .\heal-report `
--test-command npx playwright test tests\checkout.spec.ts
```

`heal` 会保存原文件、补丁、命令日志和退出码。重跑失败默认自动回滚;定位器
出现在多个文件、目标文件已有未提交修改、错误类型不属于定位器问题时会直接
阻断,不会显示假成功。详见 [修复验证契约](docs/VERIFIED_HEAL.md)。

## v6.3 快速入口

只安装本地诊断与基准核心:

```powershell
python -m pip install agent-failure-doctor
failure-doctor doctor
failure-doctor diagnose .\examples\failed_runs\proxy_network_error --out .\report
failure-doctor heal --help
failure-doctor bench
failure-doctor console
```

默认 CLI 只突出 `diagnose`、`bench`、`console`、`doctor` 和 `advanced`
五个入口。原有完整命令没有删除,可通过 `failure-doctor advanced` 查看,
默认 CLI 只突出 `diagnose`、`heal`、`bench`、`console`、`doctor` 和 `advanced`
六个入口。原有完整命令没有删除,可通过 `failure-doctor advanced` 查看,
也仍可直接调用原命令。

Web、Android、PaddleOCR 和企业集成按运行机器选装:
Expand Down
51 changes: 51 additions & 0 deletions docs/VERIFIED_HEAL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
# Verified Playwright Locator Repair

`failure-doctor heal` joins diagnosis, one exact source edit, a real rerun, and
before/after evidence into one conservative command.

## Contract

The command will modify source only when all of these are true:

- The failure report is `selector_drift`, `playwright_strict_mode_violation`, or
`selector_syntax_error`.
- The target is a supported Python or JavaScript/TypeScript source file inside
the authorized project.
- The old locator occurs exactly once in the selected file.
- The target is not a symlink.
- A Git target has no pre-existing uncommitted change unless the caller passes
`--allow-dirty-target` explicitly.
- A real verification command is supplied with `--test-command`.

## Result states

| Status | Meaning |
| --- | --- |
| `planned` | Dry run only; no source was changed. |
| `verified` | The exact patch was applied, the real command exited `0`, and the command did not alter the patched source. |
| `failed_rolled_back` | The real command failed and the original source was restored. |
| `failed_patch_retained` | The real command failed and `--keep-failed-patch` explicitly retained the edit. |
| `manual_restore_required` | The test command changed the target source, so automatic rollback was blocked to avoid overwriting concurrent work. |

Only `verified` means the repair passed its execution contract. A generated
patch, a diagnosis, or an edited file is never reported as a successful repair.

## Example

```powershell
failure-doctor heal .\test-results\checkout-failure `
--project . `
--target tests\checkout.spec.ts `
--old-locator "button.old-submit" `
--new-locator "button[data-testid=submit]" `
--out .\outputs\checkout-heal `
--test-command npx playwright test tests\checkout.spec.ts
```

Review `heal_report.json`, `heal_report.md`, `locator.patch`, `backup/`, and
`runs/after-rerun/`. Run with `--dry-run` first when the replacement needs human
review.

This command does not solve challenges, bypass access controls, change browser
fingerprints, extract credentials, or infer a replacement selector from private
production pages.
78 changes: 77 additions & 1 deletion failure_doctor/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,7 @@
from failure_doctor.deploy.cli import add_deploy_parser, handle_deploy
from failure_doctor.stability.cli import add_stability_parser, handle_stability
from failure_doctor.lite import capability_report, default_benchmark_output, print_capability_report
from failure_doctor.heal import HealBlocked, execute_verified_locator_repair
from failure_doctor.visual_runtime.adapter import adapt_visual_artifacts
from failure_doctor.visual_runtime.compare import compare_visual_runs
from failure_doctor.visual_runtime.loader import load_visual_run, validate_visual_run
Expand Down Expand Up @@ -90,6 +91,8 @@ def main(argv: list[str] | None = None) -> int:
args = parser.parse_args(raw_args)
if args.command == "diagnose":
return diagnose_inputs(args)
if args.command == "heal":
return heal_inputs(args)
if args.command == "plan":
return plan_from_report(args)
if args.command == "verify":
Expand Down Expand Up @@ -216,6 +219,33 @@ def build_parser() -> argparse.ArgumentParser:
diagnose.add_argument("--plugin", default=None, help="Optional enabled diagnosis-rule plugin id")
diagnose.add_argument("--plugins", default=".failure-doctor-plugins", help="Plugin workspace")
diagnose.add_argument("--adapter", default=None, choices=["android-apk"], help="Optional specialized evidence adapter")
heal = sub.add_parser(
"heal",
help="Apply one scoped Playwright locator repair and prove it by rerunning a real command",
)
heal.add_argument("input", help="Failure artifact or existing diagnosis report directory")
heal.add_argument("--project", required=True, help="Authorized Playwright project directory")
heal.add_argument("--old-locator", required=True, help="Exact stale locator text in source")
heal.add_argument("--new-locator", required=True, help="Exact replacement locator text")
heal.add_argument("--target", default=None, help="Optional source file relative to project")
heal.add_argument("--out", required=True, help="Output directory for backup and verification evidence")
heal.add_argument("--dry-run", action="store_true", help="Write the proposed patch without changing source")
heal.add_argument(
"--allow-dirty-target",
action="store_true",
help="Allow modifying a source file that already has uncommitted changes",
)
heal.add_argument(
"--keep-failed-patch",
action="store_true",
help="Keep the patch when verification fails instead of restoring the backup",
)
heal.add_argument(
"--test-command",
nargs=argparse.REMAINDER,
required=True,
help="Real command and its arguments, for example --test-command npx playwright test tests/cart.spec.ts",
)
plan = sub.add_parser("plan", help="Generate a fix plan from a diagnosis report directory")
plan.add_argument("report", help="Path to a report directory containing diagnosis.json")
plan.add_argument("--out", required=True, help="Output fix plan directory")
Expand Down Expand Up @@ -477,12 +507,13 @@ def build_parser() -> argparse.ArgumentParser:

def print_lite_help() -> None:
print(
"""usage: failure-doctor {diagnose,bench,console,doctor,advanced} ...
"""usage: failure-doctor {diagnose,heal,bench,console,doctor,advanced} ...

Local-first Web Agent benchmark and failure diagnosis.

quick commands:
diagnose Diagnose a trace, log, screenshot, or failure directory
heal Patch one Playwright locator, rerun, and keep it only when verified
bench Run the built-in evidence-based Web Agent benchmark
console Open the local Web console
doctor Check optional capability profiles without network access
Expand All @@ -493,6 +524,51 @@ def print_lite_help() -> None:
)


def heal_inputs(args: argparse.Namespace) -> int:
out_dir = Path(args.out)
input_path = Path(args.input)
before_report = out_dir / "before_report"
try:
has_existing_diagnosis = input_path.is_dir() and (input_path / "diagnosis.json").is_file()
if has_existing_diagnosis:
diagnosis = _load_report_diagnosis(input_path)
else:
diagnose_args = argparse.Namespace(
input=str(input_path),
out=str(before_report),
run_id=None,
kb=None,
hybrid_reasoning=False,
reasoner="mock_reasoner",
plugin=None,
plugins=".failure-doctor-plugins",
adapter=None,
)
if diagnose_inputs(diagnose_args) != 0:
return 2
diagnosis = _load_report_diagnosis(before_report)
report = execute_verified_locator_repair(
project=Path(args.project),
diagnosis=diagnosis,
old_locator=str(args.old_locator),
new_locator=str(args.new_locator),
target=Path(args.target) if args.target else None,
test_command=list(args.test_command),
out_dir=out_dir,
dry_run=bool(args.dry_run),
allow_dirty_target=bool(args.allow_dirty_target),
keep_failed_patch=bool(args.keep_failed_patch),
)
except (HealBlocked, FileNotFoundError, OSError, ValueError) as exc:
print(f"Heal blocked: {exc}")
return 2
print("Failure Doctor Verified Repair")
print(f"Status: {report.get('status')}")
print(f"Verified: {str(bool(report.get('verified'))).lower()}")
print(f"Output: {out_dir}")
return 0 if report.get("verified") or report.get("status") == "planned" else 1


def diagnose_inputs(args: argparse.Namespace) -> int:
_configure_stdio()
input_path = Path(args.input)
Expand Down
Loading
Loading