Part of #4, which says why: the pack documents installing on Cursor, and every published number was measured on Claude Code. The seam is in: run.sh asks, one script per tool answers, and evals/README.md, "Measuring another agent", is the contract such a script keeps. This issue is one runner.
What to build
One executable file, evals/agents/cursor.sh, that turns one Cursor CLI print-mode run (agent -p in docs/other-agents.md) into the stream the scorer reads. It is called as cursor.sh ARM MODEL PROMPT inside a throwaway git repository; evals/agents/claude.sh next to it is the reference. The pack must be installed the way docs/other-agents.md describes (the skills in ~/.agents/skills/, the routing block in AGENTS.md at the project root, which your runner writes into the throwaway repository), because what is measured is whether the installed pack switches on by itself.
The one judgment in the file
The scorer counts one Skill line for every time the agent opened one of the pack's skills. Say in the runner's header what that is on Cursor, most likely a read of the skill's SKILL.md, or whatever the transcript shows, and convert exactly that. A number in results.md will mean what your header says.
How to run it
EVAL_AGENT=cursor EVAL_MODEL=<a model id Cursor accepts> EVAL_ONLY="activation negatives" bash evals/run.sh
then python3 evals/score.py --print ~/.claude/open-steps/evals/<today>. Your agent appears as its own column, labelled cursor:<model> until models.md has your row. The quality arms and the premortem are out of scope here: they lean on Claude Code's --allowedTools and on a fresh subagent. The runner may exit 2 for with and without, with one line on stderr.
Done when
bash evals/test.sh passes with one new CASE: a fixture stream under evals/fixtures/ in your runner's shape that the scorer labels by your models.md row.
shellcheck evals/agents/cursor.sh is clean.
- The pull request carries the runner, the
models.md row, the fixture and its test, and the day's transcripts handed over outside git (a zip attached to the pull request, or a link); the maintainer scores them and writes results.md.
- A number that was not measured is absent, never estimated. An arm that could not run is listed as not run.
Not in scope
No AI judging, no change to cases.md, no transcripts in the repository, no runner for a second tool in the same pull request.
Cost, on your own account
The activation and off-topic phrases are 28 runs per repetition, and N_RUNS=3 is the published setting, so 84 short runs. Note in the pull request which model and which plan you ran on. Cursor needs a paid plan; say which you used.
Part of #4, which says why: the pack documents installing on Cursor, and every published number was measured on Claude Code. The seam is in:
run.shasks, one script per tool answers, andevals/README.md, "Measuring another agent", is the contract such a script keeps. This issue is one runner.What to build
One executable file,
evals/agents/cursor.sh, that turns one Cursor CLI print-mode run (agent -pindocs/other-agents.md) into the stream the scorer reads. It is called ascursor.sh ARM MODEL PROMPTinside a throwaway git repository;evals/agents/claude.shnext to it is the reference. The pack must be installed the waydocs/other-agents.mddescribes (the skills in~/.agents/skills/, the routing block inAGENTS.mdat the project root, which your runner writes into the throwaway repository), because what is measured is whether the installed pack switches on by itself.The one judgment in the file
The scorer counts one
Skillline for every time the agent opened one of the pack's skills. Say in the runner's header what that is on Cursor, most likely a read of the skill'sSKILL.md, or whatever the transcript shows, and convert exactly that. A number inresults.mdwill mean what your header says.How to run it
then
python3 evals/score.py --print ~/.claude/open-steps/evals/<today>. Your agent appears as its own column, labelledcursor:<model>untilmodels.mdhas your row. The quality arms and the premortem are out of scope here: they lean on Claude Code's--allowedToolsand on a fresh subagent. The runner may exit 2 forwithandwithout, with one line on stderr.Done when
bash evals/test.shpasses with one new CASE: a fixture stream underevals/fixtures/in your runner's shape that the scorer labels by yourmodels.mdrow.shellcheck evals/agents/cursor.shis clean.models.mdrow, the fixture and its test, and the day's transcripts handed over outside git (a zip attached to the pull request, or a link); the maintainer scores them and writesresults.md.Not in scope
No AI judging, no change to
cases.md, no transcripts in the repository, no runner for a second tool in the same pull request.Cost, on your own account
The activation and off-topic phrases are 28 runs per repetition, and
N_RUNS=3is the published setting, so 84 short runs. Note in the pull request which model and which plan you ran on. Cursor needs a paid plan; say which you used.