Part of #4, which says why: the pack documents installing on Gemini CLI, and every published number was measured on Claude Code. The seam is in: run.sh asks, one script per tool answers, and evals/README.md, "Measuring another agent", is the contract such a script keeps. This issue is one runner.
What to build
One executable file, evals/agents/gemini-cli.sh, that turns one Gemini CLI headless run into the stream the scorer reads. It is called as gemini-cli.sh ARM MODEL PROMPT inside a throwaway git repository; evals/agents/claude.sh next to it is the reference. The pack must be installed the way docs/other-agents.md describes (the skills in ~/.agents/skills/, the routing block in ~/.gemini/GEMINI.md), because what is measured is whether the installed pack switches on by itself.
The one judgment in the file
The scorer counts one Skill line for every time the agent opened one of the pack's skills. Say in the runner's header what that is on Gemini CLI, the call that activates a skill or a read of its SKILL.md, whichever the transcript shows, and convert exactly that. A number in results.md will mean what your header says.
How to run it
EVAL_AGENT=gemini-cli EVAL_MODEL=<a model id Gemini CLI accepts> EVAL_ONLY="activation negatives" bash evals/run.sh
then python3 evals/score.py --print ~/.claude/open-steps/evals/<today>. Your agent appears as its own column, labelled gemini-cli:<model> until models.md has your row. The quality arms and the premortem are out of scope here: they lean on Claude Code's --allowedTools and on a fresh subagent. The runner may exit 2 for with and without, with one line on stderr.
Done when
bash evals/test.sh passes with one new CASE: a fixture stream under evals/fixtures/ in your runner's shape that the scorer labels by your models.md row.
shellcheck evals/agents/gemini-cli.sh is clean.
- The pull request carries the runner, the
models.md row, the fixture and its test, and the day's transcripts handed over outside git (a zip attached to the pull request, or a link); the maintainer scores them and writes results.md.
- A number that was not measured is absent, never estimated. An arm that could not run is listed as not run.
Not in scope
No AI judging, no change to cases.md, no transcripts in the repository, no runner for a second tool in the same pull request.
Cost, on your own account
The activation and off-topic phrases are 28 runs per repetition, and N_RUNS=3 is the published setting, so 84 short runs. Note in the pull request which model and which plan you ran on. When the hooks were tested on Gemini CLI, the free tier's daily quota ran out after two sessions (docs/other-agents.md), so plan for more than one day.
Part of #4, which says why: the pack documents installing on Gemini CLI, and every published number was measured on Claude Code. The seam is in:
run.shasks, one script per tool answers, andevals/README.md, "Measuring another agent", is the contract such a script keeps. This issue is one runner.What to build
One executable file,
evals/agents/gemini-cli.sh, that turns one Gemini CLI headless run into the stream the scorer reads. It is called asgemini-cli.sh ARM MODEL PROMPTinside a throwaway git repository;evals/agents/claude.shnext to it is the reference. The pack must be installed the waydocs/other-agents.mddescribes (the skills in~/.agents/skills/, the routing block in~/.gemini/GEMINI.md), because what is measured is whether the installed pack switches on by itself.The one judgment in the file
The scorer counts one
Skillline for every time the agent opened one of the pack's skills. Say in the runner's header what that is on Gemini CLI, the call that activates a skill or a read of itsSKILL.md, whichever the transcript shows, and convert exactly that. A number inresults.mdwill mean what your header says.How to run it
then
python3 evals/score.py --print ~/.claude/open-steps/evals/<today>. Your agent appears as its own column, labelledgemini-cli:<model>untilmodels.mdhas your row. The quality arms and the premortem are out of scope here: they lean on Claude Code's--allowedToolsand on a fresh subagent. The runner may exit 2 forwithandwithout, with one line on stderr.Done when
bash evals/test.shpasses with one new CASE: a fixture stream underevals/fixtures/in your runner's shape that the scorer labels by yourmodels.mdrow.shellcheck evals/agents/gemini-cli.shis clean.models.mdrow, the fixture and its test, and the day's transcripts handed over outside git (a zip attached to the pull request, or a link); the maintainer scores them and writesresults.md.Not in scope
No AI judging, no change to
cases.md, no transcripts in the repository, no runner for a second tool in the same pull request.Cost, on your own account
The activation and off-topic phrases are 28 runs per repetition, and
N_RUNS=3is the published setting, so 84 short runs. Note in the pull request which model and which plan you ran on. When the hooks were tested on Gemini CLI, the free tier's daily quota ran out after two sessions (docs/other-agents.md), so plan for more than one day.