Skip to content

A runner for Gemini CLI, so activation has numbers there too #38

Description

@kharmanskyi

Part of #4, which says why: the pack documents installing on Gemini CLI, and every published number was measured on Claude Code. The seam is in: run.sh asks, one script per tool answers, and evals/README.md, "Measuring another agent", is the contract such a script keeps. This issue is one runner.

What to build

One executable file, evals/agents/gemini-cli.sh, that turns one Gemini CLI headless run into the stream the scorer reads. It is called as gemini-cli.sh ARM MODEL PROMPT inside a throwaway git repository; evals/agents/claude.sh next to it is the reference. The pack must be installed the way docs/other-agents.md describes (the skills in ~/.agents/skills/, the routing block in ~/.gemini/GEMINI.md), because what is measured is whether the installed pack switches on by itself.

The one judgment in the file

The scorer counts one Skill line for every time the agent opened one of the pack's skills. Say in the runner's header what that is on Gemini CLI, the call that activates a skill or a read of its SKILL.md, whichever the transcript shows, and convert exactly that. A number in results.md will mean what your header says.

How to run it

EVAL_AGENT=gemini-cli EVAL_MODEL=<a model id Gemini CLI accepts> EVAL_ONLY="activation negatives" bash evals/run.sh

then python3 evals/score.py --print ~/.claude/open-steps/evals/<today>. Your agent appears as its own column, labelled gemini-cli:<model> until models.md has your row. The quality arms and the premortem are out of scope here: they lean on Claude Code's --allowedTools and on a fresh subagent. The runner may exit 2 for with and without, with one line on stderr.

Done when

  • bash evals/test.sh passes with one new CASE: a fixture stream under evals/fixtures/ in your runner's shape that the scorer labels by your models.md row.
  • shellcheck evals/agents/gemini-cli.sh is clean.
  • The pull request carries the runner, the models.md row, the fixture and its test, and the day's transcripts handed over outside git (a zip attached to the pull request, or a link); the maintainer scores them and writes results.md.
  • A number that was not measured is absent, never estimated. An arm that could not run is listed as not run.

Not in scope

No AI judging, no change to cases.md, no transcripts in the repository, no runner for a second tool in the same pull request.

Cost, on your own account

The activation and off-topic phrases are 28 runs per repetition, and N_RUNS=3 is the published setting, so 84 short runs. Note in the pull request which model and which plan you ran on. When the hooks were tested on Gemini CLI, the free tier's daily quota ran out after two sessions (docs/other-agents.md), so plan for more than one day.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestevalsThe measurement harness and the published numbershelp wantedExtra attention is neededother agentsCodex, Cursor, Gemini CLI and anything past Claude Code

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions