Skip to content

feat(evals): add eight TRAIL benchmark tasks and seed annotations reliably - #16655

Draft
ehutt wants to merge 4 commits into
mainfrom
ehutt/trail-benchmark-new-tasks
Draft

ehutt wants to merge 4 commits into
mainfrom
ehutt/trail-benchmark-new-tasks

Conversation

@ehutt

@ehutt ehutt commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Before: The TRAIL benchmark asked 10 questions about the research-assistant project. None of them covered latency, time windows, trace-level annotation scores, annotation score filters, span hierarchy, or questions that need several lookups chained together. Each fixture seed also silently dropped a different random subset of annotations (one seed kept 495 of 581 trail_errors, another kept 514).

After: The benchmark has 8 more tasks, ranging from easy to hard. A fresh seed keeps all 581 trail_error span annotations and all 585 trace scores.

Task Level Tests Reference
llm-latency-p95 easy latency percentiles 27.5 s
avg-reliability easy trace-level annotations 2.39
slowest-tool easy per-tool latency VisitTool, ~55.7 s per call
busiest-window medium 10-minute time buckets 16:40–16:50 UTC, 67 traces
high-impact-traces medium filtering annotations by score 96 of 117 traces
llm-call-failure medium finding one error among 1,508 LLM spans prompt flagged for usage policy
subagent-delegation medium span hierarchy 47 traces
costliest-trace-question hard cost by trace, then root-span input the Carl Nebel citation question

Four of these (p95 LLM duration, 10-minute buckets, high-impact errors, average reliability) come from the questions in the Phoenix MCP SQL code-mode blog post that the benchmark did not cover.

How: Each task follows the existing layout and grader. Its solve.py reads Phoenix through evals.harbor.verifiers.phoenix_api. The Python client cannot read trace annotations, so a new trace_annotation_scores helper reads them through REST. In scripts/load_patronus_trail.py, the loader now posts annotations after every span is queryable. Before, it posted them right after each trace's spans, and Phoenix dropped annotations whose span or trace it had not inserted yet.

Validation: -a oracle -e docker on the 8 new tasks plus top-error-category, against a fixture reseeded with the fix: 9 of 9 trials scored 1.0. pytest tests/unit/harbor passes.

Existing staged fixtures need RESEED=1 make harbor-stage before avg-reliability and high-impact-traces will match.

…iably

New tasks cover latency percentiles, trace-level annotation scores, time
bucketing, annotation score filters, a single failing LLM span, sub-agent
hierarchy, and a cost-to-input multi-hop question.

The TRAIL loader posted annotations before Phoenix had ingested their spans,
so each seed dropped a different subset. It now posts them after every span
is queryable.
@mintlify

mintlify Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated
arize-phoenix 🟢 Ready View Preview Oct 1, 2026, 12:42 AM

💡 Tip: Enable Automations to automatically generate PRs for you.

This branch was successfully deployed

1 active (outdated) deployment
staging — f83cadd7 Deployed Oct 1, 2026 by mintlify[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: 📘 Todo

Development

Successfully merging this pull request may close these issues.

1 participant