Watch · Error-analysis workbench
Error-analysis workbench
Read your agent runs one at a time and mark each pass or fail. Open a run, inspect its span waterfall, tag a failure mode with a one-line critique, and watch the failure modes accrete into a frequency table that tells you what to fix first.
Where to find it
- Localhost:
/agent-workbench.html?repo=<id> - API:
GET /api/agent/quality?repo=<id>(runs),GETandPOST /api/agent/labels(labels),/api/agent/failure-clusters(taxonomy),/api/agent/traces/<sessionId>(waterfall),/api/loop/judge-alignment(judge panel) - Keyboard: ⌘ K then type
workbench - Navigation: Remediation in the sidebar, then Error-analysis workbench under All tools, in the Approvals and workflows group. On repoops.ai the same group opens
/team/error-workbench.
What it does for you
Runs, priciest first.The center column lists your agent sessions sorted by cost, with the source tool, tool count, and any PR, CI, or revert chips. Reading the expensive runs first is how you find what is worth fixing.
Span waterfall per run.Click a run to see its trace as a nested waterfall (session, turn, tool), each span sized by cost when it is priced and by token count otherwise, so you can spot the tool call or turn that ran up the bill.
Label pass or fail in one keystroke.Press
p for pass, f for fail, write a one-line critique, then press Cmd or Ctrl plus Enter to save from the critique box (a plain Enter saves when the box is not focused). On a fail the failure-mode list starts on ci-fail, so set it before you save.Failure modes rank themselves.Your labels accrete bottom-up into a "fix this first" taxonomy, ranked by frequency, with a coverage line (how many sessions are labeled) that marks the sample as thin below 5 scored sessions or under 20 percent coverage.
Judge alignment.The panel compares the automated judge with you on the runs that carry both a human label and a quality score: TPR, TNR, accuracy, and a confusion matrix. Until there are such runs it says so and asks for more labels.
Save a run as a regression test.The Save as regression test button records the shape of a run's trace (its tool-usage signature) as a case in the
routing-smoke golden dataset under docs/evals/, the dataset the eval gate scores routing proposals against.Configure
The taxonomy runs deterministically with no key. The optional Merge similar (BYOK) toggle folds near-duplicate clusters using your own ANTHROPIC_API_KEY; it is off by default and fail-closed, so with no key, or on any call failure, the deterministic taxonomy shows unchanged. Only category names and truncated critiques are sent, never transcripts or session ids.
Read more
Last updated