Watch · Agent traces
Find the run that failed and what it cost
RepoOps turns the Claude Code sessions it already captured into a worklist. Each row carries a verdict derived from CI, revert, quality and cost, so the failing runs sort to the top. Open one for its span waterfall, its cost-ranked steps and its per-turn detail.
For: the engineer triaging a batch of agent runs, and the lead who wants the bad ones first
What it does, and why it helps
The trace model reads the per-turn cycles the telemetry parser already pulled from this machine's Claude Code transcripts and persists them as a span tree of session, turn and tool in the local store. The tab reads the newest 200 sessions for the selected repository inside the window you pick, and stamps each one with a verdict: Reverted / defect, Failed CI, Low quality, Over budget, Shipped clean or Unscored. The first match wins in that order. Low quality is a score under 0.5 on the objective rubric. Over budget means the run sits in this repository's own costly tail, the cost at the 75th percentile of non-zero session costs, unless you type a figure into Over $.
Open a session and it renders in one of three lenses. Waterfall draws the span tree, with turn bars scaled by real duration when the timestamps are there and tool bars scaled by cost or tokens, because no per-tool clock is captured. Cost-ranked steps sorts every turn and tool step by what it cost. Turn detail shows the prompt, the response and, when deep capture is on, the ordered tool inputs and outputs. A second, separate column is yours: a human verdict of Pass, Fail or Needs review, a failure-mode tag and a free-text critique, appended to the repository brain. Fails roll up into the Failure modes, fix this first panel, ranked by how often a labeled run hit them.
The pain. A week of agent runs is a pile of JSONL. The run that failed CI, the run that got reverted and the run that cost eleven dollars all look the same in the file, so nobody opens the file.
The point of view. A run's outcome is a column, not a memory. Derive it from what git and CI already said, put the worst first, and keep the reviewer's own judgment in a separate column so it is never mistaken for a measurement.
What gets easier. Triage. Sort by verdict, cost, quality or tool count; filter to Failed CI, Reverted / defect, Over budget or Low quality; search the loaded rows by session id, source tool, PR number or verdict label.
When it helps. After a batch of agent work in a repository whose transcripts are on this machine, when the question is which run to open first and what it spent getting there.
Its limits. Per-tool latency and live streaming spans are not captured, so a tool bar never means time. A subagent's spans are not stitched into its parent's waterfall; agent-to-agent timing comes from the attested spawn log in the fan-out panel. The verdict and quality columns stay empty until you press Score cost vs quality, which needs git and gh on this machine. The list is capped at 200 sessions and the window at 90 days.
Understand it in 30 seconds
Read the narration
- 0:00 Eight agent runs finished overnight.
- 0:02 The file does not say which one hurt.
- 0:06 RepoOps reads the outcome git and CI already recorded.
- 0:10 A revert outranks a failed check.
- 0:13 Open the run and the waterfall shows its shape, turn by turn.
- 0:17 Cost-ranked steps name the call that burned the budget.
- 0:23 Label it yourself, pass or fail.
- 0:26 That label is a record, not prevention.
Synthetic example. Read the guide
Where to find it
Where to find it
- Desktop:
localhost:4000, then Attribution in the sidebar, then Agent traces under All tools, in the Evidence readers group. - Hosted:
repoops.ai/team/agent-traces, from Attribution in the sidebar. - Keyboard: ⌘ K, then type “Agent traces”.
When to use it
The overnight batch, and which run to open first
Situation. Eight agent sessions ran against one repository yesterday. The list loads with every row reading Unscored, because nothing has linked them to git yet.
What you do. Leave Since (days) at 30, press Score cost vs quality, and keep the Sort segment on Verdict. Read down from the top.
What you see. The badges fill in. A row shows the session id, its verdict, the turn and tool counts, the cost to three decimals, the PR number and a q score out of 100. The meta line reads how many of how many sessions are showing, how they are sorted, and where the over-$ cut landed. Four quadrant cells appear: Expensive and poor, Expensive and good, Cheap and poor, Cheap and good.
What it establishes. The order is derived, not guessed. Reverted beats Failed CI beats Low quality beats Over budget. A session with no CI status and no quality score reads Unscored, never a pass.
The run that cost far more than the rest
Situation. One session is an outlier on cost and you want to know which call burned the budget before you change anything.
What you do. Type a figure into Over $, set Show to Over budget, open the session, and switch the lens to Cost-ranked steps. Press the lesson button beside the step you want to keep.
What you see. Every turn and tool step is ranked with a bar, its dollars to four decimals and its token count. Save as lesson opens a form with three required fields (trigger, what went wrong, fix) pre-filled from the session, and the meta line names the lesson id it wrote to the brain.
What it establishes. The costly call is named and recorded. A lesson is a written record, not a control: nothing in this tab stops the next run from making the same call. Recurrence is read from later captured sessions, and with none in the window it is not assessable.
Before you start
- Supported versions
- RepoOps desktop v0.3.1, the release this guide was read against. The tab reads Claude Code transcripts already on this machine, so a repository with no captured sessions renders an empty list rather than an error.
- Where it runs
- Local: the Agent traces tab under Attribution, with Run forensics as its subtab. Hosted: repoops.ai/team/agent-traces is the Attribution page. A production or performance case in its queue opens on that page, on its Attribution section, and the chain cells sit beside the queue. Below them, open on arrival, it lists the captured sessions your team streamed (session, repo, cost, tokens, turns, PR, verdict) and draws a span waterfall for the one you open. The specialist readers sit under All tools, one closed list.
- Permissions
- Local: none beyond access to this machine. Hosted: a signed-in member. The team is resolved from your membership, never from a URL parameter, and a scoped member sees only their allowed repository keys.
- Connections
- Score cost vs quality shells out to git in the repository checkout (8 second timeout) and to gh for the pull request's checks (12 second timeout). Streaming to the hosted page needs a bound device and the telemetry stream opt-in.
- Plan
- The local tab has no plan gate. The hosted ingest route refuses a free or lapsed team with 402 and the message that a paid plan (Solo Hosted or above) is required to stream agent traces.
Configure it
- Pick the repository and the window.
The tab reads ?repo=<id> and says No repo selected without one. Since (days) offers 7, 30 and 90, and 30 is the default. Refresh forces a full re-walk of the telemetry cycles; every other load refreshes the trace model in the background, throttled to 45 seconds, and paints the last cached payload first.
- Fill the outcome columns.
Score cost vs quality runs the outcome linker and the quality scorer locally. The linker maps a session to its commits, its pull request and that pull request's CI conclusion; the scorer blends CI (weight 0.4), not-reverted (0.3), churn (0.1) and, when a caller supplies them, a judge sub-score (0.2) and verifier independence (0.15). Each weight applies only when its signal is present, so the score renormalizes rather than inventing a number.
- Stamp the session id on your commits so the linkage is exact.
The prepare-commit-msg git hook writes a Session-Id trailer when a session id is in the environment. Claude Code exports CLAUDE_CODE_SESSION_ID; REPOOPS_SESSION_ID takes precedence if you set it. With no trailer the linker falls back to commits authored inside the session's start and end window, which is a heuristic, not an identity.
- Set the over-budget cut, or leave it to the repository.
Blank means the cut is the cost at the 75th percentile of this repository's non-zero session costs, so Over budget means costly for this repository rather than costly in the abstract. A figure in Over $ overrides it for the whole list, and the meta line prints the cut in use.
- Turn on deep capture if you need the Turn detail lens.
REPOOPS_DEEP_CAPTURE=1 in the data directory's .env records the ordered tool_use arguments and tool_result payloads per turn. It is off by default. Without it a turn card still shows the prompt, the response, the duration and the tokens, and the lens prints a hint naming the variable.
- Label the run while you have it open.
The labeling bar takes Pass, Fail or Needs review. On a fail the failure-mode select opens with a closed list: ci-fail, revert, tool-thrash, wrong-file, oversized, duplicate, ignored-norms, message-mismatch, other. Save label appends the row to the repository brain. Save as error queues a proposed errors.md block through the brain pipeline and never edits errors.md.
- Only then, stream to the team.
Settings, Privacy & keys, Telemetry privacy, the Stream to your team checkbox. One flag gates both the activity stream and the agent-trace stream. Until a machine is bound to a team the checkbox is replaced by a note saying nothing is uploaded.
| Setting | Where | A sensible choice | Why it matters |
|---|---|---|---|
Since (days) | Agent traces, the controls bar | 30 (the default); 7 and 90 are the other options | Any value is coerced to a number, falls back to 30 when missing or not positive, and is capped at 90, so no request can force a re-parse of all history. |
Over $ | Agent traces, the worklist bar | blank, so the cut is this repository's costly tail | Blank uses the 75th percentile of non-zero session costs. A typed figure replaces it for every row at once. |
REPOOPS_DEEP_CAPTURE | the data directory's .env | 1 when you need tool inputs and outputs, otherwise unset | Off by default. Only a session recorded while it was on carries tool IO; turning it on does not backfill. |
REPOOPS_DEEP_CAPTURE_MAX_CHARS | the data directory's .env | 8000 (the default) | Per-field character cap on deep-capture payloads, so one huge tool result cannot dominate the store. |
REPOOPS_CAPTURE_FIELD_MAX_CHARS | the data directory's .env | 65536 (the default) | Character cap on the always-captured prompt and response text; a clipped field is marked as clipped. |
telemetry.capture_mode | Settings, Privacy & keys, Telemetry privacy | Metadata only (the default) | In metadata mode the prompt, response and tool text are dropped at egress and the lens says so; structure, durations and errors are kept. Content (redacted) keeps the text. A team policy can pin metadata only, and then the Content radio is disabled. |
telemetry.stream_enabled | Settings, Privacy & keys, Telemetry privacy, Stream to your team | off until you want the hosted page to render | No span leaves the machine until this is on and the device is bound. It is the same flag the activity stream uses. |
CLAUDE_CODE_SESSION_ID | the environment the commit runs in | leave the value Claude Code exports in place | The git hook stamps it as a Session-Id trailer, the exact link between a session and its commits. REPOOPS_SESSION_ID overrides it; with neither set the hook is a no-op. |
REPOOPS_NEON_RETENTION_DAYS | the hosted environment | unset, which is 30 | The daily retention cron hard-deletes hosted span rows older than this. Thirty days is the default and the ceiling: a lower value keeps less at rest without a redeploy, and a higher one runs at 30. |
What you should see
A repository with captured sessions, nothing scored yet
Configuration. A tracked repository whose Claude Code transcripts are on this machine, window 30, Score cost vs quality not pressed.
Expect. A row per session with the turn count, the tool count and the cost, each badged Unscored. Clicking one opens the waterfall with the session, its turns and its tools.
Verify. The meta line names how many of how many sessions are showing and how they are sorted. A session id the store does not hold returns session not traced, never a fabricated timeline.
After the git pass
Configuration. The same repository, Score cost vs quality pressed once.
Expect. Verdict badges fill in, the four cost-and-quality quadrant cells appear, and the meta line reads how many of the sessions were scored and where the cost split fell.
Verify. A run whose CI succeeded and was not reverted reads Shipped clean. A run with no CI status and no quality score stays Unscored. A run whose pull request failed reads Failed CI even when its cost is high, because the precedence puts the defect first.
Nothing captured yet
Configuration. A tracked repository this machine has never run Claude Code in.
Expect. No traced sessions yet. Run Claude Code in this repo, then Refresh. The fan-out panel reads No attested fan-out for this repo yet, and the drift panel reads No accepted baseline yet.
Verify. Every empty state names the missing input rather than showing a zero. A fan-out with attested spawns but no timing reads untimed, not 1.0, because an untimed fan-out is not a sequential one.
Data and cost
- What is captured
- Spans and a one-row rollup per session in this machine's SQLite store (data.db under %APPDATA%\repo-dashboard on Windows, ~/.repo-dashboard elsewhere): the session, turn and tool spans with tokens and cost, and per session the turn and tool counts, cost, commit SHAs, PR number, CI status, revert flag, churn lines and quality score. Re-syncing a session replaces its spans, so a re-scan is idempotent. Human labels append to .claude/brain/agent-labels/<date>.jsonl in the repository; a saved lesson goes to the brain's lessons/lessons.jsonl; Save as error writes a proposal block under .claude/brain/proposed/.
- Who can see it
- Local by default. With Stream to your team on and the device bound, each pass sends at most 50 sessions with their spans plus the labels from the last 30 days. A turn span's local name is a prompt preview, so it is replaced with turn plus its sequence number before it leaves; span attributes are reduced to model, turns, count and toolCount; every remaining string goes through the redactor. The ingest route re-redacts server-side and records a server_redacted witness when it finds anything the client missed.
- How long it is kept
- Local: no sweep. Spans are replaced when a session re-syncs, and the tab reads the newest 200 sessions inside a window capped at 90 days, but nothing deletes older rows from data.db. Hosted: the daily retention cron hard-deletes span rows older than REPOOPS_NEON_RETENTION_DAYS, 30 by default and at most. The hosted per-session rows are not in that sweep, so an older session still lists with its cost and outcome, and its span waterfall reads Expired with the date the detail left.
- What leaves the machine
- Nothing leaves the machine until Stream to your team is on. The tab makes no model call: the verdict, the quality score and the failure-mode ranking are all computed without a key. Score cost vs quality runs git locally and calls GitHub through gh for the pull request's checks. The quality rubric has a judge weight, and this path passes no judge score.
- What it costs
- No RepoOps charge and no token spend on the local path. The dollar figures on the rows are what your Claude Code runs already cost, read from the captured token counts. Streaming to the hosted dashboard needs a paid plan.
When the result differs
| Symptom | Likely cause | Next action |
|---|---|---|
| The tab says No repo selected. | It was opened without a repo in the query string. | Pick a repository in the dashboard shell, or add ?repo=<id> to the URL. |
| No traced sessions yet. | No Claude Code transcript for this repository on this machine, or the cycles have not been walked since the last run. | Run Claude Code in the repository, then press Refresh, which forces the full re-walk instead of the throttled background one. |
| Every row reads Unscored. | The outcome linker and the quality scorer have not run for this window. | Press Score cost vs quality. It needs git and gh on the PATH and a checkout it can read; an offline gh times out at 12 seconds and leaves CI status null. |
| Turn detail says tool inputs and outputs are not captured. | REPOOPS_DEEP_CAPTURE is not set for the session that was recorded. | Set it to 1 in the data directory's .env, restart, and re-run a session. Sessions recorded earlier keep the shape they were recorded in. |
| The turn cards show structure and no text. | Capture mode is metadata, which drops prompt, response and tool text at egress. | Switch to Content (redacted) in Settings, Privacy & keys. If your team pins metadata only, the radio is disabled and the pin note says so. |
| The hosted page says No traces yet. | No desktop in the team has streamed, or the team is on a free or lapsed plan. | Turn on Stream to your team in the desktop Settings on a bound machine. A free team's ingest is refused with 402 before anything is written. |
| The fan-out panel reads untimed. | The attested spawn log holds records with no timing, so no burst can be measured. | Nothing to fix here. Spawns are grouped into bursts on a 30 minute idle gap, and a record that fails signature verification is excluded and counted in the footer. |
- Disable
- Untick Stream to your team in Settings, Privacy & keys, to stop the hosted copy; the streamer then treats the machine as unbound and does nothing. Unset REPOOPS_DEEP_CAPTURE to stop recording tool inputs and outputs. The tab itself has no on and off switch, because it reads what telemetry already captured.
- Roll back
- Not provided. Labels, lessons and proposed errors are appends. Saving a second label for the same session is what the tab reads back, and the earlier row stays in .claude/brain/agent-labels/<date>.jsonl. A proposed error is reversed by not accepting it in the brain acceptor pass.
- Revoke access
- Disconnect the device in the desktop app to revoke its device token; the streamer reports not bound and sends nothing. On the hosted side a scoped member's repository allow-list is what narrows the session list, and it is default-deny.
- Delete
- Not provided for a real repository. The one function that clears a repository's spans and outcome rows has a single caller, the demo-data clear, and it is pinned to the demo repository id. The local rows sit in data.db in the RepoOps data directory. On the hosted side the retention cron deletes span rows past the window; the per-session rows are not in that sweep.
Related tasks
Maintenance evidence
- Feature id
agent-traces(spine leafagent-traces)- Owner
- K-EVAL program (E0.1 the trace model, E1 the viewer, E2 quality scoring), with the Loop Engineering Phase 2 worklist and labeling bar, the Agent Forensics Kernel drift panel, and the W5-TRACES hosted parity slice. Guide: LDG-0717.
- Supported product version
- RepoOps v0.3.1
- Last verified
- 2026-09-15, read against origin/main at 52366bb6d; labels read from the served tab source (public/agent-traces.html), the capture and streaming labels from public/settings.html, the hosted columns and empty state from website/app/team/(home)/agent-traces/page.tsx, and the caps and defaults from the route and lib modules cited in this file's header.
- Example fixtures
- No fixture file. The behaviour is exercised inline: lib/agent-verdict.test.mjs for the verdict precedence and the auto cut, lib/agent-trace.sync.test.mjs for a telemetry fixture walked into a span tree, lib/agent-trace-detail.test.mjs for the turn detail and the per-tool error flag, lib/agent-trace-streamer.test.mjs for the egress projection, lib/agent-labels.test.mjs and lib/agent-labels.rollup.test.mjs for the labels and their rollup, lib/spawn-topology.test.mjs for the burst split, lib/routes/agent.test.mjs for the routes, and website/app/team/(home)/agent-traces/page.test.tsx for the hosted page.
- Source references
public/agent-traces.html,public/settings.html,lib/agent-verdict.mjs,lib/agent-trace.mjs,lib/agent-trace-detail.mjs,lib/agent-labels.mjs,lib/agent-save-error.mjs,lib/outcome-linker.mjs,lib/quality-score.mjs,lib/since-days.mjs,lib/spawn-topology.mjs,lib/agent-trace-streamer.mjs,lib/routes/agent.mjs,bin/db/schema.mjs,website/app/team/(home)/agent-traces/page.tsx,website/app/api/agent-traces/ingest/route.ts,website/app/api/cron/events-retention/route.ts- Documentation review
- Independent review requested on the slice pull request; not yet recorded.
- Video review
- Story script written 2026-09-15; render and review pending in the same slice.
Last updated