Spend · Spend efficiency

Find which spend lever is leaking

RepoOps grades six levers of AI spend (Defaults, Routing, Caching, Context, Visibility, ROI) from signals it already measures, averages the graded ones into one letter, and shows a dash for every lever no signal backs. Four of the six open a pull request against the tracked repository; the other two link to the tab that holds their signal.

For: the engineer who owns a repository's AI spend, and the lead who reads the published grade

What it does, and why it helps

The scorecard reads one tracked repository over a window, thirty days by default. It detects nothing of its own. Defaults grades the share of priced surfaces with no downgrade recommendation against them. Routing grades the share of those surfaces that carry a task-phase mapping. Caching grades the measured cache hit rate. Context grades the count of cc. findings from the rule engine, each costing about a tenth of the scale. Visibility grades the share of measured spend that carries a tenant tag rather than sitting in the untagged bucket. ROI grades the count of surfaces in the bad quadrant, where cost rose and quality fell. Each fraction maps to a letter from A+ down to F, and the composite is the average of the levers that got one.

A lever whose signal is missing reads a dash and says the signal is insufficient to grade it, and the composite is the average of the rest. With none graded the whole grade box is a dash and the page says so. The four trend tiles behave the same way: the thirty-day token and spend deltas, the cost per merged pull request and the cache hit rate each read a dash rather than a zero when the history is too short to compute them. The ranked fix list puts the priced fixes first, and only the Defaults rows carry a price, because that figure is the weekly spend already flowing through the surface the row names, scaled from the aggregate's own window. The ROI guard row always sorts last. Defaults, Routing, Caching and ROI open a pull request from the row; Context and Visibility send you to the tab that carries their signal.

The pain. A bill gives you a total. It does not say which of the six things that set the total moved, and each of those six has its own tab.

The point of view. A grade is worth reading only when a signal backs it. A lever RepoOps cannot measure shows a dash, not a low grade, and a fix carries the spend measured on its surface rather than a promised saving.

What gets easier. Choosing what to change next. Six rows, a letter each, the priced fixes ranked first, and four rows that open the pull request without leaving the page.

When it helps. A tracked repository whose own application writes ai-calls rows, on a machine that holds the Claude Code history the Context lever reads. Opening a pull request also needs an origin remote, a clean working tree and an authenticated gh command line in that repository.

Its limits. It caps nothing, throttles nothing and merges nothing. Only the Defaults rows carry a dollar figure. The realized-savings ledger measures a landed change only where it moved a surface to a different tier, so a caching, task-phase, context or ROI change leaves no row there. The Caching action edits the repository's CLAUDE.md, which is the cacheable prefix it can reach, not a cache_control block in your application code.

Understand it in 30 seconds

30.1 s, captions on. Narration: Microsoft Zira Desktop (provisional voice; an approved narration source is pending).Transcript
Read the narration
  1. 0:00 The bill went up.
  2. 0:01 The bill does not say which lever moved.
  3. 0:06 Six levers set the cost.
  4. 0:08 RepoOps grades each one from telemetry you already hold.
  5. 0:13 A lever with no signal reads a dash, never a low grade.
  6. 0:17 The priced fix names the surface and opens its pull request.
  7. 0:23 Change one lever.
  8. 0:24 The saving is measured after the change lands.

Synthetic example. Read the guide

Where to find it

Where to find it

  • Desktop: localhost:4000, then LLM Cost Metrics in the sidebar, then Spend efficiency under All tools, in the Budget and routing group.
  • Hosted: repoops.ai/team/efficiency, from LLM Cost Metrics in the sidebar, then Spend efficiency under All tools, in the Budget and routing group.
  • Keyboard: ⌘ K, then type “Spend efficiency”.

When to use it

A surface running on a heavier model than it needs

Situation. Defaults grades low and the row reads that some surfaces are right-sizable to a cheaper model class. The Top fixes list puts one of them first with a weekly dollar figure beside it.

What you do. Press Open PR on that row. The route takes the top downgrade-model recommendation for the surface the row named, writes it into the repository's routing config and hands it to the routing proposer.

What you see. The row answers with a link to the pull request and the line that the saving is measured after it lands. The pull request changes .claude/brain/llm-routing.json on a branch named routing-update and its body carries the config diff, the rationale and the eval-gate result.

What it establishes. A change you can read before it merges. Nothing merged: the proposer opens the pull request and restores the branch you were on. If the repository had no routing config, the pull request creates a starter one and the rationale says so, because those tier models are defaults rather than anything measured for this repository. Once the change lands and both seven-day windows hold at least five calls, the realized-savings ledger reports the measured cost-per-call delta, which can be negative.

A routing change that cost more and scored worse

Situation. ROI names a count of routing changes in the bad quadrant, and under it one line per change: the surface, the tier it moved from, the tier it would move back to, the cost regression in dollars and the quality drop.

What you do. Read the named surface first, then press Open revert PR. Only a row carrying a surface can propose; a row with no recommendation behind it offers Review instead.

What you see. The revert is built against what the routing config holds now, not against the ledger's memory of it. The answer names the surface, both tiers, the cost regression and the quality drop. The guard flags a change only when it measurably regressed on both axes, so a window still measuring on either one is never flagged.

What it establishes. A second routing pull request that moves the surface back, for you to read. It is a new change, not an undo of the old one, and the decision is appended to the repository's llm-routing-decisions.jsonl either way.

Before you start

Supported versions
RepoOps desktop v0.3.1, the release this guide was read against. The tab and its endpoint are local. The hosted page needs no desktop running once a snapshot has been published.
Where it runs
Local: LLM Cost Metrics, then Spend efficiency under All tools, in the Budget and routing group. Hosted: LLM Cost Metrics, then Spend efficiency under All tools, in the Budget and routing group (repoops.ai/team/efficiency), read-only, one card per published repository. The hosted page renders only what a desktop published, never an estimate.
Permissions
The local tab has no per-user gate; it is the machine's own dashboard. On the hosted side the page is session-authed, the team comes from your membership, and per-repository access is resolved before the store read, so a scoped member never sees a grade from a repository outside their list and a personal repository never rolls up. A crafted repo parameter falls back to the default pick.
Connections
No model key. The three routing actions and the caching action need an origin remote, a clean working tree and an authenticated gh command line in the tracked repository. Publishing needs this device bound to a team and the toggle on.
Plan
No plan gate. The pricing capability map has no efficiency row; reading the hosted page follows the hosted dashboard sign-in.

Configure it

  1. Pick one repository.

    The Repo picker opens on all tracked repos, which is the fleet index: one row per repository, worst grade first, with a note saying how many of six levers are graded. Levers, fixes and the usage attribution are per repository, and the fleet view says so in each card rather than rendering an empty one.

  2. Read the dashes before the letters.

    A dash means the signal is absent, not that the lever is bad. The grade note under the letter says how many levers the composite averaged, and the line beside it says how many still leak.

  3. Open a pull request from a lever, or follow the row to its tab.

    Defaults, Routing and ROI post to their propose route and the answer is printed in the row, link or reason. Caching posts to the cache-alignment proposer. Context opens Context health and Visibility opens FleetView, because those two have no proposer of their own.

  4. Turn on publishing only if the team should read the grade.

    Settings, then the section What this machine publishes to your team, then Spend Efficiency scorecard. With it on, the server publishes every ready repository on its hourly tick. The Publish to team button on the tab does the same thing for one repository, now.

  5. Tune the two context caps if your sessions are a different shape.

    The Context lever counts findings, and the over-long-session finding fires at 500,000 tokens of fresh input or 1,000 tool calls in one session. Both are environment overrides in the data directory's .env file.

SettingWhereA sensible choiceWhy it matters
repothe tab's Repo picker, and ?repo= on the endpointone tracked repository, not the all tracked repos defaultThe default option answers the fleet question instead: a grade table with no levers, no fixes and no attribution.
sinceDaysthe query on GET /api/spend-efficiency30, the defaultCoerced by the shared clamp: missing or not a positive number reads as 30, anything larger than 90 reads as 90. The tab never sends it, so the grade on screen is always the thirty-day one.
dryRunthe query on each propose route1 while you are reading what a lever would doReturns the built config and a null pull request URL. It touches no git and no gh.
efficiencypublish.enabledSettings, What this machine publishes to your team, Spend Efficiency scorecardoff until the team should read the gradeThe server's hourly tick returns at once unless the value is exactly 1 and the device is bound, so the hosted page stays on its empty state.
REPOOPS_CONTEXT_FRESH_INPUT_CAPthe data directory's .env500000 (the unset default)The fresh input, input plus cache creation, a single session may reach before the Context lever counts it over-long.
REPOOPS_CONTEXT_TOOL_CALL_CAPthe data directory's .env1000 (the unset default)The other half of the same test; either one crossing is enough to count the session.
tenantthe rows your application writes to .claude/brain/ai-calls/a tenant on every call you can attributeVisibility grades the share of measured spend that is not in the untagged bucket, so an untagged repository grades low however well it is routed.
ⓘ
To stop or undo
There is nothing running to stop. The scorecard computes on request and writes no file of its own. Untick Spend Efficiency scorecard in Settings to stop the hourly publish. A pull request a lever opened is an ordinary pull request: close it. The proposer restores the branch you were on and, on success, appends a proposal-opened row to the repository's .claude/brain/llm-routing-decisions.jsonl.

What you should see

A repository with signal

Configuration. One tracked repository with ai-calls rows in the window and Claude Code history on this machine.

Expect. A letter in the grade box with the count of graded levers under it, four trend tiles, six lever rows each with a letter and a dot meter, a ranked fix list with the priced rows first, and the usage attribution card by model and by MCP server.

Verify. The line under the tiles names the composite and how many levers still leak. The attribution card covers the last 24 hours, states the turn count, the session count, the token total and the cost, and names two things it cannot derive from per-turn rows rather than filling them in: a per-subagent split and which skill ran.

A repository with no signal

Configuration. A tracked repository with no ai-calls rows and no local Claude Code telemetry in the window.

Expect. Every lever reads a dash with the note that the signal is insufficient to grade it, the grade box reads a dash, and the narrative says no lever has a backing signal yet.

Verify. The fix list says there are no fixes this window, and the attribution card names the directory it reads from on this machine. No fabricated letter appears anywhere, and the composite stays null rather than defaulting.

The fleet view

Configuration. The picker left on all tracked repos, which clears the repo parameter.

Expect. One table, worst grade first, with a coloured dot and a link per repository and a note saying how many of six levers are graded or that no lever has a backing signal yet. A repository whose read failed reads unreadable, separately from one with no signal.

Verify. The three cards below say the levers, the fixes and the attribution are per repository and to pick one above. None of them sits on a loading placeholder.

Data and cost

What is captured
Nothing new is captured. The scorecard reads files already on disk: the calls your application logged to the repository's .claude/brain/ai-calls/ day files, this machine's Claude Code history for the context findings and the attribution card, the repository's llm-routing-decisions.jsonl and realized-savings.jsonl, and the session outcomes. The rendered answer is cached in your browser, IndexedDB with a localStorage fallback, so a revisit paints the last known answer while the fresh one loads.
Who can see it
Local unless you turn publishing on. The published snapshot is shape only: the repository key, the composite grade, the graded-lever count, the six levers as id, name, grade and tone, the four trend numbers, and at most five fix titles, each redacted and cut to 120 characters. Dropped before it leaves: the lever descriptions, the per-lever signals object, every per-fix surface name and every per-fix dollar figure. The hosted side re-validates the same allowlist, rejects any raw-content key, and refuses a body over 16 KiB.
How long it is kept
The hosted store keeps one row per team and repository, overwritten in place on every republish. No expiry and no purge route; see Delete below. Locally the read window is the sinceDays clamp, thirty days by default and ninety at most, over files you own.
What leaves the machine
The publish is the only RepoOps egress, over the device token, and tenancy is resolved from that token rather than from anything the desktop sends. Opening a pull request talks to your own git remote and your own gh command line. The scorecard itself calls nothing outside the machine.
What it costs
No model call. Every grade re-reads a file already on disk, and the eval gate a routing pull request runs scores its golden dataset with a deterministic string comparison, not a model. The only spend this feature moves is the spend a merged pull request changes.

When the result differs

SymptomLikely causeNext action
The grade box reads a dash and the page says no lever has a backing signal yet.This repository has no priced ai-calls rows and no local Claude Code telemetry in the window.Pick a repository with telemetry. The attribution card names the directory on this machine that it reads.
The page shows a table of every repository instead of six levers.The Repo picker is on its default, all tracked repos, which is the fleet index.Pick one repository. Levers, fixes and attribution are all per repository.
A sinceDays value in the page URL changes the published grade but not the one on screen.The tab sends the window only when you press Publish to team; the scorecard fetch never sends it.Read the on-screen grade as the thirty-day default, or call GET /api/spend-efficiency?repo=<id>&sinceDays=<n> for another window.
Open PR answers that no pull request was opened.The route printed its own reason: no surface is right-sizable in the window, no landed change is in the bad quadrant, the cache lint found no dynamic token in the cached prefix, the repository has no CLAUDE.md, the routing config is invalid, or the working tree has uncommitted changes.Read the reason in the row. The general refusal also names what opening one needs: an origin remote, a clean working tree and an authenticated gh in the tracked repository.
The Routing lever refuses with a note that this repository already routes by task phase.A phases block is already in the routing config, and the proposer refuses rather than overwriting a block someone reviewed.Edit phases in .claude/brain/llm-routing.json by hand. Planning defaults to the heavy tier and execution to the light tier.
The Caching pull request edits CLAUDE.md rather than adding cache_control.The Caching action runs the cache-alignment lint over the repository's CLAUDE.md, which is the cacheable prefix RepoOps can edit. The lever text describes the move; the button edits the file it can reach.Expected. Treat the pull request as a prefix-stability fix and make the cache_control change in your own application code.
The hosted Efficiency page shows its empty state.The device is not bound to a team, the Spend Efficiency scorecard toggle is off, or no tick has run since it was turned on.Bind the device, turn the toggle on in Settings, and wait for the hourly tick. Publish to team on the tab publishes one repository now and reports the reason as-is when it cannot.
Disable
Untick Spend Efficiency scorecard in Settings to stop the hourly publish. There is nothing else to disable: the scorecard runs only while you are looking at it, and every pull-request action is a button press.
Roll back
Not provided for a merged change. RepoOps does not undo a routing pull request; revert it in git. Before the merge, closing the pull request is the reversal. The ROI lever's revert is a new routing pull request that moves a surface back, not an undo of the one that moved it.
Revoke access
Disconnect the device in the desktop app to forget the binding and revoke the device token, which stops this publisher and every other one. The gh credential the proposers use is your own; RepoOps neither stores nor revokes it.
Delete
Not provided. No route deletes a published rollup. The row lives in the hosted spend_efficiency_rollup table, keyed by team and repository, and a republish overwrites it in place, so unbinding the device stops new rows rather than removing the last one. Locally there is nothing to delete: the scorecard stores no file of its own, and the browser copy goes with the site data.

Maintenance evidence

Feature id
spend-efficiency (spine leaf efficiency)
Owner
The Spend Efficiency Scorecard program in docs/roadmap.md (2026-06-29), plan docs/plans/spend-efficiency-scorecard.md: the scorecard, the five lever proposers (Model Right-Sizing, Task-Aware Routing, Cache Optimizer, Context Hygiene Coach, Cost-Quality Guard) and the Hosted Spend Efficiency tab. No ledger id owns the feature itself. Guide: LDG-0717.
Supported product version
RepoOps v0.3.1
Last verified
2026-09-15, read against origin/main at 52366bb6d; labels read from the served tab source (public/spend-efficiency.html and public/settings.html), defaults and refusal reasons read from the route and lib modules, and the hosted contract read from the store and ingest route. Not checked on a running instance.
Example fixtures
The one real fixture is the served tab itself: test/spend-efficiency-levers.test.mjs reads public/spend-efficiency.html and asserts the Defaults and Caching rows reach a proposer. Every other signal shape is inline: lib/spend-efficiency.test.mjs (the honesty contract and the composite), lib/spend-efficiency-fix-actions.test.mjs (which fix rows may propose), lib/spend-efficiency-propose-routing.test.mjs with lib/routes/spend-efficiency-propose-routing.test.mjs and lib/routes/spend-efficiency-propose-revert.test.mjs (the proposers, with git and gh injected), lib/routes/cache-optimization.test.mjs, lib/cna/spend-efficiency-publish.test.mjs (the redaction contract), and on the hosted side website/lib/cna/spend-efficiency-store.test.ts, website/lib/cna/spend-efficiency-store.load-error.test.ts, website/app/api/publish/spend-efficiency/route.test.ts and website/app/team/(home)/efficiency/page.test.tsx.
Source references
lib/spend-efficiency.mjs, lib/routes/spend-efficiency.mjs, lib/spend-efficiency-propose.mjs, lib/routes/cache-optimization.mjs, lib/routing-proposer.mjs, lib/realized-savings.mjs, lib/cc-usage-attribution.mjs, lib/rules/cc.session-fragmentation.mjs, lib/since-days.mjs, lib/cna/spend-efficiency-publish.mjs, lib/routes/publish-toggles.mjs, public/spend-efficiency.html, public/settings.html, website/lib/cna/spend-efficiency-store.ts, website/app/team/(home)/efficiency/page.tsx
Documentation review
Independent review requested on the slice pull request; not yet recorded.
Video review
Narrated story rendered and published 2026-09-26 (render 17c84266af6d, LDG-1014) with the breadcrumb LLM Cost Metrics, which lists the feature under Moved here, checked against main at 7aab4cd82 with LDG-1014 part 1. Six frames, the captions and the transcript were reviewed by the authoring agent, not an independent reviewer; the audio was not listened to by a person. Narration is the provisional Windows voice until LDG-0721.

Last updated