Spend · Billing Guard reference

Check a bill against your own token counts

RepoOps recomputes what each day of AI work should have cost from the token counts captured on this machine, compares that against what the provider billed, and reports each day as matching your records, not matching, or not checked. It detects and reconciles. It never caps, blocks or throttles spend.

For: the engineer who owns a repository's telemetry, and the person who has to take a disputed line back to the provider

What it does, and why it helps

Token counts are the primitive and dollars are derived. For every day in the window, the engine prices your local Claude Code telemetry against the versioned rate card and calls that the expected cost, then joins it to the provider's own usage and cost reports by day and model. The comparison runs inside a tolerance band, never on cent equality, because a provider disclaims its client-side cost figures as approximations. Each billed day then reads one of three ways on the tab: matches your records, doesn't match, or not checked. A day your machine did not capture is not checked, and it is never counted as a problem.

Ten detectors run over that join, and every finding carries exactly one of them. Four compare the bill against your local telemetry and are marked scope sensitive: variance-band, ghost-usage, duplicate-burst and cache-share-divergence. Four compare the provider's own surfaces against each other or against the rate card, so a coverage gap cannot produce them: rate-check, delta-alarm, cold-start-spend and contradiction. Two are exact joins inside the local request log and need no provider report at all: duplicate-request and ghost-prompt. A finding only ever carries the verdict verified-mismatch, which means both sides were present and disagreed beyond the band. Everything the engine could not check lands in coverage instead, as a day it could not verify or as a named blind spot.

What you do with a finding is yours. The findings table names the detector and shows what you expected, what you were billed and the difference; Draft dispute opens your mail client with those numbers and the raw evidence, and sends nothing; Export evidence bundle downloads the same JSON the command line writes, so finance reads one format. A daily sweep at 02:00 reconciles the last seven days and pages on three signatures only. The hosted team page shows what a connected desktop synced, because the Admin key stays on the machine by design.

The pain. The invoice arrives as a total. The tokens that produced it were counted on your own machine, and almost nobody puts the two side by side. The auditing startup Vaudit reported 1.7 million dollars in overcharges across 34 million dollars of audited AI invoices in June 2026, a self-reported figure that Anthropic and OpenAI dispute.

The point of view. A bill is a claim, and a claim you have the inputs for is one you can recompute. Recompute the day from your own token counts before you argue about it, and treat a day you could not check as unchecked rather than clean.

What gets easier. Disputing. Every finding names the detector that fired, both numbers and the difference, and the evidence bundle carries the token counts behind them, so the conversation with a provider starts from arithmetic rather than a feeling that the bill looks high.

When it helps. You hold an organization Admin key for the provider you want to check, and the machine captured telemetry for the days in question. Without a key it still recomputes your own side and marks every day not checked.

Its limits. Local telemetry covers one checkout. An organization Admin key bills the whole organization, so the four scope-sensitive detectors read low whenever other people, other worktrees or other projects share that key, and the tab says so whenever a Claude account is connected. AWS Bedrock and Azure OpenAI cost APIs carry no per-model token count at all, so their findings can only ever be day level. Billing Guard caps no spend.

Understand it in 30 seconds

30.1 s, captions on. Narration: Microsoft Zira Desktop (provisional voice; an approved narration source is pending).Transcript
Read the narration
  1. 0:00 The invoice arrives as one total.
  2. 0:02 Nobody recomputes what it should have been.
  3. 0:06 RepoOps prices your own token counts against the rate card.
  4. 0:10 Then it compares, day by day.
  5. 0:13 Ten detectors name what they found, with both numbers.
  6. 0:17 A day it cannot check reads not checked, never clean.
  7. 0:23 Export the evidence, or draft the dispute.
  8. 0:26 Billing Guard detects; it never caps spend.

Synthetic example. Read the guide

Where to find it

Where to find it

  • Desktop: localhost:4000, then LLM Cost Metrics in the sidebar, then Billing Guard under All tools, in the Evidence readers group.
  • Hosted: repoops.ai/team/billing-guard, from LLM Cost Metrics in the sidebar, then Billing Guard under All tools, in the Evidence readers group.
  • Keyboard: ⌘ K, then type “Billing Guard”.

When to use it

What was billed, beside what was measured

Situation. You open LLM Cost Metrics on the desktop with a repository chosen in the scope bar, and you want the bill and your own receipts in one place without a number that quietly adds them together.

What you do. Read the panel titled Billed and measured, in this scope, above the cost queue. Open Receipts or Billed days under any figure to see what it is made of.

What you see. Six figures. Provider billed comes from the provider cost report Billing Guard reads, with the provider, the time it was read and the window; it covers the provider account, not only this repository. Measured usage is the case cost receipts in the scope, with the newest receipt time, each currency listed by its own code. Estimates carry the words estimate, never invoice evidence. Unknown rate counts receipts that were not priced. Unallocated counts measured receipts with no spend purpose. Costly failed attempts prices the builds that ended failed, and reads not reported where no build receipt is in scope. A receipt row gives the provider and model, the source record, the time, the rate version, the case and repository it is allocated to, whether it was measured or estimated, and the amount.

What it establishes. The two figures stay two figures. They cover different scopes, so the page computes no difference between them and points to Billing Guard, which reconciles the bill against this machine's telemetry per day and model and says how many days reconciled, how many have a finding and how many were not checked. A receipt attached to two cases is counted once, and a shared parent is not summed beside its shares. Open a case and its Cost section lists Investigation, Remediation, Validation and Production impact, each measured apart from estimates and not recorded where no receipt names that purpose. The panel also says what a budget does on each surface: Billing Guard never caps, a monthly LLM budget set on the desktop holds that machine's spending actions once the month's measured spend reaches it, and hosted budget alerts are alerts, not caps. The hosted LLM Cost Metrics page carries the same panel over the team's case receipts. Its billed figure is the provider bills desktops attached to cases, one entry per bill however many cases it sits on, and reads not reported where none is attached, because the hosted dashboard holds no provider cost report. Its costly failed attempts count the receipts that say how their run ended (desktop scope runs and builds, hosted validation runs) and read not reported where none does. The Costs page CSV export follows the repository you scoped the page to.

A billed day that does not match your records

Situation. The Anthropic account is connected, the tolerance band is the default 2 percent, and the tab is reading the last 30 days.

What you do. Open LLM Cost Metrics in the sidebar, then Billing Guard under All tools, in the Evidence readers group. Read the day chips, then the findings table. Press Draft dispute on the row you want to argue, and Export evidence bundle for the attachment.

What you see. The row names the day, the model, what we expected, what you were billed, the difference, and a badge naming the detector. The mail draft repeats those numbers with the raw token counts underneath. The bundle is JSON with kind repoops.billing-guard-report.

What it establishes. A dispute with arithmetic attached. It does not establish that the provider is wrong: if the detector is a scope-sensitive one, the coverage gap is the first thing to rule out, and the page names which detectors that applies to.

Spend that restarts after a quiet spell

Situation. A credential was set and metered billing resumed. Yesterday billed nothing, today bills real money, so a day-over-day ratio has nothing to divide by.

What you do. Leave the daily sweep alone. It runs in the 02:00 pass over the last seven days and pages through the notifications spine.

What you see. One notification per finding per seven-day dedup window, keyed by provider, day, model and signature. Only cold-start-spend, delta-alarm and contradiction page; the other seven stay on the tab and in the command line output.

What it establishes. The shape that a ratio alarm is blind to gets its own floor, 10 dollars by default. A quiet page is not a clean bill: the sweep no-ops when no provider is configured, and it is off entirely when the alerts flag is set to 1.

Reading it for a team, and finding the cost of a case

Situation. Several repositories run the desktop app. You want reconciliation status across them, and you want to know what a given incident case cost.

What you do. Open Billing Guard on the hosted dashboard. Read the tiles and the per-repository table, then the Cases with a cost section, which is on both the hosted page and the desktop tab.

What you see. Per repository: last sweep, findings, warn, info, scope-sensitive, top signature and failed provider pulls. A repository with no sweep for 48 hours is badged stopped, and one that never reported reads not reported. A case row shows measured beside estimated, never summed, with none measured rather than a zero.

What it establishes. You can tell a repository that reconciles clean from one that is not being checked. The LLM Cost Metrics primary page on the desktop carries the pillar's cost-class queue, which says whether the cost source (the production events gateway) is missing, idle, stale or failed rather than reading zero, and reads a measured zero only when that source is current and has carried cost events. The hosted LLM Cost Metrics page carries the same cost-class queue under your repository scope.

Before you start

Supported versions
RepoOps desktop v0.3.1, the release this guide was read against. The same engine ships in the command line shell (bin/billing-guard.mjs, zero model tokens by design) and in the Claude Desktop extension under installer/claude-desktop-billing-guard/. The daily sweep needs the app running in the 02:00 window.
Where it runs
Local: the Billing Guard tab, under All tools on the LLM Cost Metrics primary page, which opens with the billed and measured panel and then the cost queue. Hosted: the Billing Guard page shows a 30-day rollup of the billing_guard.finding and billing_guard.sweep events a connected desktop synced, plus the same Cases with a cost table. Hosted: LLM Cost Metrics in the sidebar opens with the same billed and measured panel, then the applied-action spend; a cost case's link opens the case there, and /team/llm-cost redirects to it.
Permissions
Storing or removing a provider key takes this machine's operator write confirmation; the read that renders the tab is deliberately not gated, so the page can always say what state it is in. On the hosted side a member sees only the repositories their team narrowing allows.
Connections
An Anthropic organization Admin key, an OpenAI organization admin key, or both, pasted on the tab and validated live against that vendor before anything is stored. AWS Bedrock and Azure OpenAI reconcile too, but only through their environment variables: neither is offered a paste box, because a multi-variable credential set is not one box. Optionally, Claude Code's own OTLP logs export pointed at this machine, which unlocks the two request-level detectors.
Plan
No plan gate. The tab, the command line shell and the sweep are local, and the pricing capability map has no Billing Guard row. Seeing the rollup on the hosted dashboard follows the hosted dashboard tiers.

Configure it

  1. Connect a provider account from the tab.

    Under Your provider accounts, paste an Anthropic Admin key or an OpenAI admin key and press Connect. RepoOps checks the shape, then checks the key with that vendor before storing it, so a key it accepts is one the vendor accepted. The key is stored on this machine, never returned to the browser, and never shown again: every response carries configured, source and a masked hint. An environment key always wins over a stored one, and Disconnect cannot remove an environment key.

  2. Set the tolerance band.

    The Tolerance band control sits next to Reconciliation and starts at 2 percent. It is the width inside which a billed day counts as matching your records. A band tighter than the provider's own approximation noise produces noise rather than findings.

  3. Read the measured baseline before you move the band.

    The line under the control reports your typical drift across fully checked days: the median, the ninetieth percentile and the worst. Once three checked days exist it also suggests a band, the ninetieth percentile with a 25 percent margin, rounded up to the nearest half point and never below 1 percent. It is read only. RepoOps never applies it for you.

  4. Narrow the Anthropic reports if you hold a workspace-scoped key.

    Setting the workspace id in the data directory's .env makes Anthropic's usage and cost reports cover that workspace instead of the whole organization, which is the closest an Admin key gets to local telemetry's single-folder scope. It is Anthropic only; the other three clients ignore it.

  5. Read the ledger, then act on a row.

    Day chips first, findings table second, coverage summary third. The two disclosures under the summary are worth opening: the days that could not be checked, each with its reason, and the blind spots, which are things the engine can see but cannot cross-check and so never calls mismatches. Draft dispute and Export evidence bundle are the two actions.

  6. Decide whether the daily pages stay on.

    The sweep runs in the 02:00 pass, reconciles the last seven days against every configured provider, and pages only on cold-start-spend, delta-alarm and contradiction, each deduped for seven days. Setting the alerts flag to 1 stops the pages and leaves the tab and the command line untouched.

SettingWhereA sensible choiceWhy it matters
ANTHROPIC_ADMIN_KEYBilling Guard tab, Your provider accounts, Connect Anthropic account, or this machine's environmentan Admin key from Create Admin Key, starting sk-ant-A regular API key cannot read organization usage and cost, so it is refused. Without any key the findings table says nothing to report yet and every day reads not checked.
OPENAI_ADMIN_KEYBilling Guard tab, Connect OpenAI accountan organization admin key, sk-admin-Validated live on connect. The box refuses an sk-ant- key outright so an Anthropic credential is never sent to OpenAI. Reading a large organization has not been exercised against a real account, and the tab says so.
toleranceBilling Guard tab, the Tolerance band control, and the query parameter behind it2 (the default), or the suggested band once three checked days existThe width inside which a billed day counts as matching. Below the band nothing is a finding; above it, variance-band fires.
REPOOPS_BILLING_GUARD_TOLERANCE_PCTthe data directory's .envunset, which reads 2The default band for the read and the command line shell when no explicit value is passed. An explicit value on the request still wins.
REPOOPS_BILLING_COLD_START_FLOOR_USDthe data directory's .env10 (the default)A day that follows a zero-billed day and bills at least this much fires cold-start-spend. Zero is refused rather than honored, because a floor of zero pages on every first billed day after any gap.
REPOOPS_BILLING_GUARD_ALERTS_OFFthe data directory's .envunsetSet to 1 to stop the daily pages. The sweep already no-ops when no provider is configured.
REPOOPS_ANTHROPIC_ADMIN_WORKSPACE_IDthe data directory's .envunset, unless you hold a workspace-scoped Admin keyNarrows Anthropic's reports to one workspace so they sit closer to local telemetry's scope. Anthropic only.
REPOOPS_BILLING_TTL_MSthe data directory's .env300000 (the default); 0 turns the memo offHow long a resolved Admin cost report is reused before another pull goes out.
AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_REGIONthe environment onlyset all three if you bill through BedrockCost Explorer is the cost source, and it has no per-model token count, so Bedrock findings are day level and always carry a precision-ceiling note.
AZURE_TENANT_ID, AZURE_CLIENT_ID, AZURE_CLIENT_SECRET, AZURE_SUBSCRIPTION_IDthe environment onlyset all four if you bill through Azure OpenAICost Management is cost only, an even narrower signal than Cost Explorer, so the same day-level ceiling applies.
OTEL_LOGS_EXPORTERClaude Code's own environment, never this repository'sotlp, pointed at this machine's logs receiver with the tracked repo idUnlocks duplicate-request and ghost-prompt, which supersede the statistical duplicate-burst on any day the store covers. The receiver half is verified; three attempts to make a one-shot claude run export logs produced no records, so treat the emitting half as unconfirmed.
ⓘ
To stop or undo
Press Disconnect to remove a stored key from this machine: the comparison stops, every day goes back to not checked, and your own records stay. A key set in the environment is not yours to remove from here, and the page says where it came from instead of offering a button that cannot work. Setting the alerts flag to 1 stops the daily pages on their own. Billing Guard caps no spend, so stopping it changes only what you can see.

What you should see

Connected, and the days line up

Configuration. One provider connected, the band at 2 percent, the tab reading its fixed 30-day window.

Expect. A chip per billed day, a count of how many match your records, how many do not and how many were not checked, and an empty findings table reading that nothing looks wrong and every day it could check matches your own records.

Verify. The four coverage tiles read Days of your records, Days you were billed, Mismatches found and Days not checked. The badge row says what it had to work with: your records, usage from the provider, the bill from the provider, and whether the usage resolved per model or whole day only.

A day it cannot check, and a blind spot it will not call a mismatch

Configuration. Any of the above, on a day the machine captured nothing, or a day carrying Priority Tier or code-execution charges.

Expect. The day reads not checked rather than matching or not matching. The reason appears under the disclosure titled with the count, which says in the same line that missing information is never treated as a problem. A structural blind spot goes to a second disclosure instead, and is never a finding.

Verify. The per-day reason names what was missing, for example billed cost with no local telemetry coverage for this day. A Bedrock or Azure window carries the precision-ceiling note saying which detectors that provider can never produce.

A provider pull that failed

Configuration. Two providers configured, one of whose usage or cost pull errored for this window.

Expect. An amber panel above the day chips saying the comparison below is incomplete, with the failure listed. The other provider's comparison is unaffected: one failure never takes down another.

Verify. The panel spells out the consequence, that anything the unreachable provider bills is missing from the numbers, so a clean result there does not mean a clean bill. Hosted, the same repository shows a failed pull count in its row.

Data and cost

What is captured
Per day and model: the token counts already captured from local Claude Code telemetry, and the provider usage and cost reports pulled for the window. Optionally, per-request rows under the repository brain's claude-code-usage-requests directory, one file per day, deduplicated by request id. The findings, the coverage and the day ledger are computed per request by the read and are not stored by it. The daily sweep writes one notification per paged finding, one billing_guard.finding event per finding it sees, and one billing_guard.sweep event per completed sweep, findings or not.
Who can see it
Local by default. The hosted page shows only the events a connected desktop synced, grouped by repository rather than by developer, because the sweep is a per-repository daily job. The Admin key never leaves the machine, never reaches the browser and never appears in an error string: the connect route returns configured, source and a masked hint and nothing else. At rest the key is protected by the data home's owner-only permissions, which is the honest guarantee here, not encryption.
How long it is kept
Not provided. There is no retention knob for anything Billing Guard writes. The per-day request files and the telemetry stay where they are written. The hosted page reads a 30-day window and caps a read at 5000 events, and says so on the page when that cap cut the window.
What leaves the machine
Your Admin key goes to that vendor's own read-only usage and cost endpoints and nowhere else, on a bounded request. Draft dispute opens your mail client with a draft; it sends nothing, and RepoOps never contacts a provider on your behalf. The evidence bundle downloads to your machine. Syncing the sweep events to a connected team is the only RepoOps egress.
What it costs
No model call. Detection is deterministic code, which is why the command line shell is described as zero model tokens by design. A resolved Admin cost report is reused for the memo window, five minutes by default, before another pull goes out.

When the result differs

SymptomLikely causeNext action
Every day reads not checked.No provider is configured, or the machine captured no telemetry for those days.Connect an account under Your provider accounts. If a key is already connected, open the days it could not check and read the reason on each.
The findings table says nothing to report yet.No provider key, so there is no invoice side to compare against.Connect an Anthropic or OpenAI account. Until then the page is showing your own side only, which is the honest half.
A pasted key is refused before anything is stored.The shape check, then the live check with the vendor. Neither stores an invalid key.Read the message. An Anthropic key must start with sk-ant- and come from Create Admin Key; an OpenAI admin key starts sk-admin-, and a project key cannot read organization usage.
The OpenAI box refused an Anthropic key.Deliberate. The two paste boxes sit one under the other, and the narrowed check stops a live Anthropic credential being sent to OpenAI as a bearer token.Paste it into the Anthropic box above.
An amber panel says a provider could not be reached.That provider's usage or cost pull failed for this window.Read the listed error. Treat the result below it as incomplete rather than clean, and re-read after the provider recovers.
Bedrock or Azure findings never name a model.Neither cost API exposes a per-model token count. It is a permanent API limit, not a gap in the reconciliation.Expect only the day-level detectors there, and read the precision-ceiling note the coverage summary carries.
The table has findings but nothing paged.Only cold-start-spend, delta-alarm and contradiction page. The other seven are visible on the tab and in the command line output by design.Read the tab. If even those three are silent, check whether the alerts flag is set to 1 and whether the app was running at 02:00.
The hosted page says no desktop is reporting yet.No connected desktop has synced a sweep. The Admin key stays on the machine, so there is no hosted path to turn this on.Connect a desktop. A repository appears once its daily sweep syncs, whether or not that sweep found anything.
A hosted repository row is badged stopped.Its last sweep is more than 48 hours old, so it is not being checked.Check that the desktop is running and the daily pass fired. A row that never reported reads not reported instead, which is a different problem.
Disable
Disconnect on the tab for a stored key, or set the alerts flag to 1 for the daily pages alone. An environment key is removed where it is configured, not from here.
Roll back
Nothing to roll back. Billing Guard changes no spend, opens no ticket and writes nothing to your provider, so there is no effect to reverse. A tolerance band you regret is one number on the tab, and it applies on the next read.
Revoke access
Disconnect removes the key from this machine. Revoking the credential itself is done where it was issued, in the Anthropic Console or on the OpenAI admin keys page, and RepoOps cannot do it for you. The key is never returned to the browser, so nothing here can leak it back.
Delete
Not provided. No route deletes a finding, the per-day request files under the repository brain's claude-code-usage-requests directory, or the synced billing_guard.finding and billing_guard.sweep events. The hosted page reads a 30-day window, so a row leaves it by ageing out of that window.

Maintenance evidence

Feature id
billing-guard (spine leaf billing-guard)
Owner
Billing Guard program (WS-1 engine and WS-2 detectors, WS-3 and WS-4 surfaces, WS-5 shells, the 2026-07-13 hardening, WS-E1 to WS-E7 and WS-H1 to WS-H4c; docs/plans/billing-guard.md), plus the P5 cost-lens slice, LDG-0688. Guide: LDG-0717.
Supported product version
RepoOps v0.3.1
Last verified
2026-09-15, read against origin/main at 52366bb6d; labels read from the served tab source (public/billing-guard.html, public/llm-cost.html and public/lib/pillar-landing.js) and from the hosted page source; the ten detector signatures counted in lib/billing-guard.mjs and checked against website/lib/billing-guard-detectors.ts; defaults, caps and env keys read from lib/billing-guard.mjs, lib/routes/billing-guard.mjs, lib/since-days.mjs and .env.example.
Example fixtures
lib/billing-guard.test.mjs and lib/billing-guard.cold-start.test.mjs (the engine and its detectors over fixtures), lib/billing-guard-providers.test.mjs (the per-provider registry and the isolated pull failure), lib/billing-guard-alerts.test.mjs (the sweep, the dedup key, the finding and sweep events), lib/billing-guard-dispute.test.mjs (the draft), lib/billing-guard-tap-contract.test.mjs (the request-level detectors end to end over the documented ingest), bin/billing-guard.test.mjs, installer/claude-desktop-billing-guard/server.test.mjs, and on the hosted side website/lib/billing-guard-rollup.test.ts, website/lib/billing-guard-detectors.parity.test.ts (which fails when the root adds a detector the hosted catalog does not carry), website/app/team/(home)/billing-guard/sweep-tiles.test.ts and page.test.tsx.
Source references
lib/billing-guard.mjs, lib/routes/billing-guard.mjs, lib/billing-guard-providers.mjs, lib/billing-guard-alerts.mjs, lib/billing-guard-dispute.mjs, lib/provider-credentials.mjs, lib/routes/provider-credentials.mjs, lib/otel/log-receiver.mjs, lib/since-days.mjs, bin/billing-guard.mjs, public/billing-guard.html, public/llm-cost.html, public/lib/cost-summary.js, website/components/cost-summary.tsx, website/lib/ai-incident-cases/cost-summary.ts, public/lib/case-sections.js, lib/incident-source-contract.mjs, website/lib/billing-guard-detectors.ts, website/lib/billing-guard-rollup.ts, website/app/team/(home)/billing-guard/page.tsx
Documentation review
Independent review requested on the slice pull request; not yet recorded.
Video review
Story script written 2026-09-15; render and review pending in the same slice.

Last updated