LLM Cost Metrics

See the cost of reaching
a supported answer.

Follow captured usage through investigation, verification, and follow-up. Distinguish priced usage, estimates, and activity whose cost is unknown.

One investigation

Account for the work behind the conclusion.

Keep usage attached to the case and the session that produced it. A rollup and a receipt that describe the same activity are counted once.

Usage receipt / illustrativeSynthetic example
Deterministic evidence projectionNo model call
Assisted diagnosisCaptured usage
Missing model rateUnpriced

A cost shows the price version it was computed with. An unpriced model reads unpriced, never a fallback rate.

Start with one question

What happened in this change?

Follow an example from the instruction to the verified fix, then install RepoOps and open your own.

Implementation details & evidence

Know what the work and its resolution cost.

Investigate model usage by repository, developer, and session. Follow measured spend back to work, keep unknown model rates unpriced, and reconcile comparable records against provider billing.

Billing Guard detects and reconciles billing discrepancies. It does not cap or throttle spend.

The surfaces

Five ways spend ties back to work

Telemetry and Claude Code usage

Telemetry is a combined view of tokens and spend across every tracked repo, day by day, captured locally and retained as long as you want. Claude Code usage drills into per-session token spend and breaks the cost down by model, with heuristic and LLM-driven recommendations on where you are overspending. Spend to outcome splits the same dollars per developer. Together they mean you stop guessing what AI cost this week.

Cost per PR

Merged pull requests get priced as a group, attributed LLM cost across the merged PRs in a window divided by their count, so spend ties to the work it produced rather than a monthly total. You can see the dollars behind a session and a developer. Billing Guard separately reconciles your telemetry against real Anthropic billing.

Financial Model

The full model: revenue, cost of goods, margin floors, and the levers that move them. A team sees the economics of its AI usage, not just the invoice, so a spend decision can be made on the numbers.

Efficiency scorecard

A composite Spend-Efficiency grade over six levers: defaults, routing, caching, context, visibility, and ROI. Routing savings open as a pull request, then get measured after they land, so a saving is measured rather than promised.

defaults / routing / caching / context / visibility / ROI

Billing Outcomes

A per-client, one-page invoice export with proof-of-work line items, so you can defend every billable AI minute to a client or to finance. Each line ties a dollar to the work it shipped.

How it helps you

  • Know what AI cost this week, per developer and per repo, with a defensible split for chargeback.
  • Tie measured spend to the pull request and outcome it shipped, not a lump-sum bill.
  • Turn a cheaper-model insight into a reviewable pull request, then confirm the saving landed.
  • Hand a client or your CFO an invoice that shows what the spend shipped.