Docs

Price your app's LLM calls against the bill

RepoOps reads the per-call cost stream your app writes into its own repository, slices it by tier, surface, model, day, repository, session, tenant and prompt, and keeps two things apart that most cost tools sum: the saving a rule forecasts, and the saving the cost stream later measured.

For: the engineer who owns an app's LLM bill, and the operator who sets its routing

What it does, and why it helps

Your app writes one JSON line per model call into .claude/brain/ai-calls/YYYY-MM-DD.jsonl in its own repository: the surface, the tier, the model, the token counts, the cost it computed, and a 120-character prompt preview. RepoOps reads those files live from the working tree of every tracked repository, over a trailing 90 days unless you ask for another window, and aggregates them. The tab defaults to all repositories combined, because app LLM telemetry is cross-repo and the aggregator emits none of its own.

Two sections on the tab answer different questions and are never added together. Routing recommendations is a set of thresholded rules over the window: a heavy-tier surface whose average output is small, a cache hit rate under 30 percent, the same prompt repeated, a session spending five times the median. Each carries a forecast, and the tab labels it a projection. Realized savings joins a routing change that landed to the cost stream on either side of it, and reports the real cost per call before and after. A change that cost more is marked regressed rather than dropped. A change with too little data yet says it is still measuring and how many more days it needs.

The pain. The model bill moved and the invoice cannot say which call moved it. The dashboard that can say is usually a forecast engine, so the number you act on is a model of a saving rather than a saving.

The point of view. A forecast and a measurement should never share a total. Show the forecast as a forecast, then measure the same change against the recorded cost stream afterwards, and say plainly when the measurement is not available yet.

What gets easier. Finding the one surface, model, session or tenant carrying the spend, and settling an argument about whether last month's routing change paid for itself.

When it helps. An app that logs its own LLM calls to the convention, on a machine that tracks that repository. It is the production-app view; Claude Code session spend lives on Claude Code usage, and the cross-provider ledger on Provider spend.

Its limits. It prices what your app logged. A call your app does not log is invisible here, and cost_usd is computed by your app at write time, so the tab is an estimate until you connect an Anthropic Admin key and read the drift. It ranks single calls, not prompts: a cheap prompt run a thousand times shows up in the duplicate-prompt recommendation, not in Hot prompts. Realized savings needs one repository selected, because it joins one decision ledger to one cost stream.

Understand it in 30 seconds

30.1 s, captions on. Narration: Microsoft Zira Desktop (provisional voice; an approved narration source is pending).Transcript
Read the narration
  1. 0:00 The model bill rose.
  2. 0:01 Nothing on it says which call did it.
  3. 0:06 A forecast is not a saving.
  4. 0:08 RepoOps measures cost per call before and after.
  5. 0:13 Seven days each side, five calls minimum.
  6. 0:16 A change that cost more is marked regressed, and the tab says so.
  7. 0:23 Then check your estimate against the Anthropic bill.
  8. 0:26 Decide on the measured number.

Synthetic example. Read the guide

Where to find it

Where to find it

  • Desktop: localhost:4000, then LLM Cost Metrics in the sidebar, then App LLM calls under All tools, in the Evidence readers group.
  • Hosted: repoops.ai/team/app-llm-calls, from LLM Cost Metrics in the sidebar, then App LLM calls under All tools, in the Evidence readers group.
  • Keyboard: ⌘ K, then type “App LLM calls”.

When to use it

One surface is carrying the bill

Situation. Spend is up and nobody can name the cause. Several repositories are tracked and the app tags each call with a surface and a tier.

What you do. Open App LLM calls, leave Scope on All repos (combined), and read By surface (top 25 by cost) from the top row down. Then read Routing recommendations above it.

What you see. The summary cards show calls, estimated cost, input and output tokens, cache read, cache creation, and the share of calls that hit cache. A heavy-tier surface averaging under 800 output tokens over at least 5 calls raises a downgrade recommendation naming the surface and the tier pair.

What it establishes. You have the surface, its call count and its share of the window. The forecast beside it is a projection, and the tab says so; nothing has been saved yet.

Checking whether last month's routing change paid

Situation. A tier mapping changed in one repository weeks ago and the recommendation that prompted it predicted a saving per call.

What you do. Select that single repository in Scope and read Realized savings. Leave it alone if the row says measuring.

What you see. A measured row prints the before and after windows (7 days each by default), the calls and total cost in each, the cost per call in each, and the realized delta multiplied by the measured call volume. A change whose cost per call rose is labelled regressed and printed in red. A thin row says how many of the 5 required calls it has and roughly how many more days it needs.

What it establishes. A measured number, or an honest statement that there is not one yet. Neither is the forecast that opened the change, and the two sections sit apart on the page for that reason.

Before you start

Supported versions
RepoOps desktop v0.3.1, the release this guide was read against. The app whose calls you want to see must write the ai-calls convention documented in docs/llm-routing-convention.md.
Where it runs
Local: LLM Cost Metrics, then App LLM calls under All tools, in the Evidence readers group. Hosted: LLM Cost Metrics, then App LLM calls under All tools, in the Evidence readers group (/team/app-llm-calls), renders the published ai-calls metadata for the repositories you can see, grouped by tier and by surface, plus the routing findings the desktop pushed and the team-side event-store spend.
Permissions
The local tab is the localhost dashboard and takes no sign-in. The hosted page needs a signed-in team member, and the read is narrowed to the repositories that member can open. The billing reconciliation panel on the hosted page is owner and admin only, resolved on the page and again on the route behind it.
Connections
For production traffic from a deployed app: a telemetryExport block in repos.config.json pointing at that app's GET /api/ai-calls/log/export, with a bearer token. For the estimate-versus-actual row: an Anthropic Admin key, connected on the Billing Guard tab or set as ANTHROPIC_ADMIN_KEY. For the gateway probe: OPENROUTER_API_KEY. For the hosted page: a bound device publishing brain snapshots.
Plan
The desktop tab has no plan gate; the pricing capability map has no App LLM calls row. Seeing the same data rolled up on the hosted dashboard follows the hosted dashboard tiers, and cross-repository, cross-developer aggregation is the Team line on that map.

Configure it

  1. Have the app write the convention.

    One JSON line per call into .claude/brain/ai-calls/YYYY-MM-DD.jsonl, carrying ts, surface, tier, model, the tokens block, cost_usd and a request_id. The directory belongs in .gitignore: RepoOps reads the working tree, not a commit. Without tier and surface the totals still add up, but every per-tier and per-surface view has nothing to group by.

  2. Point RepoOps at the deployed app, if the calls happen in production.

    Implement GET /api/ai-calls/log/export on that deploy behind a bearer token, then give the repository a telemetryExport block in repos.config.json with the same url and token. The puller asks for rows since its cursor, pages at 5,000 rows up to 50 pages, and appends each row to the day file it belongs to. An implemented endpoint with no telemetryExport block sends nothing, because the puller skips a repository that has none.

  3. Read By surface first, then the recommendations.

    By surface is already ranked by cost. Routing recommendations sits above it and is labelled as forward estimates. A recommendation that names a tier pair is the one the LLM routing control tab can turn into a pull request against that repository's own llm-routing.json.

  4. Select one repository before reading Realized savings.

    The measured before and after join one repository's routing decision ledger to that repository's own cost stream. On All repos (combined) the section says to pick one rather than joining ledgers that do not belong together.

  5. Connect an Anthropic Admin key for the estimate-versus-actual row.

    Connect Anthropic account on the Billing Guard tab, or set ANTHROPIC_ADMIN_KEY (prefix sk-ant-admin01-, which a regular API key is not). The environment value wins over a connected one. The comparison reads its own 30-day estimate over the same scope, so it never divides a 90-day estimate by a 30-day bill.

  6. Tune a threshold only when a rule is firing wrongly.

    Every cost rule ships defaults and a validated schema. Open the Rules tab, edit the rule's fields and press Save params; the value persists to config/rules/<id>.json and merges onto the defaults. The playground on the same tab tests a candidate against real data without saving it.

SettingWhereA sensible choiceWhy it matters
telemetryExportrepos.config.json, on the repository entry{ url, token, intervalMs }, with the token as ${env:AI_CALLS_EXPORT_TOKEN}No block means the puller skips that repository, so a working export endpoint still sends nothing.
intervalMsinside the telemetryExport block300000 (the default, 5 minutes)How often the puller asks for new rows. Each request is bounded at 30 seconds.
AI_CALLS_EXPORT_TOKENthe .env file, referenced from repos.config.jsonthe same bearer token the deployed app checksIt keeps the credential out of the committed config file.
ANTHROPIC_ADMIN_KEYthe .env file, or Connect Anthropic account on Billing Guardan Admin key (sk-ant-admin01-), from the organization that owns the API key the app calls withWithout it the Actual Anthropic billing section says it is not connected. A key from another organization reports a real bill as $0.00, and the panel says so above the numbers rather than calling it agreement.
REPOOPS_BILLING_TTL_MSthe .env file300000 (the default)How long a fetched cost report is reused. The tab reloads every 30 seconds, and the Admin report changes at most daily.
OPENROUTER_API_KEYthe .env fileunset unless you want the probePrice against gateway sends one short request through the gateway on each of the two models and compares the cost it reports. Without the key the button says not measured and names the operator step, and the table stays modeled.
cost.downgrade-tier-heavyRules tab, the rule card, Save paramsmaxAvgOutput 800, minCalls 5 (the defaults)Raise maxAvgOutput and more heavy-tier surfaces are proposed for the mid tier; raise minCalls and a low-volume surface stops being proposed at all.
cost.enable-prompt-cachingRules tab, the rule card, Save paramsmaxHitRate 0.3, minCalls 50 (the defaults)The global cache-health finding. Under a 30 percent hit rate across more than 50 calls it says the shared system block is not being cached.
cost.user-cost-outliersRules tab, the rule card, Save paramsmedianMultiple 5, absoluteFloorUsd 1, minUsers 3, minCallsPerUser 10 (the defaults)Flags a session spending a multiple of the median, which is the usual shape of a retry loop. The median is used rather than the mean so one heavy session cannot hide the rest.
REPOOPS_TELEMETRY_MAX_FILE_MBthe .env file512 (the default)The byte cap on one day-file read. Over the cap the reader streams a bounded prefix of whole lines and logs that it did, so a very large day reads as a subset.
REPOOPS_AI_CALLS_CACHE_MAX_MBthe .env file128 (the default)The memory budget for the parsed day-file cache. Past days are immutable, so only today's file is re-read.
ⓘ
To stop or undo
There is no switch on the tab. To stop the pull, remove the telemetryExport block from the repository entry. To stop the bill read, unset ANTHROPIC_ADMIN_KEY and disconnect Anthropic on Billing Guard. To stop the gateway probe calling out, unset OPENROUTER_API_KEY. To stop the app writing calls at all, stop writing the file: RepoOps only reads it.

What you should see

A tracked app that logs the full convention

Configuration. Scope on All repos (combined), calls carrying surface, tier and model.

Expect. Six summary cards, then By tier, By surface (top 25 by cost), By model, Model upgrades, By day, By repo, By user (top 15 by cost), By tenant (top 25 by cost) and Hot prompts (top 25 by cost). Model upgrades lists only older Claude models in the same family as a current one, ranked by spend.

Verify. The header line under the title reads the window, the call count, the estimated cost, the token split and the cache hit rate. Every bucket table repeats calls, input, output, average in, average out, cache read and cost, so a row's share of the window is readable without arithmetic.

A repository with no app telemetry

Configuration. Scope narrowed to a repository that only carries Claude Code session events.

Expect. The tab says that repository is not sending app LLM call telemetry yet, links the setup guide, and dashes out every section. It does not render a page of zeros. Rows written by the Claude Code hook are skipped at parse time, so they never count as zero-token calls on an unknown surface.

Verify. The subtitle names the repository and, when there were session events, says how many. Switch Scope back to All repos (combined) and the sections return.

Nothing measured yet, or a change that regressed

Configuration. One repository selected, a routing change landed less than a week ago or with fewer than 5 calls in the after window.

Expect. The row reads measuring, with the calls it has of the 5 it needs and roughly how many more days at the current rate. Once both windows qualify, the row prints the before and after cost per call. If the cost per call rose, the row is labelled regressed and the realized total is negative.

Verify. The banner above the rows prints the total measured to date, how many changes are measured and how many are still measuring. Nothing in that section is a forecast; the forecasts are in the section above it.

Data and cost

What is captured
One line per call, written by your app, in its own repository at .claude/brain/ai-calls/YYYY-MM-DD.jsonl: ts, request_id, session_id, optional tenant, surface, tier, model, fallback, the tokens block (input, output, cache_read, cache_creation), cost_usd, duration_ms, max_tokens, stop_reason, prompt_chars, a 120-character prompt_preview, cwd and ok. RepoOps writes none of it; it reads the file and, with a telemetryExport block, appends rows pulled from your deploy. The Tool-step cost section additionally reads the session transcripts in the aggregator's own checkout, which is why it is blank on a machine that mirrors repositories rather than working in them.
Who can see it
Local by default, on localhost. With a bound device publishing brain snapshots, the ai-calls files are published as metadata: each line is parsed, ten named prose keys are dropped recursively (transcript, prompt, user_prompt, assistant_text, raw, raw_content, messages, completion, content, text), the result is run through the secret redactor, and the snapshot is capped at 1 MiB per file, 5 MiB in total and 500 files. On the hosted page a member sees only the repositories they can open, and the reconciliation figures are owner and admin only.
How long it is kept
Not provided. No retention window and no purge route exist for the day files; they live at <repo>/.claude/brain/ai-calls/YYYY-MM-DD.jsonl for as long as you keep them. The tab reads a trailing 90 days by default, and the bill comparison a fixed 30, but that is a read window, not a deletion policy.
What leaves the machine
The tab itself calls nothing out. The Anthropic Admin read goes to api.anthropic.com on your Admin key, memoized for five minutes. Price against gateway sends one short request per model through OpenRouter on your key, which costs what those two calls cost. The puller calls your own deploy. The published snapshot is the only RepoOps egress, and only on a bound device. One honesty note: prompt_preview is not one of the ten keys the snapshot stripper drops, so the 120-character preview a row carries is published in the snapshot file even though the hosted page does not render it.
What it costs
Reading and aggregating is local file work and costs nothing. The gateway probe is the only paid call RepoOps makes here, and only when you press the button with a key set. The cost_usd on every row is computed by your app at write time, which is why the tab calls its own figure an estimate and offers the drift row against the real bill.

When the result differs

SymptomLikely causeNext action
The tab says no app LLM call telemetry has been captured yet.No tracked repository has a .claude/brain/ai-calls/ directory with rows in the window.Point the app at the convention, or add a telemetryExport block so the puller fills the files from your deploy.
A repository shows calls but $0.00 and zero output tokens.Its JSONL carries session events rather than app LLM calls.Switch Scope to All repos (combined) or another repository. The tab names this case rather than rendering zeros.
By tier is all one row called unknown.The rows carry no tier field.Add tier to the call payload. Older rows keep none, so the group stays until the window rolls forward.
Actual billing reads $0.00 while the estimate is real.The Admin key belongs to a different organization than the API key the app calls with.Create an Admin key in the organization that owns that API key, set ANTHROPIC_ADMIN_KEY to it and restart. The panel prints this reason above the numbers.
Realized savings asks for a single repository.Scope is on All repos (combined), and the join needs one decision ledger and one cost stream.Pick the repository whose routing changed.
Price against gateway reports not measured.OPENROUTER_API_KEY is unset, or the gateway answered without a usable cost.Read the reason the button prints. The savings stay modeled; no number is invented in either case.
Tool-step cost says the tool steps cannot be sized.There are no session transcripts on this machine.Nothing to fix from the tab. It reads the aggregator's own checkout, and a mirror has no transcript directory.
Disable
Remove the telemetryExport block to stop the pull, unset ANTHROPIC_ADMIN_KEY and disconnect Anthropic on Billing Guard to stop the bill read, and unset OPENROUTER_API_KEY to stop the probe. The tab has no on and off control of its own, because it only reads files your app already writes.
Roll back
Not provided. RepoOps does not revert a routing change. A landed change is a commit against that repository's own .claude/brain/llm-routing.json, so revert it in git; the realized-savings pass then measures the reverted state on its next run.
Revoke access
The export bearer token lives on your deploy as AI_CALLS_EXPORT_TOKEN and is referenced from repos.config.json by ${env:VAR}: rotate it in both places. The Anthropic Admin key is revoked in the Anthropic console, and disconnecting on Billing Guard removes the stored copy. Unbinding the device stops any further snapshot reaching the hosted page.
Delete
Not provided. No route deletes a captured call. The files are <repo>/.claude/brain/ai-calls/YYYY-MM-DD.jsonl, with the puller cursor in .puller-state.json beside them; removing a day file removes it from every local view. The hosted side holds whatever the desktop last published for that repository.

Maintenance evidence

Feature id
app-llm-calls (spine leaf app-llm-calls)
Owner
App LLM call telemetry (docs/features/app-llm-calls.md; Phase 5 token optimization, with the realized-savings ledger and the PL3 inline-gateway slice). Guide: LDG-0717.
Supported product version
RepoOps v0.3.1
Last verified
2026-09-15, read against origin/main at 52366bb6d; labels read from the served tab source (public/app-llm-calls.html, public/rules.html, public/settings.html), defaults read from lib/ and .env.example, the hosted behaviour read from website/app/team/(home)/app-llm-calls/page.tsx and website/lib/app-llm-calls.ts.
Example fixtures
lib/ai-calls.test.mjs, lib/ai-calls.rollup.test.mjs and lib/ai-calls.events.test.mjs cover the aggregation and the tenant bucket; lib/realized-savings.test.mjs covers the measured and measuring paths and the gateway probe; lib/model-versions.test.mjs pins every entry in LATEST to a priced row; lib/rules/cost.test.mjs and lib/rules/cost.downgrade-model.test.mjs cover the thresholds; lib/tool-call-cost.test.mjs and lib/tool-call-cost-read.test.mjs cover the carry model; website/lib/app-llm-calls.test.ts covers the published-metadata parser. scripts/seed-telemetry.mjs writes a demo stream.
Source references
lib/ai-calls.mjs, lib/routes/ai-calls.mjs, lib/realized-savings.mjs, lib/routes/recommendations.mjs, lib/anthropic-billing.mjs, lib/routes/anthropic-billing.mjs, lib/model-versions.mjs, lib/telemetry-puller.mjs, lib/brain-publisher.mjs, lib/rules/index.mjs, public/app-llm-calls.html, docs/llm-routing-convention.md, website/lib/app-llm-calls.ts
Documentation review
Independent review requested on the slice pull request; not yet recorded.
Video review
Narrated story rendered and published 2026-09-26 (render 505f098e1db8, LDG-1012) with the breadcrumb LLM Cost Metrics, checked against main at b0bb02812. Six frames, the captions and the transcript were reviewed by the authoring agent, not an independent reviewer; the audio was not listened to by a person. Narration is the provisional Windows voice until LDG-0721.

Last updated