← All posts
July 20, 20263 Min ReadRepoOps Team

Prove it

AI coding runs on unverified claims. This week we shipped the tools that check them, and pointed the first one at our own marketing.

Scroll the AI coding timeline for ten minutes and count the claims. Ninety percent fewer tokens. Fully autonomous overnight. Beats last month's model on the benchmark. A screenshot of a very big number. Almost none of it ships with a way to check it.

That is the real state of the field. A lot of impressive numbers, and almost no receipts.

We spent the week building the receipts. The first thing we checked was our own copy.

The claim economy

The incentive is not a mystery. A claim spreads and a verification does not. "I cut my tokens 90%" gets retweeted. "Here is the holdout that proves the answer stayed correct" does not. So the field optimizes the claim and skips the check.

A number you cannot check is not a result. It is a mood.

Checks, not claims

Take token compression, the loudest pitch of the season: shrink the context, keep the quality. We did not build a compression engine. We built the thing that checks whether the saving cost you anything. A partner asserts a token saving, and RepoOps joins it to a quality signal and flags the bad quadrant, saved money and lost quality. The popular tool's own quality check validates the shape of the output, not whether the answer stayed right. That gap is the whole product.

Then the token diet itself. Everyone is posting their context-trimming wins, so we shipped one verdict, "is my token diet real," that reads your own telemetry and tells you whether the diet moved the bill or just moved a chart. Subagent token attribution, so a delegated read shows up as a measured saving instead of a guess. A cache-share check that reconciles what you observed against what you were actually billed.

And the loop. The discourse agrees the verifier is the bottleneck now. Good, so we shipped a signal that asks whether your verifier is independent or quietly grading its own homework. We shipped a per-skill trust ledger, so a skill earns trust from its record instead of your optimism. We shipped standing goals that re-check a finished invariant every day, because "done" decays.

Everyone is selling autonomy. We spent the week selling proof.

We ran the check on ourselves

Here is the part that is easy to skip, and we did not skip it. We pointed the same lens at our own product and marketing, and it found things.

We reworded four overreaching claims on the homepage. We dropped a support number we could not back. We made the product fail closed instead of open when a security key is unset and when a quota check errors. We closed dozens of launch-audit findings the same week we shipped features, because a product that checks other people's claims cannot ship its own unbacked ones.

The fastest way to lose the right to say "prove it" is to skip it on your own page.

The velocity is not the point

We also built the next four layers of the thing we wrote about last week, graph, eval, memory, and trust, in a day. Autonomous agents did the building. Anyone can point agents at a repo and watch them type, so that is not the part worth your attention.

The part worth your attention is that every one of those changes landed with a receipt: what ran, what verified it, what it cost, and a memory entry that says why. Nothing merged without passing the gate. The speed is only safe because the accountability was not optional.

That was the whole thesis, and this week was the demo. RepoOps proves you did the work, and remembers why. When the work is done by agents, that stops being a nicety and becomes the only thing between you and a fast, confident, unverifiable mess.

The wedge

The AI coding world is about to flood with agent output, most of it arriving with a claim and no receipt. The teams that win will not be the ones running the most agents. They will be the ones who can prove, for any change, that it did what it said, and who know what it cost.

It is free, it is local, and it does not phone home. Point RepoOps at the agents you already run. It reads what they cost, checks what they shipped, and remembers why, so the next time someone says "prove it," you have an answer instead of a screenshot.

repoops.ai