← All posts
July 20, 20264 Min ReadRepoOps Team

After graph, the ladder keeps climbing: eval, memory, trust

Prompt to context to harness to loop to graph. Each rung makes a bigger unit programmable. Here are the next three, and the three things every rung so far has left out.

For four years the job of getting useful work out of a model has had a name, and the name keeps changing. That is not fashion. It is the unit of work getting bigger.

Prompt engineering was about wording: give the model a role, break the task into steps, ask it to think first. The unit was one message. Context engineering moved the question to what the model can see. Harness engineering connected it to a place it could act. Then, in June 2026, loop engineering changed the unit again: stop typing prompts, design the loop that prompts the agent for you. Discover, plan, execute, verify, repeat. The line that stuck was that the verifier is the bottleneck now, not the generator. And on July 18, graph engineering landed: loops are one agent's behavior, graphs are the org structure connecting many agents, wired into nodes, edges, and shared state.

Read that top to bottom and the shape is hard to miss. Word, then agent over time, then a whole organization of agents. Each rung is someone realizing the last unit was too small.

What every rung so far has left out

Three things stay invisible on that ladder, and all three matter more, not less, as the unit grows.

Cost. A loop runs. A graph of loops runs. Nobody on the ladder asks what it spent. Once you are running an organization of agents, "which of these was worth running" is the first question anyone with a budget asks, and the ladder has no way to answer it.

Proof. Loop engineering names the verifier as the bottleneck, then throws the verifier's output away. The check passes or fails, and the result evaporates when the loop ends. There is no durable record that says this ran, this is what verified it, and this is what it cost. The proof is discarded at the moment it becomes valuable.

Compounding. The loop lives inside one project. The graph lives inside one run. When either finishes, what did the organization learn? On the ladder as drawn, nothing survives. Every run starts where the last one started.

These are the difference between a demo and a business. They are also the ground RepoOps has stood on the whole time. We do not run your loops. We are the memory, the meter, and the proof underneath whoever does.

The next three rungs

If the ladder climbs by zooming out, the next three rungs are not hard to call. They zoom out along a new axis: what becomes the scarce, versioned asset once the unit is a running organization.

Eval engineering, next. Once the model is a commodity and the verifier is the bottleneck, the verifier stops being a throwaway command and becomes an asset you version and track. If loop engineering is designing the loop the agent runs, eval engineering is designing the bar the loop has to clear. Quality becomes something you can name, store, publish, and hold a run against.

Memory engineering, after that. When graphs of agents run loops against versioned bars, the scarce input is no longer the model, the structure, or the check. It is accumulated, verified experience. What to remember, what to forget, what to reuse, and how knowledge moves between agents and across time. The brain becomes the programmable asset. The winner here is whoever already joins knowledge to the cost of learning it.

Trust and market engineering, last. When agent organizations start working across company lines, one org's graph adopting another's specialist structure, buying its evals, licensing its verified memory, the product becomes provenance and price. Where an artifact came from, whether you can trust it, and what it costs to run.

Eval is the quality of one loop. Memory is the experience of the whole organization. Trust is the value that moves between organizations. Each is a further zoom-out, and each turns one of the three ignored things into a product.

What RepoOps is building

We do not build a loop runner, and we do not build a graph execution engine. Tools that run graphs already exist. We observe, remember, verify, and price the runs. Four layers, in the order the discourse will reach them.

Graph parity. A multi-agent run becomes one mission: one execution across many agents with its own cost rollup, its own shared-state timeline, and one verified outcome receipt. Not N disconnected transcripts. One accountable thing with a cost and a verdict.

Eval engineering. The verifier becomes a versioned bar: a named case set plus a pass threshold, tracked over time, with a view that trends pass rate per loop and catches a regression before it ships. Verified memory turned into a leaderboard, and a bar a team can adopt instead of reinventing.

Memory compounding. A whole graph run contributes one durable lesson that carries its cost, and a single view shows what the organization has learned and what it cost to learn it. Cost-fused memory is the one thing the ladder cannot produce on its own, because nothing on the ladder joins knowledge to spend.

Provenance and market. Proven graph structures and evals become tradeable, each carrying a provenance receipt and a cost benchmark, so a buyer sees what an artifact does and what it costs to run before adopting it. When memory crosses from one org to another, its lineage travels with it.

The wedge

The bet is not that graphs matter. Everyone can see that graphs matter. The bet is that the three things the ladder ignores, cost, proof, and compounding, become the whole game once the unit is a running organization.

RepoOps proves you did the work, and remembers why. Every layer above is a way of doing those two things at a larger scale.

It is free, it is local, and it does not phone home. Point RepoOps at the agent you already run. It reads what your loops and graphs cost, verifies what they produced, and remembers what the work taught you, so the next run starts ahead of the last one.

repoops.ai