Watch · Runtime health

See which install is flapping right now

Every RepoOps install samples its own runtime every five minutes. When a signal crosses its threshold, a pass writes a finding with the numbers behind it. Runtime health groups those findings per install over your team's event store, so you read which machine is climbing instead of a fleet average.

For: the team lead watching a fleet of installs, and the engineer who owns the machine a finding names

What it does, and why it helps

The desktop app samples its own resident memory, its pid and this tick's duration every five minutes and appends them as runtime.metric events. A watcher derives three signals from those samples: the steepest resident-memory climb that does not straddle a restart, how often the process recycled per hour, and an endpoint p95. The self-audit pass classifies each signal green, yellow or red and fires only when one is yellow or red. A healthy install writes nothing at all.

A finding that fires carries six fields: a title, a narrative body, the named root cause, a fix hint drawn from a frozen menu of real levers, the evidence map of sampled numbers, and a severity. It is written as an action.proposed event, which always waits for a person. With cloud sync on, that event streams to your team's event store stamped with the install's machine hash, and this page groups the last 30 days of findings into one card per install: the count, the newest finding's severity and title, its root cause, its measured samples, its suggested fix, and the measured before and after of a fix somebody applied. When no fix was applied, or none could be measured, the card says that instead of scoring a win.

The pain. A fleet of installs degrades one machine at a time. A team-wide average hides it, and the person who notices is whoever happens to be sitting at the slow laptop. Finding the machine means opening each one.

The point of view. A health claim is worth what its numbers are worth. Group the findings by the install that produced them, carry the samples that fired each one, and when the outcome of a fix was not measured, say so rather than inferring an improvement.

What gets easier. Naming the machine. One card per install, newest finding first, with the root cause, the sampled numbers, the remediation the pass chose, and a before and after table when a fix was applied and re-sampled.

When it helps. A team on a paid hosted plan whose installs are bound to it with cloud sync on, and at least one install whose resident memory is climbing or whose process is recycling often.

Its limits. It reports; it fixes nothing. Two of the three signals can fire today: nothing in the shipped code writes the endpoint-timing samples the endpoint p95 needs, so that source stays absent. The thresholds are fixed constants with no per-account override. Findings carry no repository scope, so a member narrowed to a subset of repositories sees none of them. The read is capped at 5,000 rows over 30 days and says so when it fills.

Understand it in 30 seconds

30.1 s, captions on. Narration: Microsoft Zira Desktop (provisional voice; an approved narration source is pending).Transcript
Read the narration
  1. 0:00 One machine in the fleet is climbing.
  2. 0:02 The team view shows an average.
  3. 0:06 Group every finding by the install that fired it.
  4. 0:09 Keep the numbers with it.
  5. 0:13 Each card names the root cause, the measured samples, and the suggested fix.
  6. 0:18 An applied fix shows before and after.
  7. 0:23 Unmeasured is said plainly, never scored a win.
  8. 0:26 The guide has the caps.

Synthetic example. Read the guide

Where to find it

  • Hosted: repoops.ai/team/runtime-health, from Team & settings in the sidebar, then Runtime health under Workspace utilities.
  • Desktop: hosted only.

When to use it

One machine in the fleet is climbing

Situation. Two engineers report the desktop app feeling slow at the end of the day. The fleet has nine bound installs and you do not know which ones are affected.

What you do. Open Runtime health. Read the install cards top to bottom; they are ordered by the most recent finding.

What you see. Two cards carry findings. The newest reads severity ERROR with the title Self-audit: memory-recycle flap, the root cause line memory-recycle flap, and the Measured chips carry rss-slope-mb-per-min and recycle-rate-per-hour with their values. The Suggested fix line names the levers from the pass's frozen menu.

What it establishes. You know which installs fired and on which numbers. The other seven installs streamed no finding, which means their signals stayed green, not that nobody looked.

A fix was applied and you want the number

Situation. Somebody approved the finding on the affected machine last week and raised the memory ceiling. The question is whether the signal moved.

What you do. Open the same install card and read the block under the finding.

What you see. The card reads one of five things: no fix applied yet, a fix applied whose before and after could not be sampled, or a verdict sentence with the Signal, Before, After and Change table. The verdict is improved, not improved, or measured with no significant change.

What it establishes. An answer or an honest absence. A delta of exactly zero is recorded as no measured change, never as an improvement, and a source that appears in only one of the two snapshots is dropped rather than filled in.

Before you start

Supported versions
RepoOps desktop v0.3.1, the release this guide was read against, on each install whose findings should appear. The hosted page is the repoops.ai deployment; it needs no local install to render.
Where it runs
Hosted only: repoops.ai/team/runtime-health, in the Attribution group of the sidebar. The desktop app carries a pointer row under On repoops.ai that opens the same page. The desktop equivalent of one install's own finding is the Self-audit section on the Brain health tab.
Permissions
Any signed-in member of the team can open the page; there is no owner or admin gate on it. A member the admin narrowed to a subset of repositories will see nothing here, because these findings carry no repository scope and an active include-filter drops them. It opens from Team & settings, under Workspace utilities.
Connections
Each install bound to the team (Connect to team in the desktop Settings) with cloud sync on, so its events reach the team store. ANTHROPIC_API_KEY on an install only if you want the narrated version of its findings; without it the finding is deterministic. CRON_SECRET on the hosted deployment if the hosted runtime should audit its own cron fleet.
Plan
A paid hosted plan: Solo Hosted or above, with the subscription active, trialing or past due, and the reader inside the seats the team pays for. A free local install never reaches this page.

Configure it

  1. Bind each install and turn cloud sync on.

    Desktop Settings, Connect to team, then the Cloud sync (hosted) card and the Enable hosted cloud sync checkbox. Cloud sync is off by default for an unbound device; a device with a team binding and a device token syncs by the team default unless the stored teamCloudSyncDefault is false. Without it the install produces findings that nobody else ever sees.

  2. Leave the metric tick alone.

    The five-minute sample is always on and costs nothing: resident memory in MB with the pid, uptime and heap, plus that tick's own duration. It is bounded to that cadence on purpose, never per request, so the audit cannot become the cost it measures.

  3. Run a pass, or arm the schedule.

    The scheduled pass is off unless REPOOPS_SELF_AUDIT_SCHEDULE is exactly 1 in the data directory's .env. Until then the only trigger is the Run a pass now button on the desktop Brain health tab, or POST /api/self-audit/run, which works either way.

  4. Open the page on the hosted dashboard.

    repoops.ai/team/runtime-health. Open Team & settings in the sidebar, then Runtime health under Workspace utilities. The deep link and the search box find it too.

  5. Approve a finding if you want the before and after.

    The finding is always held for a person. Approve it on the desktop Actions queue tab, under Held actions awaiting human approval. Approving writes the approval audit event; the deferred step fires on the next approvalQueueTick, within five minutes, re-samples the same signals and writes the outcome this page reads. Nothing about that step changes a setting or a file.

  6. Set CRON_SECRET on the hosted deployment.

    The hosted runtime audits its own cron fleet every six hours through GET /api/cron/self-audit and writes a finding into each team that streamed an event in the last seven days, under the install label hosted runtime (repoops.ai). With CRON_SECRET unset the route answers 401 and that bucket never appears.

SettingWhereA sensible choiceWhy it matters
cloudSyncDesktop Settings, Cloud sync (hosted), Enable hosted cloud syncon for every install whose findings the team should seeAn install with it off pushes nothing, so it is absent from the page rather than shown as healthy.
teamCloudSyncDefaultstored account settings, no desktop controlleave unset, so a bound device with a device token syncs by defaultSet to false it turns that default off, and only an explicit cloudSync true still pushes.
REPOOPS_SELF_AUDIT_SCHEDULEthe data directory's .env on each install1 to run the pass on the five-minute tick, unset to run it by handAny value other than the string 1 reads as off; the on-demand route stays available either way.
ANTHROPIC_API_KEYthe data directory's .env on each installset only if you want a narrated findingWith no key the finding is deterministic and free. With a key the pass makes at most two calls on your key, and a narration the judge does not pass is discarded in favour of the deterministic text.
REPOOPS_DAEMON_MAX_EVENTS_PER_TICKthe data directory's .env on each installunset (5000)Bounds the events one push tick reads into memory; a backlog drains across ticks rather than in one read.
CRON_SECRETthe hosted deployment's environmentsetThe hosted self-audit cron rejects an unauthenticated call, so an unset secret means the hosted install bucket never appears.
REPOOPS_NEON_RETENTION_DAYSthe hosted deployment's environmentunset, which is 30Hosted telemetry rows older than this are hard-deleted daily. Thirty is the default and the ceiling, so a lower value shortens what the page's 30-day window can show and a higher one changes nothing.
ⓘ
To stop or undo
Untick Enable hosted cloud sync in the desktop Settings, or press Disconnect under Connect to team, and that install stops pushing; the disconnect message says the binding is forgotten here and revoked on the server, and activity history is retained. Rows already in the team store stay until the retention sweep deletes them. Unsetting REPOOPS_SELF_AUDIT_SCHEDULE stops the scheduled pass and leaves the five-minute sample running.

What you should see

A flapping install, cloud sync on

Configuration. A bound install with cloud sync on, the schedule armed, resident memory climbing faster than 300 MB a minute or the process recycling more than four times an hour.

Expect. Within a few ticks a card appears for that install, labelled install plus the first ten characters of its machine hash, with the severity ERROR, the title, the root cause, the Measured chips and the Suggested fix line.

Verify. The two count cards at the top read Installs with findings and Findings (30d). The same finding is on that machine's own Brain health tab, under Self-audit (this dashboard's own runtime).

A quiet fleet

Configuration. Every install bound and syncing, every signal green.

Expect. No self-audit findings streamed yet, with the note that an install's fleet health shows up once its pass fires and cloud sync is on.

Verify. Press Run a pass now on an install's Brain health tab. A healthy install reports no finding rather than writing one, so the empty state is the correct answer, not a broken read.

A fix approved and measured

Configuration. A finding approved on the desktop Actions queue, with the evidence snapshot still on the proposal and fresh samples readable at apply time.

Expect. The install card gains a verdict sentence and a table with Signal, Before, After and Change, followed by the note that every signal here is lower is better, so a negative change is the improvement.

Verify. The Change column is signed. A source that was sampled before but not after is missing from the table rather than shown as zero, and a table with no rows says no signal was sampled both before and after.

Data and cost

What is captured
Per install, one runtime.metric sample every five minutes (resident memory in MB with the pid, uptime and heap, plus that tick's own duration), appended to .claude/brain/events/ under the repository brain. A pass that fires appends one action.proposed event carrying the title, the body, the root cause, the fix hint, the evidence map and the severity. Approving it appends the outcome, with the before and after rows when they could be sampled.
Who can see it
Local until cloud sync is on for that install. Then the already-redacted event store streams to the team over the device token, and any member of the paid team can open the page. A member narrowed to a subset of repositories sees none of these rows, because they carry no repository scope. Personal repositories are excluded from the read by the same narrowing.
How long it is kept
The page reads a 30-day window, capped at 5,000 rows of the two kinds it needs, and prints a lower-bound note when the cap is reached. Hosted rows are hard-deleted after REPOOPS_NEON_RETENTION_DAYS (default and ceiling 30) by the daily events-retention cron. The local day files under .claude/brain/events/ have no purge route.
What leaves the machine
With cloud sync on, gzipped batches of the redacted event store go to the ingest endpoint (https://www.repoops.ai/api/events/ingest by default), at most 1 MB of NDJSON per batch before compression and 5,000 events per tick, once an hour. The ingest gate is 60 requests a minute with a 1.5 times burst and 100,000 events a day per account by default. With ANTHROPIC_API_KEY set, a pass that fires also calls Anthropic on your own key before the finding is written.
What it costs
The sampling, the watcher and the deterministic finding make no model calls. With a BYOK key, a pass that fires makes at most two calls on your key, one for the narration and one for the judge that scores it, and a pass on a healthy install makes none. The hosted read is a query over your team's own rows with no per-view charge.

When the result differs

SymptomLikely causeNext action
The page says No self-audit findings streamed yet.No install has streamed a finding: cloud sync is off, the device is unbound, or every signal is green.Check Connect to team and Enable hosted cloud sync in the desktop Settings, then press Run a pass now on the Brain health tab.
Runtime health is not in the sidebar.It is a workspace utility, so the sidebar does not list it.Open Team & settings in the sidebar, then Runtime health under Workspace utilities, or open /team/runtime-health directly.
A card reads unattributed install.The event reached the store with no source machine hash.Expected for a row pushed before the stamp. It keeps its own bucket and is never merged into another machine's.
The page says the window filled its 5,000-row budget.The team pushed more proposal and outcome events in 30 days than the read is allowed to fetch.Read the two counts as a lower bound. The newest findings are the ones shown.
No fix applied to this finding yet, but you approved it.The deferred step had not fired when the page was read; it runs on the approval queue tick, every five minutes.Wait a tick and reload. If it never arrives, check the Actions queue tab for the held row.
Fix applied; before/after could not be measured.The proposal carried no evidence snapshot, or no fresh sample could be read at apply time.Nothing to fix on the page: it is the honest result. Confirm the install is still sampling on its Brain health tab.
No endpoint latency finding ever appears.Nothing in the shipped code writes the endpoint-timing samples that signal is derived from; the source is defined and unwired.Expect findings from the two memory signals only. The doc pages that list endpoint latency regression describe the watcher's intent, not a signal that can fire today.
The hosted runtime (repoops.ai) bucket never appears.CRON_SECRET is unset so the six-hourly route answers 401, or the cron fleet is healthy and wrote nothing.Set CRON_SECRET on the deployment. A healthy fleet writing nothing is the intended zero-noise behaviour.
Disable
Untick Enable hosted cloud sync, or press Disconnect under Connect to team, on each install that should stop pushing. Unset REPOOPS_SELF_AUDIT_SCHEDULE to stop the scheduled pass. The five-minute metric sample has no off switch short of stopping the app.
Roll back
Not provided. The page is read-only and the verb applies no fix, so there is nothing to undo: its rollback returns recommendation-only verb; nothing was applied, so nothing to roll back (lib/actions/verbs/runtime-self-audit.mjs). A remediation you applied by hand is reverted by hand.
Revoke access
Disconnect in the desktop app forgets this device's binding and revokes the device token on the server; the confirmation says the team's activity history is retained. Narrowing a member's repository access in the team admin removes the whole page's content for them, because these findings carry no repository scope.
Delete
Not provided. No route deletes one finding. Hosted runtime rows in the events table are hard-deleted after REPOOPS_NEON_RETENTION_DAYS (default 30) by the daily events-retention cron; on the install the row sits in .claude/brain/events/ under the repository brain.

Maintenance evidence

Feature id
runtime-health (spine leaf runtime-health)
Owner
Terminal-telemetry self-audit program (P0 to P3, plus the hosted producer added by the 2026-08-23 audit as G098). Guide: LDG-0717.
Supported product version
RepoOps v0.3.1
Last verified
2026-09-15, read against origin/main at 52366bb6d; labels read from the served tab source (website/app/team/(home)/runtime-health/page.tsx, public/settings.html, public/brain-health.html, public/actions.html) and the thresholds, caps and env flags from the modules cited in this file's header.
Example fixtures
No fixture file. The behaviour is exercised inline by website/app/team/(home)/runtime-health/page.test.tsx (the rendered card, including the before and after table), website/lib/self-audit-fleet-view.test.ts (grouping, the unattributed bucket, the outcome parse), website/lib/hosted-self-audit.test.ts (the hosted producer and its dedupe), lib/rules/runtime.self-audit.test.mjs (the deterministic finding and the AI fallback), lib/actions/verbs/runtime-self-audit.test.mjs (the measurement) and lib/perception/runtime-health-watcher.test.mjs (the derived signals and the omitted endpoint row).
Source references
website/app/team/(home)/runtime-health/page.tsx, website/lib/self-audit-fleet-view.ts, website/lib/hosted-self-audit.ts, website/app/api/cron/self-audit/route.ts, website/app/api/cron/events-retention/route.ts, website/lib/dashboard-nav.ts, lib/perception/runtime-health-watcher.mjs, lib/rules/runtime.self-audit.mjs, lib/actions/verbs/runtime-self-audit.mjs, lib/actions/approval-queue.mjs, lib/sync/cloud-pusher.mjs, public/settings.html, public/brain-health.html
Documentation review
Independent review requested on the slice pull request; not yet recorded.
Video review
Story script written 2026-09-15; render and review pending in the same slice.

Last updated