Getting started · Connect production evidence

Trace a production signal to its repository

A collector beside your running service sends typed metrics, logs, traces and events to RepoOps over a short-lived credential it exchanges for each delivery window. The hosted Production evidence page shows the latest observations for the repositories you can see, and marks the release, commit and session the collector reported as reported context rather than verified fact.

For: the engineer investigating a production signal, and the team owner or admin who approves the workload that sends it

What it does, and why it helps

This is the one evidence source with no provider behind it. Nothing is pulled: your collector writes version-1 JSON envelopes and pushes them. An owner or admin approves one exact GitHub Actions identity per repository, by audience, subject and optionally a reusable workflow reference. The workload then exchanges a fresh ID token at POST /api/vm-credentials/issue for a token scoped to that one repository and the events-ingest endpoint. Installing the GitHub App does not authorize a workload, wildcards are refused, and two policies that both match deny rather than pick one.

On the collector host, enqueue redacts an envelope and writes it to a persistent spool, and drain pushes it to the gateway. A file is deleted only after the receiver reports exactly one accepted or deduplicated record, so a network failure costs a retry and not the evidence. The receiver keeps the source's own name, event id and timestamp, drops every field outside the allowlist for that kind, redacts allowed text, and derives the row id from the team, repository, source, source event id and kind, so a replay lands once. A production, performance, security or cost event can also open an incident case; a deployment event is an anchor, not a problem. Whatever the collector said about a session, the card reads Unverified.

The pain. A metric moves in production and the context sits in three places: the release in one tool, the agent run in another, and the person who would know is offline. The desktop app only publishes evidence from a machine that is awake, so the window you most want to read is the one nobody captured.

The point of view. A collector's word about a commit or a session is a place to start looking, not a finding. Keep the source's own identifiers and timestamps, label everything else reported, and investigate the change before judging the agent that wrote it.

What gets easier. Getting the evidence at all. The sender runs on the collector host with its own persistent spool, needs no desktop process and no repository checkout, and holds a credential that expires in minutes rather than a long-lived key.

When it helps. A paid hosted team with the repository connected through the GitHub App, a collector that can write the version-1 JSON envelope, and a host that can present a GitHub Actions OIDC ID token.

Its limits. It is a JSON mapping interface, not an OTLP receiver: it accepts newline-delimited JSON, not native protobuf. It does not verify a reported session and it does not prove that a prompt caused an incident. A host with a different identity provider needs explicit operator configuration and cannot reuse a GitHub identity it does not hold. Nothing on this page decides recurrence, because it reads no captured sessions.

Understand it in 30 seconds

30.1 s, captions on. Narration: Microsoft Zira Desktop (provisional voice; an approved narration source is pending).Transcript
Read the narration
  1. 0:00 A metric spikes in production.
  2. 0:02 Nobody can say which change did it.
  3. 0:06 Collect beside the service, not on a laptop.
  4. 0:09 One approved workload exchanges a ten-minute token.
  5. 0:13 Each card keeps the collector's own source, time and release.
  6. 0:17 The reported session reads unverified, because nothing has verified it.
  7. 0:23 Investigate the change before judging the agent.
  8. 0:26 The record says who reported what.

Synthetic example. Read the guide

Where to find it

  • Hosted: repoops.ai/team/production-evidence, from Attribution in the sidebar, then Production evidence under All tools, in the Evidence readers group.
  • Desktop: hosted only.

When to use it

A latency signal with a release beside it

Situation. Checkout latency moved and the collector for that service already sends metrics and traces. You want the source record before anyone opens a pull request.

What you do. Open Production evidence, type the repository into the Repository field, and press Filter evidence. Open Inspect source fields on the card that matches the window.

What you see. Each card carries Repository / service / environment, Source observation (the source name, its event id and the source timestamp), Reported release / commit, and Reported session followed by Unverified. A metric card also shows an Observed metric figure with its unit. The list says how many observations it is showing, up to 100, ordered by source timestamp.

What it establishes. You have the collector's own record of what happened and when, and the release it named. Inspect reported session and cost looks that identifier up within your access, under a line that says a matching run does not verify causation.

A workload that should stop sending

Situation. A repository moved, or the workflow that collects was replaced. Its approval is still active and its credentials still mint.

What you do. On the same page, in Approved production workloads, press Revoke workload access on that policy.

What you see. The status line reads that access is revoked and existing workload credentials are refused immediately. The policy row reads Revoked, and the change is written to the audit log as production.workload.revoke.

What it establishes. No new credential can be minted under that policy, and an already-minted one fails its next use, because every use rechecks the policy, the installation and the repository. Evidence already stored stays where it is.

Before you start

Supported versions
RepoOps v0.3.1, the release this guide was read against. The sender is part of the installed RepoOps runtime (bin/repoops-production.mjs and lib/), not a separate package, and runs on the collector host.
Where it runs
Hosted only. The desktop app has no tab for it. One local command reads the same batch on a developer machine: node bin/repoops-production.mjs observe <repo-root> <ndjson-file> writes one prod.error per envelope the shared detector qualifies, and nothing schedules it.
Permissions
Reading takes team membership and a paid seat, narrowed by the repositories you may see. Approving or revoking a workload takes the owner or admin role, a paid seat and a same-origin request, and both actions are written to the audit log as production.workload.approve and production.workload.revoke.
Connections
The GitHub App installed, the repository connected to the team, and the installation not suspended. On the collector side: a host that can present a GitHub Actions OIDC ID token, a persistent directory for the spool, and one drain worker for that directory. A temporary CI workspace is not durable storage.
Plan
A paid hosted plan (Solo Hosted or above) with subscription status active or trialing. Below that the page shows Access required and the ingest route answers 402.

Configure it

  1. Approve the exact workload.

    On /team/production-evidence, in Approved production workloads, pick the Connected repository, enter the Exact OIDC audience and the Exact OIDC subject, add the Exact reusable workflow reference (optional) when a reusable workflow does the sending, then press Approve this exact workload. The issuer is fixed at GitHub Actions and the repository is bound to its immutable GitHub id. A value with a wildcard, a line break or leading whitespace is refused.

  2. Exchange the workload's ID token for a credential.

    POST /api/vm-credentials/issue with the fresh ID token and the repository. The response carries vmToken, expiresAt, repoScope and endpointScope. The credential lasts ten minutes, covers one repository and the events-ingest endpoint, cannot be refreshed, and one ID token mints at most one credential. Acquire a fresh one for each delivery window.

  3. Give the sender a persistent spool.

    node bin/repoops-production.mjs enqueue /persistent/repoops-spool observations.ndjson reads one envelope per line, redacts it, and writes it to the spool directory, which it creates with owner-only permissions. Each file is written and then renamed, so a restart mid-write loses nothing. An envelope over 64 KiB is refused, and so is one whose source identity or repository context changed under redaction.

  4. Drain it on a schedule.

    node bin/repoops-production.mjs drain /persistent/repoops-spool sends the queue, oldest first, at most 100 files per run and one envelope per request. Supply the credential through REPOOPS_VM_INGEST_TOKEN in the worker environment; tokens never enter the spool. Schedule one worker per spool directory.

  5. Read the evidence, then follow it.

    Filter evidence narrows to one repository. Inspect source fields shows the typed data the receiver stored. Inspect reported session and cost opens Run forensics on the identifier the collector reported, within your own access.

  6. Revoke when the identity changes.

    Revoke workload access ends the approval. Removing the repository from the team or suspending the installation has the same effect, because both are rechecked on every credential use.

SettingWhereA sensible choiceWhy it matters
REPOOPS_VM_INGEST_TOKENthe drain worker's environmenta credential fetched fresh for each delivery windowThe drain sends it as the bearer. It is never written to the spool, so a queued file carries no credential.
REPOOPS_VM_INGEST_ENDPOINTthe drain worker's environmentunset, which sends to https://repoops.ai/api/events/ingestAnything that is not a public HTTPS URL is refused, including loopback, private and metadata addresses. There is no development exemption.
audienceProduction evidence, Approved production workloads, Exact OIDC audiencethe audience your workflow requests, for example repoops-production-your-teamChecked against the approved policy and again against the signed token. A mismatch denies.
subjectProduction evidence, Approved production workloads, Exact OIDC subjectthe workflow's exact subject claim, for example repo:owner/repo:environment:productionThis is what binds a credential to one identity rather than to a repository in general.
workflowRefProduction evidence, Approved production workloads, Exact reusable workflow reference (optional)set it when a reusable workflow sends, leave it empty otherwiseWhen set, the token's workflow reference must equal it, so another workflow in the same repository cannot send.
REPOOPS_VM_OIDC_ISSUER, REPOOPS_VM_OIDC_AUDIENCE, REPOOPS_VM_OIDC_TEAM_IDthe hosted deployment's environmentunset unless an operator provisions a non-GitHub identity providerAll three set enables the older operator path. With any of them unset and no approved policy, the issue route answers 404 and mints nothing.
REPOOPS_VM_CRED_TTL_SECthe hosted deployment's environment600 (the default)The operator path's credential lifetime, clamped to 900 seconds. A policy-minted credential always uses 600.
REPOOPS_NEON_RETENTION_DAYSthe hosted deployment's environmentunset, which is 30The daily sweep deletes stored production events older than this. Thirty is the default and the ceiling. An org retention window can only tighten it.
ⓘ
To stop or undo
Press Revoke workload access on the policy. No new credential mints under it and a live one fails its next use. Stopping the drain worker stops delivery but leaves the queued files on the collector host, where removing them is your filesystem operation, not something RepoOps does.

What you should see

A collector that is delivering

Configuration. One approved policy, a drain worker on a schedule, the repository connected to a paid team.

Expect. The drain reports ok with a count acknowledged and nothing pending. The page lists the newest observations, up to 100, ordered by the source timestamps the collector supplied.

Verify. The card's Source observation field shows the collector's own name and source event id. Re-running the same file adds no rows, because the id is derived from the team, repository, source, source event id and kind.

An envelope the gateway will not take

Configuration. An envelope missing a required field, a metric with no finite value, a log with a severity outside debug, info, warn, error and fatal, or a secret-shaped value in a source identity field.

Expect. Either enqueue refuses at the collector, or the receiver returns the record in its rejected list and the drain stops with the status and that list. The file stays queued.

Verify. Read the reason, correct the envelope at its source, and re-run. A retry alone does not make invalid data valid.

Nothing for this filter

Configuration. A repository with no delivered envelopes, or a filter that matches none.

Expect. Connect a production collector, with the line that no supported production envelopes are available for this repository filter, and a link to the collector setup.

Verify. That state means nothing arrived, not that nothing happened. Check the drain worker's own output before reading a quiet page as a quiet service.

Data and cost

What is captured
One row per observation in the hosted event store: the kind (prod.metric, prod.log, prod.trace or prod.event), the source timestamp, the repository scope, the reported commit and session, the source name and source event id, service, environment, release, and the typed data for that kind. Fields outside the allowlist for the kind are dropped before storage, allowed text is redacted, and the stored attribution says unverified with the basis that correlation is not causation. A prod.trace span can declare its sampling in data.sampling (policy none, head or tail, and a rate), data.traceState (the OpenTelemetry th threshold is read) or data.traceFlags; RepoOps stores only the derived policy and rate, and a span that declares nothing is stored as unknown. On a case, a trace that was sampled upstream, or whose spans name a parent span that is not stored, reads partial in the Runtime event payload row with the reason, never full. The case's Coverage matrix also shows the trace join: which stored traces and request ids reach the case, how (trace id, request id, or a session, commit or release the collector reported) and with what confidence, with a span or request that arrived twice counted once.
Who can see it
Team members with a paid seat, narrowed by repository access. Personal repositories stay outside team views. The authorization predicates run in SQL before the limit, and a request returns at most 100 rows.
How long it is kept
The daily events sweep deletes production event rows older than REPOOPS_NEON_RETENTION_DAYS, 30 days by default and at most. An org owner can record a tighter window, 1 to 3650 days, which a second daily sweep enforces on member teams; it can only tighten the global window, never relax it. There is no per-signal retention control.
What leaves the machine
Sending is the point, so everything queued leaves the collector host: gzipped newline-delimited JSON to the hosted gateway, over public HTTPS only. The collector redacts before the queue touches disk, over the secret, aws_key, jwt, email, oauth_token, slack_token, card and ssn classes, and refuses an envelope whose source identity or repository context changed under redaction rather than sending a broken join. Full session transcripts do not travel this path; they stay on the development evidence pipeline with its own controls.
What it costs
No model call is made on this path, so it adds no LLM cost. It draws on the team's ingest allowance: 60 requests a minute with a 1.5x burst, and 100,000 events a day, both per account, both overridable per account, and both answered with 429 when exceeded.

When the result differs

SymptomLikely causeNext action
The credential exchange answers 404.No workload policy is approved and the operator path is not configured, so the route is dormant.Approve a workload on /team/production-evidence, or have an operator set the three REPOOPS_VM_OIDC variables.
The exchange answers 401, workload identity is not authorized.The token's audience, subject, repository id, repository name or workflow reference does not match exactly one active policy, or more than one policy matches.Compare the approved audience and subject with what the workflow requested. An ambiguous match denies by design; revoke the duplicate.
The exchange answers 409.That ID token was already exchanged, or the approval stopped being valid between the check and the mint.Request a fresh ID token. One assertion mints at most one credential.
Ingest answers 402.The team is not on a paid plan, or the acting seat is not entitled.Check the plan and the seat count on the billing page, then retry the drain. The queued files are still there.
The drain reports not ok with a status and a rejected list.The receiver refused an envelope, or acknowledged a count other than one, which is not permission to discard the file.Read the reason, fix the envelope at its source, and re-run the drain.
The drain fails saying a public HTTPS ingestion endpoint is required.REPOOPS_VM_INGEST_ENDPOINT points at loopback, a private or metadata address, plain HTTP, or a URL carrying credentials.Point it at the hosted endpoint, or unset it to use the default.
The page says the evidence could not be loaded.The read failed. The page renders that rather than an empty success.Retry shortly. Your stored evidence has not been changed.
A card reads Unverified although you can find the run.The label is a statement about provenance, not a lookup result.Follow Inspect reported session and cost if you want the run. A matching run does not verify causation.
Disable
Revoke every approved policy and stop the drain worker. With no active policy and no operator configuration, the credential route answers 404 again and nothing can be minted.
Roll back
Not provided. No route withdraws an accepted observation. Delivery is one way, and the stored row in the hosted event store is the record.
Revoke access
Revoke workload access on the policy row, which is immediate: the exchange and every later use of an already-minted credential recheck the policy, the installation and the repository ownership. Removing the repository from the team or suspending the installation revokes access the same way.
Delete
Not provided for one observation. There is no per-row delete route. Rows leave on the daily retention sweep (REPOOPS_NEON_RETENTION_DAYS, 30 days by default) or sooner under an org retention window, which permanently deletes a member team's rows older than the window.

Maintenance evidence

Feature id
production-evidence (spine leaf production-evidence)
Owner
Hosted incident pipeline program (IRP-02, the production-events gateway, the fifth observation source), with the workload-identity path from CNA.4b. Guide: LDG-0717.
Supported product version
RepoOps v0.3.1
Last verified
2026-09-15, read against origin/main at 52366bb6d; labels read from the served tab source, website/app/team/(home)/production-evidence/page.tsx and website/components/production-workload-settings.tsx; the sidebar location and the advanced-view toggle read from website/lib/dashboard-nav.ts and website/components/team-sidebar.tsx; the envelope fields read from website/lib/production-events.ts against docs/examples/production-evidence.ndjson.
Example fixtures
docs/examples/production-evidence.ndjson and docs/examples/production-evidence.schema.json, four synthetic signals for one service. The behaviour is exercised by website/lib/production-events.test.ts, website/app/api/team/production-evidence/route.test.ts, website/app/api/team/production-workloads/route.test.ts, website/app/team/(home)/production-evidence/page.test.tsx, website/lib/vm-credentials.test.ts, lib/sync/production-sender.test.mjs and lib/prod-adapters/production-events.test.mjs.
Source references
website/app/team/(home)/production-evidence/page.tsx, website/components/production-workload-settings.tsx, website/lib/production-evidence.ts, website/lib/production-events.ts, website/lib/production-workload-policy.ts, website/lib/vm-credentials.ts, website/app/api/vm-credentials/issue/route.ts, website/app/api/team/production-workloads/route.ts, website/app/api/events/ingest/route.ts, website/lib/ai-incident-cases/gateway-intake.ts, lib/sync/production-sender.mjs, bin/repoops-production.mjs, lib/incident-source-contract.mjs
Documentation review
Independent review requested on the slice pull request; not yet recorded.
Video review
Story script written 2026-09-15; render and review pending in the same slice.

Last updated