Docs
Tell a live daemon from a frozen heartbeat
Use the cloud receipt time to decide whether a device is alive, then use its capture and publish timestamps to locate the stalled stage. A daemon that died cannot refresh the last green status it sent.
For: the team member checking whether each bound machine is reporting, capturing, and publishing current data
What it does, and why it helps
Daemon status is the hosted rollup of one health row per bound device. It derives running, stale, offline, or stopped from the server-stamped Last report time plus the device report. It then shows Telemetry push, Brain publish, repository count, the latest reported error, quiet reason, and version drift as separate evidence.
The page is read-only. It does not start, stop, restart, update, remove, or rebind a daemon. Run those operations on the machine through the desktop tray, the Settings toggle, or the local daemon and service commands.
The pain. A dead process cannot send a final failure state. If a dashboard trusts only its last running flag, the row can stay green after the machine or process has disappeared.
The point of view. Treat server receipt time as the liveness clock. Treat capture, telemetry push, brain publish, errors, quiet reasons, and version as separate clues about what happened after the process checked in.
What gets easier. Comparing every bound device in one table, spotting silence, distinguishing a deliberate stop from an offline process, finding a blocked push, and seeing a build behind the published feed.
When it helps. Use it when hosted data looks old, a machine was restarted, a token or plan changed, a team disabled cloud sync, or one device appears to feed fewer repositories than expected.
Its limits. A fresh health report proves that the status endpoint received metadata. It does not prove a recent telemetry or brain publish. A device that never reported has no row and triggers no dark alert. Feed lookup failure hides version drift rather than guessing.
Understand it in 30 seconds
Read the narration
- 0:00 A dead daemon cannot update the green running flag it reported before failing.
- 0:06 Judge liveness from cloud receipt time, then inspect each pipeline timestamp and quiet reason separately.
- 0:13 After five and a half minutes of silence, a device reads stale.
- 0:17 Telemetry push and brain publish get separate columns.
- 0:23 Restart the right device; do not call an old heartbeat healthy.
Synthetic example. Read the guide
Where to find it
- Hosted:
repoops.ai/team/daemon-status. It has no sidebar row: open it from the search box, or go straight to its address. - Desktop: hosted only.
When to use it
Separate a dead process from an old push
Situation. Hosted data is old, but the last device report and the last telemetry push do not have the same age.
What you do. Read Status and Last report first. Then compare Telemetry push and Brain publish for that device.
What you see. A report older than five and a half minutes reads stale. A still-running row silent past the dark cutoff reads offline. A recent report can remain running while its publish timestamp is old.
What it establishes. You can decide whether to restore the process, inspect a push gate, or wait for the hourly brain cycle without treating those paths as one health signal.
Explain why a device is quiet
Situation. A row is amber or stopped and the machine has sent a bounded reason with its status.
What you do. Read the line below the status chip before changing the service. Check for cloud sync off, no binding, an expired or revoked token, a plan without cloud push, or a network error.
What you see. The page maps the daemon reason to plain text. Older daemons that do not send reason fields show only the state chip.
What it establishes. The next action follows the reported gate. An intentional opt-out does not get diagnosed as a process crash.
Find a version lag without inventing one
Situation. One device reports a package version older than the version in the published update feed.
What you do. Check Version and devices behind feed. Update the desktop build on that machine through the supported release path.
What you see. The row shows behind plus the latest version only when the feed resolves and semantic comparison says the device is older. An unresolved feed leaves the count as a dash.
What it establishes. You can identify a reported version gap. The hosted page does not download or swap the daemon build.
Respond to a dark-device alert
Situation. A still-running row has stopped reporting for the configured dark cutoff and the hourly sweep has notified owners and admins.
What you do. Open Daemon status, identify the device by team label and shortened ID, then inspect the local service and status on that machine.
What you see. The sweep records one daemon.dark audit event and one alert per dark episode. A later heartbeat re-arms a future episode. A graceful stop does not count as dark.
What it establishes. The alert points to the machine whose cloud heartbeat stopped. It does not establish whether power, network, service, or credentials caused the silence.
Before you start
- Supported versions
- RepoOps v0.3.1. Server-timed liveness, quiet reasons, version drift, local status reconciliation, service installation, and dark-device alerts are present in this release.
- Where it runs
- Hosted: AI Security, then Check health on the Runtime collection row (repoops.ai/team/daemon-status). Desktop: no sidebar row. Local status is available through the tray and npm run daemon:status.
- Permissions
- A signed-in team member can view the team page. The health ingest authenticates a device token and resolves both team and device from that token. Local service operations need access to the machine and its user service manager.
- Connections
- The device must be bound and effective cloud sync must be on before health metadata is posted. The hosted page and dark alert also need the RepoOps website and its database. Local status works without the hosted page.
- Plan
- The code gates health push through binding and cloud sync. A plan refusal is reported as a quiet reason when the server returns it. No separate view-only plan gate was found on the team page.
Configure it
- Bind the device and opt into cloud sync.
Health push is a no-op without a device token or when effective cloud sync is off. The team and device identity come from the token, never the posted body.
- Keep capture and cloud sync as separate choices.
Run the background capture daemon controls scanning and capture. Enable hosted cloud sync controls whether already captured redacted events and health metadata can reach the team.
- Install the service for unattended use.
npm run daemon:install installs and starts a per-user launchd agent, systemd user unit, or Windows Scheduled Task, with a per-user Run-key fallback on Windows when task creation is denied.
- Verify on the machine before trusting hosted state.
npm run daemon:status reads the local status file and reconciles its process ID. A hard-killed process can leave running true in the file until this local check corrects the output.
- Read each timestamp for its own job.
Last report is cloud liveness. Telemetry push advances only when events left the machine. Brain publish advances only after a successful brain publish.
| Setting | Where | A sensible choice | Why it matters |
|---|---|---|---|
Run the background capture daemon | Desktop Settings, Capture daemon | on by default | Turning it off makes the running process inert: no transcript scan and no capture push. Re-enabling takes effect on the next cycle but cannot start a process that already stopped. |
Enable hosted cloud sync | Desktop Settings, Cloud sync (hosted) | off by default unless a team default applies | Health push requires effective cloud sync and a binding. This is separate from whether the daemon captures locally. |
REPOOPS_DAEMON_DISABLED | the data directory environment | unset; set to 1 for a host-level stop | It overrides the Settings toggle and pins the daemon inert or makes start exit. |
REPOOPS_DAEMON_CAPTURE_INTERVAL_MS | the data directory environment | 300000 by default | The daemon posts health after each capture cycle, so this cadence also bounds the normal hosted heartbeat interval. |
REPOOPS_DAEMON_BRAIN_INTERVAL_MS | the data directory environment | 3600000 by default | Brain publish has its own hourly cadence and timestamp. It is not the liveness heartbeat. |
REPOOPS_DAEMON_STALE_ALERT_HOURS | hosted runtime environment | 2 by default | The page, fleet view, and hourly dark-device sweep use this cutoff for offline state and alerting. |
REPOOPS_HEALTH_PUSH_TIMEOUT_MS | the daemon environment | 12000 by default | It bounds one status POST so a half-open connection does not hold the daemon cycle indefinitely. |
REPOOPS_HEALTH_PUSH_MAX_RETRIES | the daemon environment | 3 by default | Transient status failures retry within the push path. Hard client errors do not retry. |
REPOOPS_UPDATE_FEED_URL | the daemon and hosted feed resolver environment | https://repoops.ai/updates/latest.yml by default | Version comparison depends on the published feed. Failure resolves to no drift claim. |
REPOOPS_UPDATE_CHECK_INTERVAL_MS | the daemon environment | 86400000 by default | The local drift check runs at start and then on this cadence. It detects and surfaces an update; it does not apply one from this page. |
What you should see
Fresh process and recent push
Configuration. A bound device is reporting inside five and a half minutes, reports stale false, and has pushed events recently.
Expect. The banner reads Daemon running, the row reads running, and Last report and Telemetry push show recent ages.
Verify. Compare the hosted Last report with npm run daemon:status on that machine. Use Telemetry push, not the green chip alone, to establish recent event egress.
Fresh process, stale capture path
Configuration. The device reports recently but its own stale flag is true, or no device has a recent events push.
Expect. The page reads Daemon capture may be stale. The reason line may identify an opt-out, binding, token, plan, or network cause.
Verify. Check Last report separately from Telemetry push. A recent report with an old push is not the same as an offline process.
Silent process
Configuration. The last server receipt is older than five and a half minutes but younger than the dark cutoff, or older than the cutoff.
Expect. The row first reads stale with its age, then offline after the cutoff. The aggregate banner is amber when no fresh device remains.
Verify. Inspect the local service, process ID, and status file. The cloud page cannot name the physical cause of lost heartbeats.
Graceful stop
Configuration. The daemon completed its stop path and posted running false.
Expect. The row reads stopped even as the report ages. The dark sweep does not alert a gracefully stopped row.
Verify. Check the Cloud audit timeline for daemon.stopped and confirm the local lock is released. Do not relabel an intentional stop as offline.
No row yet
Configuration. The team has no accepted health report from any device.
Expect. The page reads No daemon has connected yet and offers the local start command or service install path.
Verify. Check binding, effective cloud sync, service state, and the local log. A never-reported device does not enter the dark sweep.
Data and cost
- What is captured
- One hosted row per team and device stores booleans, bounded reason strings, version, start and activity timestamps, repository count, up to 200 canonical repository keys, a last error capped at 500 characters, server receipt time, and the dark-alert stamp. The local status file holds a broader process and cycle snapshot.
- Who can see it
- Signed-in team members can view rows for their team. The hosted page displays the team label, the first eight device-ID characters, canonical repository keys, counts, timestamps, state, reason, error count, and version. The ingest drops unknown fields and never accepts token or local path fields.
- How long it is kept
- The hosted daemon_health table upserts one current row per team and device. A team deletion cascades its rows. No device-row expiry, manual delete route, or retention duration is provided. State-change and dark events also enter the Cloud audit store under that store's policy.
- What leaves the machine
- When bound and cloud sync is on, the daemon sends status metadata and credential-stripped canonical repository keys to the authenticated status endpoint. It does not send prompt or response text through this endpoint. The device token is the bearer credential and is not part of the body.
- What it costs
- State derivation, status ingest, version comparison, local status, and the dark sweep make no model call. Network and hosted storage usage remain. Dark alerts reuse configured email and budget-alert chat channels.
When the result differs
| Symptom | Likely cause | Next action |
|---|---|---|
| The page says No daemon has connected yet. | The daemon never posted, the device is unbound, effective cloud sync is off, or the service is not running. | On that machine run npm run daemon:status, confirm the binding, turn on Enable hosted cloud sync, and start or install the daemon. |
| A device says running but Telemetry push is old. | The process is alive while no events left the machine, or capture is blocked or idle. | Read the quiet reason and local lastError, then check cloud sync, token, plan, and whether there were new events. Do not diagnose process death from the push timestamp alone. |
| The row is stale after the machine was killed. | The last report is beyond the five-and-a-half-minute freshness window but has not yet reached the offline cutoff. | Restart the local service or inspect its log. The row becomes offline at the configured dark cutoff if reports do not resume. |
| The row says offline but no alert arrived. | The hourly sweep may not have run, no email or chat channel is available, the row already alerted for this episode, or the row never existed. | Check the cron health route, team owner and admin email delivery, the armed chat webhook, and the Cloud audit timeline for daemon.dark. |
| devices behind feed shows a dash. | The hosted update feed could not be resolved, so the page refuses to infer drift. | Check feed availability and REPOOPS_UPDATE_FEED_URL. Use the reported version as raw evidence until comparison returns. |
| Turning the Settings toggle on did not restart capture. | The toggle is re-read by a running daemon, but it cannot start a process that has already stopped. | Start it from the desktop tray, run the daemon locally, or install the OS service. |
- Disable
- Turn off Run the background capture daemon, use REPOOPS_DAEMON_DISABLED=1 for a host override, stop the process from the tray or local CLI, or uninstall the service. Turning off cloud sync stops future health egress but is not the same as stopping capture.
- Roll back
- Not provided on the hosted page. Version drift is detection only there. Restore a prior desktop or daemon build through the machine's release and service workflow if your deployment policy permits it.
- Revoke access
- Revoke or disconnect the device binding to stop that token from authenticating future status posts. The daemon marks an expired or revoked token as unbound. The hosted status page has no revoke control.
- Delete
- Not provided. There is no hosted control or route that deletes one daemon_health device row or its audit history. The current row persists until overwritten or its team is deleted.
Related tasks
Maintenance evidence
- Feature id
daemon-status(spine leafdaemon-status)- Owner
- Cloud-hosted localapp Phase A, daemon observability and daemon reliability hardening. Guide: LDG-0717.
- Supported product version
- RepoOps v0.3.1
- Last verified
- 2026-09-15, read against origin/main at 52366bb6d; labels read from the served tab source; hosted state derivation, ingest validation, local status reconciliation, service controls, health push gates, update comparison, alert schedule, and stored row shape checked in source and tests.
- Example fixtures
- lib/daemon/runner.test.mjs covers status derivation, stale state, disabled state, cycles, heartbeat, process health, and lifecycle writes; lib/daemon/health-push.test.mjs covers payload shape, binding and sync gates, canonical keys, retry, and token failure; website/lib/activity-format.test.ts covers server-timed liveness and quiet reasons; website/app/api/daemon/status/route.test.ts covers device auth, tenancy, size cap, validation, and ingest; website/app/team/(home)/daemon-status/page.test.tsx covers empty, running, stale, offline, stopped, version drift, and feed failure; website/app/api/cron/daemon-staleness/route.test.ts covers auth, dark detection, deduplication, recovery, and graceful stops; scripts/daemon-service.test.mjs and scripts/service-install/generators.test.mjs cover service installation paths.
- Source references
lib/canonical-spine.json,lib/canonical-tabs.mjs,public/settings.html,lib/account-settings.mjs,lib/daemon/runner.mjs,lib/daemon/health-push.mjs,scripts/daemon.mjs,scripts/daemon-service.mjs,website/app/team/(home)/daemon-status/page.tsx,website/lib/activity-format.ts,website/lib/daemon-health.ts,website/app/api/daemon/status/route.ts,website/app/api/cron/daemon-staleness/route.ts,website/db/schema.ts,website/vercel.json,.env.example- Documentation review
- Independent review requested on the slice pull request; not yet recorded.
- Video review
- Narrated story rendered and published 2026-09-26 (render 65417a6803d8, LDG-1012) with the breadcrumb AI Security, checked against main at b0bb02812. Six frames, the captions and the transcript were reviewed by the authoring agent, not an independent reviewer; the audio was not listened to by a person. Narration is the provisional Windows voice until LDG-0721.
Last updated