Guard · MCP security

MCP security

Five defenses over the MCP servers you have already configured. No server is ever executed to check it. Drift alerts catch a server definition changing under you. An allowlist pins the exact configuration you approved. Canaries plant a non-functional honeytoken and fire if it turns up in a configured MCP server's command, arguments, or URL. The red-team runner attacks your own setup with a tool-poisoning library. Skill supply chain signs and verifies what each skill and subagent is allowed to do.

Start here if you want to see what is exposed in a repo and whether it is contained: Security.

Where to find it

  • Localhost: /mcp-security.html
  • API: GET /api/mcp-security/alerts (drift alerts plus the allowlist verdict for every configured server), POST /api/mcp-security/alerts/dismiss, POST /api/mcp-security/allow, GET /api/mcp-security/canary, POST /api/mcp-security/redteam, GET /api/skills/manifest
  • CLI: repoops skills sign and repoops skills verify
  • Navigation: AI Security in the sidebar, then MCP security under All tools, in the MCP and tool inventory group. The page answers at its URL either way.

What it does for you

A server that changes after you approved it does not stay approved.This is the rug-pull vector: a tool you vetted is edited later. The allowlist pins the source tool, the server name and a hash of its configuration together. Change any of them and it un-pins, which forces a re-approve. An unpinned server is flagged by default and marked blocked only when REPOOPS_MCP_ALLOWLIST_BLOCK=1 is set. Drift becomes a standing watch you dismiss deliberately, not a line in a log.
A honeytoken that fires on a literal match, not a guess.A canary is a credential-shaped string that authenticates to nothing. RepoOps matches it against the MCP config surface it can reach: the command, arguments and URL of every configured server. A hit is a literal token match, so there is nothing to interpret, and it means something rewrote that config to carry your credentials. The scan stops there. A canary sitting in a .env or a notes file is never read, so planting it there proves nothing.
Attack your own setup before someone else does.The red-team runner runs a tool-poisoning attack library against your configuration and returns a per-attack verdict: fails-closed when the shipped detector caught the probe, fails-open when it did not.
Skills and subagents carry a signed capability manifest.Each one declares its tools, its network scope and its filesystem scope. When the tool set changes you get a drift alert, and gaining a network or filesystem-write capability (or widening to an unrestricted inherit) escalates it to high. Unsigned, drifted, unverifiable, and deleted-but-still-signed all fail closed.

Built vs. planned

All five sections ship today. Each manifest carries an Ed25519 signature under the per-install attestation key beside the HMAC ledger signature; a record holding only the HMAC signature verifies as weaker, not as plain verified. Either way the check catches an out-of-band change by something that does not hold the key, and a local key holder can still re-sign a file they edited. The surface says so rather than implying more. The inventory records environment variable names so a server's secrets can be named; values are never read. LLM-free throughout.

Read more

Last updated