Continuous assurance

Agent inventories go stale in weeks

Someone adds a tool on Monday and the picture you signed off in March is no longer true. Continuous assurance keeps the inventory current, runs your attack suite on every release, and tells you when an agent gains reach it should not have.

Three layers

Baseline, pipeline, runtime

Each layer catches what the one before it cannot.

01

Baseline

The assessment ends with a signed-off picture of every agent, tool, scope and data source. That baseline is what drift gets measured against.

02

Pipeline

The injection corpus, permission probes and egress tests from your assessment run on every release, so a regression is caught before the agent reaches users.

03

Runtime

In production we watch for what a test suite cannot predict: new tools, widened scopes, first-seen egress destinations and unusual tool sequences.

Agent assurance · this week

illustrative

Agents

14

Tools

63

MCP servers

5

Suite

214 tests

New tool reachable without an approval gate

billing-copilot · finance.refund · added 2h ago

Scope widened on a production agent

support-agent · mail.send:self → mail.send:all

First-seen egress destination

research-agent · outbound fetch outside allowlist

Injection regression suite passed on release

214/214 · build 2026.09.3 · 6m 12s

Capabilities

What it watches

Built around the failure modes we keep finding in assessments.

Live agent inventory

Every agent, tool, MCP server, model endpoint and data source, kept current instead of rediscovered once a year.

Drift detection

Alerts when an agent gains reach it did not have last week: a new tool, a widened OAuth scope, a new untrusted content source.

Regression testing in CI

The attack suite built during your assessment runs as a pipeline gate, with a clear pass or fail for each attack class.

Runtime anomalies

Unusual tool chains, outbound calls to first-seen destinations and spend spikes, each surfaced with the trace that caused it.

Policy conformance

Continuous checks that approval gates, allowlists, redaction and permission budgets are still in place and still enforced.

Evidence for audits

Test history and control evidence you can export for customer security reviews, auditors and internal risk committees.

How it operates

Close to your agents, light on your data

Runs where your agents run

Deployed into your cloud, or connected read-only, depending on how much you want leaving your perimeter. Scope is agreed in writing before anything connects.

Minimum necessary data

We work from configuration, metadata and traces you choose to share. Your data is not used to train models, and retention is set by you.

People behind the platform

Monitoring is backed by the same team that ran your assessment, with a quarterly red-team refresh as new attack classes appear.

Works with

OpenAI Agents SDKAnthropic Claude and MCPLangChain / LangGraphCrewAIMicrosoft Copilot StudioAmazon Bedrock AgentsGoogle Vertex AI AgentsIn-house orchestration

Assess first, then keep it honest

Continuous assurance is built on the baseline and attack suite your assessment produces. Start with a scoping call and we will tell you which agents need it.