Agent inventories go stale in weeks
Someone adds a tool on Monday and the picture you signed off in March is no longer true. Continuous assurance keeps the inventory current, runs your attack suite on every release, and tells you when an agent gains reach it should not have.
Baseline, pipeline, runtime
Each layer catches what the one before it cannot.
Baseline
The assessment ends with a signed-off picture of every agent, tool, scope and data source. That baseline is what drift gets measured against.
Pipeline
The injection corpus, permission probes and egress tests from your assessment run on every release, so a regression is caught before the agent reaches users.
Runtime
In production we watch for what a test suite cannot predict: new tools, widened scopes, first-seen egress destinations and unusual tool sequences.
Agent assurance · this week
illustrativeAgents
14
Tools
63
MCP servers
5
Suite
214 tests
New tool reachable without an approval gate
billing-copilot · finance.refund · added 2h ago
Scope widened on a production agent
support-agent · mail.send:self → mail.send:all
First-seen egress destination
research-agent · outbound fetch outside allowlist
Injection regression suite passed on release
214/214 · build 2026.09.3 · 6m 12s
What it watches
Built around the failure modes we keep finding in assessments.
Live agent inventory
Every agent, tool, MCP server, model endpoint and data source, kept current instead of rediscovered once a year.
Drift detection
Alerts when an agent gains reach it did not have last week: a new tool, a widened OAuth scope, a new untrusted content source.
Regression testing in CI
The attack suite built during your assessment runs as a pipeline gate, with a clear pass or fail for each attack class.
Runtime anomalies
Unusual tool chains, outbound calls to first-seen destinations and spend spikes, each surfaced with the trace that caused it.
Policy conformance
Continuous checks that approval gates, allowlists, redaction and permission budgets are still in place and still enforced.
Evidence for audits
Test history and control evidence you can export for customer security reviews, auditors and internal risk committees.
Close to your agents, light on your data
Runs where your agents run
Deployed into your cloud, or connected read-only, depending on how much you want leaving your perimeter. Scope is agreed in writing before anything connects.
Minimum necessary data
We work from configuration, metadata and traces you choose to share. Your data is not used to train models, and retention is set by you.
People behind the platform
Monitoring is backed by the same team that ran your assessment, with a quarterly red-team refresh as new attack classes appear.
Works with
Assess first, then keep it honest
Continuous assurance is built on the baseline and attack suite your assessment produces. Start with a scoping call and we will tell you which agents need it.