The AI agent security assessment
Four to six weeks from the first workshop to verified fixes. We map what your agents can reach, attack them the way a motivated adversary would, implement the remediation with your engineers, then re-run the same attacks to prove they fail.
Typical duration
4-6 weeks
Scope unit
Per agent or agent fleet
Environment
Staging replica, or production under agreed rules
Output
Findings, fixes, verified re-test
Four domains, tested against your agents
The same four questions structure the whole engagement, from the inventory workshop to the final re-test.
What is the agent allowed to touch?
Permissions and identity
Agents inherit credentials from whoever wired them up, and that is usually a service account with far more reach than the task needs. We map the real blast radius of every agent identity.
- What we test
- Service accounts, OAuth scopes and API keys held by each agent
- Standing access versus just-in-time, per-task credentials
- Permission inheritance from the invoking user and from sub-agents
- Secrets living in prompts, config files and tool definitions
- What a single compromised agent could reach on its worst day
How the six weeks run
Nothing in this sequence waits on a tool finishing a scan. Each phase ends with something you can act on.
- 01Days 1 to 3
Scope and inventory
We build the map nobody has yet: every agent, the tools it holds, the data it reaches and the identity it runs as.
- Workshop with the teams who built and operate the agents
- Architecture, prompt, tool and credential review
- A threat model tied to your business rather than a generic checklist
- 02Weeks 1 to 3
Adversarial assessment
We attack the agent end to end across the four domains: permissions, tool use, prompt injection and data egress.
- Hands-on red teaming against a staging replica, or production under agreed rules
- A tailored injection corpus built from your own untrusted inputs
- Every finding reproducible, with the exact transcript that produced it
- 03Week 3
Findings and risk ranking
Findings ranked by what an attacker actually gains rather than by raw severity labels, with a briefing for executives and an appendix for engineers.
- Severity tied to real blast radius and reachability
- Mapped to the OWASP Top 10 for LLM Applications, MITRE ATLAS and NIST AI RMF
- A remediation plan sequenced by effort against risk removed
- 04Weeks 3 to 6
Remediation
We implement the fixes alongside your engineers: scope reduction, tool guardrails, approval gates, injection defenses and egress control.
- Pull requests and configuration changes against your own repositories
- Least-privilege identities and just-in-time credentials for agents
- Human-in-the-loop gates placed where they buy the most safety
- 05Week 6
Verify and hand over
We re-run the same attack suite against the fixed system and show, finding by finding, what no longer works.
- A re-test report you can hand to a customer, an auditor or your board
- The attack suite handed over as a regression pack you keep
- A short enablement session for the engineering team
- 06Ongoing
Continuous assurance
Agents change every week. The suite runs on every release, and inventory drift such as new tools, new scopes or new MCP servers gets flagged.
- Injection and permission regression tests in your CI pipeline
- Drift alerts when an agent gains reach it did not have last week
- A quarterly red-team refresh as new attack classes appear
How we close the findings
Remediation happens alongside your engineers, in your repositories, through your review process. Most engagements involve some mix of the work below.
Least privilege & identity
- A distinct, scoped identity per agent instead of a shared service account
- Just-in-time, task-scoped credentials in place of standing access
- Secrets moved out of prompts, tool definitions and config
- Permission budgets that fail closed when an agent reaches past them
Tool & action guardrails
- Tool allowlists with enforced argument schemas and typed validation
- Human approval gates on irreversible, financial and outbound actions
- Dry-run and idempotency support for high-impact tools
- Depth, recursion and spend limits on chaining and sub-agents
Prompt-injection defenses
- Hard trust boundaries between instructions and retrieved content
- Provenance tagging and spotlighting for everything untrusted
- Policy enforced at the tool-call layer, not only in the system prompt
- Untrusted content handled by isolated agents that hold no privileged tools
Data & egress control
- Retrieval that enforces the source system's access control per user
- Tenant and session isolation across context, cache and memory
- Redaction and secret detection on outputs, logs and traces
- Domain allowlists on outbound calls, links and rendered content
Some fixes are product decisions rather than code changes: removing a capability, adding a confirmation step, narrowing who can invoke an agent. We bring you the trade-off with the evidence, and you decide.
What you get
Everything is reproducible, and everything is yours to keep.
Agent and tool inventory map
Every agent, tool, MCP server, data source and identity in one diagram. For most teams it is the first time anyone has seen the whole thing.
Ranked findings with reproductions
Each finding carries the transcript, the payload and the exact conditions that produce it. Nothing we cannot reproduce goes in the report.
Executive summary
A plain-English read of what an attacker could do today, written for the people who sign off on shipping the agent.
Implemented fixes
Pull requests, policies and configuration delivered straight into your repositories, so your team is not left with a to-do list.
Verification report
The same attacks, re-run after remediation, with a clear before and after for every finding.
Regression attack suite
The tests we built for you, handed over to run in CI so the fixes stay fixed.
Agent stacks we test
Start with a scoping call
Thirty minutes on your agent architecture. You leave with our read on where the exposure sits and whether an assessment is worth the spend. No obligation either way.