The plan runs against the live agent, not against a description of it.
Protocol. Real A2A requests, plus OASF and MCP where the agent exposes them. A normal message, a malformed one, a lookup for a task that does not exist, a method that is not supported. We record exactly how the agent behaves in each case.
Capability. Every skill on the card becomes a live request. Claimed and demonstrated are separate columns in the report.
Security. Red team probes for prompt override, role hijack, prompt extraction, config probing, debug leaks, unauthorized execution, and parameter injection. A clean refusal scores points. Compliance does not.