What the terms mean
- Attempt
- One adversarial test case sent to the SUT (possibly several conversational turns), then scored by the Judge. It is the denominator: 30 attempts / 0 findings means the target defended all 30.
- Finding
- An attempt the Judge ruled a genuine weakness (a fail or partial), recorded as a vulnerability. Findings are a subset of attempts.
- SUT / "System Under Test"
- The deployed Clinical Co-Pilot we attack. Each run is stamped with the SUT version it hit;
livemeans the currently-deployed target, not pinned to a specific build. - Defended %
- Share of a version's attempts the SUT passed (higher is more resilient).
Resilience over SUT versions (defended %)
loading…
SUT-version timeline — how results changed as the target was improved
loading…
Findings by severity / status
loading…
Spend & published reports
loading…
loading…
Same test, over time: pass/fail/partial trend per category · subcategory
loading…
Run history
loading…
Regression outcomes (findings re-tested on each version)
loading…
Published reports
loading…