Method
AI risk assessment: method and template
The method works in seven steps from the top down. Your company applies it to its own systems; I provide the template and support the start. The approach follows the thinking in ISO 14971 (risk management for medical devices: harm, hazard, control), applied to AI.
- HarmWho is hurt, and how badly? Money, data, health, reputation, law.
- HazardIn which situation can this harm happen?
- Failure modeHow must the AI behave for that to occur? Wrong figure, invented detail, data leak, a command hidden in a document.
- LikelihoodThe upper failure bound comes from tests, not from a guess.
- ControlWhat catches the failure? A rule, a person, a lock, a stop switch.
- TestA test as code, acceptance criterion fixed before the run.
- EvidenceResult, version, data hash, date, responsible person.
Example: AI prepares quote drafts for a trades business
Made-up example, no customer data. The thresholds are assumptions for illustration; in a project I set them with you before the first test runs.
| Harm | Failure mode | Control | Test as code | Acceptance criterion |
|---|---|---|---|---|
| Customer receives a wrong price | Model copies or computes a wrong line item | Prices only from the price list, total computed in code, a person approves | 200 cases with a known total; mutation test: a changed price list must change the total | 0 deviations in 200 cases (upper bound 1.5%) |
| One customer's data lands in the draft for another | Leak through context or an instruction smuggled into a document | Access separated per customer, document content never treated as an instruction | Injection test: 4 languages, 3 patterns, 10 runs each; foreign-data scan | 0 hits in 120 runs (12 patterns × 10 repeats; computed upper bound 2.5%, only for these patterns) |
| Draft contains a legal commitment | Model words a binding promise | Text blocks for legal wording, pattern filter, a person | 100 drafts, two reviewers with a checklist | No commitment; reviewer agreement κ ≥ 0.8 |
| The human just waves it through | Automation bias | Spot checks with planted errors | Decoy drafts: 1 in 20 contains an error on purpose | At least 90% of decoys found |
| Quality drops after a model or data change | Drift, silent model change | Version pinned, regression set on every change | All tests above on every update; drift index (PSI) | All green, PSI below 0.25 |
Why testing alone is not enough
Runs without a failure only show an upper bound on the failure rate. For rare, severe harm you therefore need controls outside the model.
| Independent runs without a failure | Shown (95% confidence) |
|---|---|
| 30 | Failure rate below 9.5% |
| 60 | Failure rate below 4.9% |
| 100 | Failure rate below 3.0% |
| 200 | Failure rate below 1.5% |
| 300 | Failure rate below 0.99% |
| 1000 | Failure rate below 0.30% |
| 3000 | Failure rate below 0.10% |
To show “below 0.1%” you need 3000 independent runs without a failure, and the bound only holds for the cases tested. With zero failures the exact one-sided 95% bound 1 − 0.05^(1/n) applies; with failures the rig uses the Wilson interval. Where a single failure weighs heavily, safeguards are put around the model: permissions, locks, approvals.
Answers
How does an AI risk assessment work?
An AI risk assessment runs in three steps: first the use and possible harm are clarified, then failure modes and controls are derived from it, and last the acceptance criteria are fixed before the first test runs. The result is a risk table that every test hangs on. Your team carries out the assessment for its systems; I provide the method and template and support the start.
How many test runs do you need to show a failure rate?
Sixty independent test runs without a failure give a failure rate below 5 percent at 95 percent confidence, and three thousand give below 0.1 percent. The runs must be independent; repeats of the same input are only approximately so. The bound only holds for the cases tested, not for unknown ones. The table on this page shows more values, computed as 1 minus 0.05 to the power 1 over n.
When is testing alone not enough?
Testing alone is not enough when a single failure weighs heavily and is rare, because rare failures can only be bounded with very many runs. Then the safety is put around the model: permissions, locks, approvals and a stop switch. Where a risk cannot be shown sufficiently, I recommend controls outside the model, or not doing it.
Contact
Send me the use case in two sentences. I will get back to you and say whether and how your project can be tested. Whether I can take the job depends on my workload.