AI Risk

Checks

Testing ML models: all angles

Each line is a question, a test as code and a criterion. The criterion is an example; in a project it comes from the risk analysis. “In the demo” means an excerpt of it really runs in the proof above (for example only sensitivity, only exact row hashes). All other lines are method; which of them a project needs follows from the risk analysis. They are not demonstrated here.

As of: 6 October 2026

AngleQuestionTest as codeExample criterion
Data
Quality and schema Do types, value ranges, required fields and duplicates hold? Schema and range tests on every data import 0 violations or a documented exception
Leakage and splitting In the demo (excerpt): T01 Is a test case in the training data, even indirectly (same person, same customer, the future)? Row-hash intersection, group and time split (the demo only does the hash intersection) 0 overlap
Representativeness and labels Does the data cover real use? Can the labels be trusted? Distribution comparison, two annotators, agreement κ κ ≥ 0.8; gap to real use documented
Provenance and rights Where does the data come from, under which licence, with which personal data? Provenance register, personal-data scanner, retention rule Provenance, licence and personal-data link of every source documented; legal assessment by a lawyer
Model
Performance with confidence interval In the demo (excerpt): T02 How often is it right, weighted by cost of error? Sensitivity and specificity with Wilson bound, not accuracy alone Lower 95% bound ≥ threshold
Subgroups and fairness In the demo (excerpt): T03 Does performance hold for every group (language, region, device, age)? Evaluation per subgroup with a minimum case count Every group ≥ threshold, otherwise excluded
Calibration and uncertainty In the demo (excerpt): T06, T08 Are the stated probabilities right? Does it know when it does not know? ECE, reliability diagram, hand-over to humans ECE ≤ 0.05; error rate of automatic cases bounded
Behaviour In the demo (excerpt): T05 Does it behave as required in known situations? Invariance, directional and minimum-functionality tests (CheckList), metamorphic tests 0 violations
Robustness In the demo (excerpt): T04 Does it survive noise, typos, format changes, language changes? Perturbed inputs, measure the flip rate Flip rate ≤ 1%
Attacks Can it be fooled, poisoned or probed? Adversarial examples, data-poisoning test, extraction test Every attack found has a countermeasure
Privacy Does it give away training data or personal data? Memorisation and membership-inference test, PII scan of outputs 0 hits in the test set
Explainability Do the reasons support the decision, or just decorate it? Attribution sanity tests, counterfactuals Explanation changes with the decision
Repeatability In the demo (excerpt): T07 Does the same state give the same result? Seeds, versions, data hash, bit-identical comparison Hash identical
System
Integration and contract What happens on a timeout, an error, an empty reply? Contract tests on the interface, error-path tests, fallback without AI Every error path ends safely
Load, time, cost Does it hold under load, and what does a request cost? Load test, p95 latency, cost per request, budget cap p95 and budget within limits
Human and process (UAT) Do real people actually use the control? Acceptance with users, decoy cases, post-training test Decoy detection rate ≥ 90%
Supply chain Where do the model and libraries come from, are they pinned? Versions and checksums pinned, bill of materials (SBOM, ML-BOM), licences documented Everything pinned, bill of materials current
Operations
Drift and monitoring Does the world change under the model? Drift index (PSI) on inputs, sampled review of outputs, alert thresholds PSI < 0.25; every alert has an owner
Incidents and stop Who stops it, how fast, how does it roll back? Stop switch and rollback as a drill, incident register Drill passed, time measured
Change and re-release What happens with a new version, a new prompt, new data? Every change triggers the whole gate No change without a green gate

Answers

What do you check in an ML model besides accuracy?

Besides accuracy you check the data (quality, leakage, representativeness), the behaviour (subgroups, calibration, robustness), the system (integration, load, the human in the process) and operations (drift, incidents, changes). The checklist distinguishes twenty checks. Which ones your model needs follows from the risk analysis.

From risk analysis to test

How do I detect drift in an AI model in operation?

You detect drift by continuously comparing the inputs in operation with the training data, for example with the Population Stability Index (PSI). A PSI above 0.25 is often used as a rule of thumb for a clear shift; you set the threshold per input. What matters is that every alert has an owner.

Why is a single metric not enough?

A single metric is not enough because different failures can produce the same number. In the demo a model squeezed to be too cautious keeps the same sensitivity and fails only on calibration and automation coverage. So data, behaviour, robustness, subgroups, calibration and operations are checked separately.

See the counter-examples

Contact

Send me the use case in two sentences. I will get back to you and say whether and how your project can be tested. Whether I can take the job depends on my workload.

admin@all-answer.com
+41 76 511 52 25