Checks
Testing ML models: all angles
Each line is a question, a test as code and a criterion. The criterion is an example; in a project it comes from the risk analysis. “In the demo” means an excerpt of it really runs in the proof above (for example only sensitivity, only exact row hashes). All other lines are method; which of them a project needs follows from the risk analysis. They are not demonstrated here.
| Angle | Question | Test as code | Example criterion |
|---|---|---|---|
| Data | |||
| Quality and schema | Do types, value ranges, required fields and duplicates hold? | Schema and range tests on every data import | 0 violations or a documented exception |
| Leakage and splitting In the demo (excerpt): T01 | Is a test case in the training data, even indirectly (same person, same customer, the future)? | Row-hash intersection, group and time split (the demo only does the hash intersection) | 0 overlap |
| Representativeness and labels | Does the data cover real use? Can the labels be trusted? | Distribution comparison, two annotators, agreement κ | κ ≥ 0.8; gap to real use documented |
| Provenance and rights | Where does the data come from, under which licence, with which personal data? | Provenance register, personal-data scanner, retention rule | Provenance, licence and personal-data link of every source documented; legal assessment by a lawyer |
| Model | |||
| Performance with confidence interval In the demo (excerpt): T02 | How often is it right, weighted by cost of error? | Sensitivity and specificity with Wilson bound, not accuracy alone | Lower 95% bound ≥ threshold |
| Subgroups and fairness In the demo (excerpt): T03 | Does performance hold for every group (language, region, device, age)? | Evaluation per subgroup with a minimum case count | Every group ≥ threshold, otherwise excluded |
| Calibration and uncertainty In the demo (excerpt): T06, T08 | Are the stated probabilities right? Does it know when it does not know? | ECE, reliability diagram, hand-over to humans | ECE ≤ 0.05; error rate of automatic cases bounded |
| Behaviour In the demo (excerpt): T05 | Does it behave as required in known situations? | Invariance, directional and minimum-functionality tests (CheckList), metamorphic tests | 0 violations |
| Robustness In the demo (excerpt): T04 | Does it survive noise, typos, format changes, language changes? | Perturbed inputs, measure the flip rate | Flip rate ≤ 1% |
| Attacks | Can it be fooled, poisoned or probed? | Adversarial examples, data-poisoning test, extraction test | Every attack found has a countermeasure |
| Privacy | Does it give away training data or personal data? | Memorisation and membership-inference test, PII scan of outputs | 0 hits in the test set |
| Explainability | Do the reasons support the decision, or just decorate it? | Attribution sanity tests, counterfactuals | Explanation changes with the decision |
| Repeatability In the demo (excerpt): T07 | Does the same state give the same result? | Seeds, versions, data hash, bit-identical comparison | Hash identical |
| System | |||
| Integration and contract | What happens on a timeout, an error, an empty reply? | Contract tests on the interface, error-path tests, fallback without AI | Every error path ends safely |
| Load, time, cost | Does it hold under load, and what does a request cost? | Load test, p95 latency, cost per request, budget cap | p95 and budget within limits |
| Human and process (UAT) | Do real people actually use the control? | Acceptance with users, decoy cases, post-training test | Decoy detection rate ≥ 90% |
| Supply chain | Where do the model and libraries come from, are they pinned? | Versions and checksums pinned, bill of materials (SBOM, ML-BOM), licences documented | Everything pinned, bill of materials current |
| Operations | |||
| Drift and monitoring | Does the world change under the model? | Drift index (PSI) on inputs, sampled review of outputs, alert thresholds | PSI < 0.25; every alert has an owner |
| Incidents and stop | Who stops it, how fast, how does it roll back? | Stop switch and rollback as a drill, incident register | Drill passed, time measured |
| Change and re-release | What happens with a new version, a new prompt, new data? | Every change triggers the whole gate | No change without a green gate |
Answers
What do you check in an ML model besides accuracy?
Besides accuracy you check the data (quality, leakage, representativeness), the behaviour (subgroups, calibration, robustness), the system (integration, load, the human in the process) and operations (drift, incidents, changes). The checklist distinguishes twenty checks. Which ones your model needs follows from the risk analysis.
How do I detect drift in an AI model in operation?
You detect drift by continuously comparing the inputs in operation with the training data, for example with the Population Stability Index (PSI). A PSI above 0.25 is often used as a rule of thumb for a clear shift; you set the threshold per input. What matters is that every alert has an owner.
Why is a single metric not enough?
A single metric is not enough because different failures can produce the same number. In the demo a model squeezed to be too cautious keeps the same sensitivity and fails only on calibration and automation coverage. So data, behaviour, robustness, subgroups, calibration and operations are checked separately.
Contact
Send me the use case in two sentences. I will get back to you and say whether and how your project can be tested. Whether I can take the job depends on my workload.