Position
Proof before claim
Ten theses the framework builds on and my work can be measured against.
- “Safe” is a measurement. One system, one area of use, one version, one number with a confidence interval. I do not say “the model is safe” without these.
- Proof before claim. Order: risk, test, measurement, claim. Whoever writes the claim first only goes looking for evidence afterwards.
- Top down. The method starts with the harm a person suffers, not with the model. Every test hangs on a risk. Every risk has at least one test or a reasoned exception.
- Tests are code. Versioned and repeatable, in a project triggered automatically on every change. A document that says “tested” is not a test.
- The test rig must be able to go red. It is tested against models broken on purpose. A test that never fails says nothing about the model.
- Too little data is a finding. 60 independent runs without a failure only show a failure rate below 5 percent (95% confidence). What cannot be shown, is taken out of scope and named.
- One metric is never enough. A model with data leakage can show flawless numbers. So data, behaviour, robustness, subgroups, calibration and operations are checked separately.
- The human sits where a mistake does harm. Human oversight is tested too: is it used, or do people click through? That is what decoy cases are for.
- Measuring continues after go-live. Drift, incidents, a stop switch. A release applies to one version and expires with every change.
- A certificate does not replace proof. An ISO/IEC 42001 certificate confirms the management system. Whether a model is safe enough in your use can only be shown by a test on the model, and then only for the cases tested.
Contact
Send me the use case in two sentences. I will get back to you and say whether and how your project can be tested. Whether I can take the job depends on my workload.