AI Risk

Position

Proof before claim

Ten theses the framework builds on and my work can be measured against.

As of: 6 October 2026

  1. “Safe” is a measurement. One system, one area of use, one version, one number with a confidence interval. I do not say “the model is safe” without these.
  2. Proof before claim. Order: risk, test, measurement, claim. Whoever writes the claim first only goes looking for evidence afterwards.
  3. Top down. The method starts with the harm a person suffers, not with the model. Every test hangs on a risk. Every risk has at least one test or a reasoned exception.
  4. Tests are code. Versioned and repeatable, in a project triggered automatically on every change. A document that says “tested” is not a test.
  5. The test rig must be able to go red. It is tested against models broken on purpose. A test that never fails says nothing about the model.
  6. Too little data is a finding. 60 independent runs without a failure only show a failure rate below 5 percent (95% confidence). What cannot be shown, is taken out of scope and named.
  7. One metric is never enough. A model with data leakage can show flawless numbers. So data, behaviour, robustness, subgroups, calibration and operations are checked separately.
  8. The human sits where a mistake does harm. Human oversight is tested too: is it used, or do people click through? That is what decoy cases are for.
  9. Measuring continues after go-live. Drift, incidents, a stop switch. A release applies to one version and expires with every change.
  10. A certificate does not replace proof. An ISO/IEC 42001 certificate confirms the management system. Whether a model is safe enough in your use can only be shown by a test on the model, and then only for the cases tested.

Contact

Send me the use case in two sentences. I will get back to you and say whether and how your project can be tested. Whether I can take the job depends on my workload.

admin@all-answer.com
+41 76 511 52 25