Checks
Testing LLMs and AI agents
On top of classic ML testing, language models add three things: the answer is not fixed, inputs can contain commands, and tools have permissions.
- Gold set and regressionFixed cases with a known answer. Runs on every change to model, prompt, data or tools.
- Repeated runsEvery case n times, because answers vary. A failure rate with an upper bound instead of “it worked once”. Security cases demand zero failures.
- GroundednessEvery statement must be backed by a passage in the source document. Code checks the quotes, not the model.
- Prompt injectionDirect and indirect attacks through documents, mails, web pages. A secret marker in the system prompt: if it shows up in an answer, it leaked.
- Tool permissionsOnly what is needed (least privilege), confirmation on write access, sandbox, cost limit, maximum number of steps.
- Tenant separation and data leaksForeign data in the context must never appear in the answer. Output filters and scanners as a second line.
- Misuse and refusalsA red-team set for misuse and a set for over-caution: refusing too much is also a failure.
- Four languagesThe same test in German, French, Italian and English, plus dialect and mixed language.
- AI as judgeOnly after calibration against human judgements (agreement κ) and never alone. The judge is tested too.
- Model changePin the version, compare old against new on every update, read the deviation report.
- Loops and costStop conditions, caps on steps, tokens and money. An agent without a brake is a risk.
- Oversight and loggingAn approval step where harm can occur. Prompt, version, tool calls and decision are logged so that you can trace later what happened.
Answers
How do I test a language model for prompt injection?
You test a language model for prompt injection by sending documents with built-in commands, in several languages, through the agent and counting how often it misuses tools, outputs a secret marker or names foreign data. Every case runs several times because answers vary. For security cases one failure counts as a rejection.
Why is a single test run not enough for a language model?
A single test run is not enough for a language model because the same input can produce different answers. An agent that breaks the rules in 4 of 100 runs passes twelve runs with a probability of just over 60 percent, but one hundred and twenty runs with less than 1 percent. So every case runs repeatedly and the bound is stated with the number of runs.
What permissions should an AI agent have?
An AI agent should only have the permissions its task needs (least privilege): read access to this customer's data only, write access only with confirmation, a cost limit and a cap on steps. Document content counts as data, never as an instruction. Whether these locks hold is checked with attack cases.
Contact
Send me the use case in two sentences. I will get back to you and say whether and how your project can be tested. Whether I can take the job depends on my workload.