All articles

AI Prompting

Evaluating Prompts Like Code: Test Suites for LLM Behavior

A prompt is a piece of production logic with no type system. Here is how to build the evaluation harness that replaces the compiler you do not have.

June 28, 2026 · 7 min read · AI Research, Quantum Tech & IT Strategy Consulting

Teams ship prompts the way nobody would ship code: edited in place, verified by trying three examples by hand, deployed with no regression suite. It works until the day a small wording change to fix one customer complaint quietly breaks a behaviour that had been correct for months. The fix is not more careful editing. It is an evaluation harness.

Start with failures, not with coverage

Do not try to build a representative dataset on day one. Build a failure corpus. Every time someone reports a bad output, capture the exact input, the actual output, and one sentence describing what was wrong with it. Within a few weeks you will have a set of cases that concentrates precisely on the behaviours your system gets wrong, which is far more useful for iteration than a balanced sample of cases it already handles.

Three tiers of assertion

Not every property needs the same checking machinery, and using the heaviest tool for everything is how evaluation suites become too slow and expensive to run. Sort your assertions into tiers and run them in order of cost.

  • Deterministic checks: schema validity, required fields, citation identifiers resolving, forbidden phrases absent, numeric ranges. Fast, free, and catches the majority of regressions.
  • Programmatic similarity: exact match or fuzzy match against reference answers for tasks with a correct answer, such as extraction and classification.
  • Model-graded rubrics: for tone, completeness, and reasoning quality. Use a rubric with explicit criteria and require the grader to cite evidence; calibrate it against human labels before you trust it.

Version the prompt as an artefact

Prompts belong in the repository, not in a database field that someone edits through an admin panel at midnight. Give each prompt a version identifier, record that identifier with every logged inference, and treat a prompt change as a code change with a diff, a reviewer, and a test run. When an incident happens, you want to be able to answer 'which version produced this?' in seconds.

Watch the variance, not just the mean

Model outputs are stochastic. A pass rate of 94 percent on a single run tells you less than you think, especially on a suite of eighty cases. Run the suite multiple times at your production temperature and look at the distribution. A change that raises average quality while doubling variance is often worse in production than one that does nothing, because users experience the tail, not the average.

The point of an evaluation suite is not to prove the system is good. It is to make regressions visible before customers find them.

Close the loop with production traffic

Sample real production inferences continuously, grade a slice of them automatically, and promote interesting cases into the test suite. This is how the suite stays relevant as usage drifts away from what you imagined when you built the thing. It also gives you an early-warning signal when a model provider updates a model beneath you — a scenario every team running on hosted models should assume will happen without notice.

None of this requires exotic tooling. A directory of test cases, a runner script, a set of assertions, and a results table checked into the repository will take a team most of the way. What matters is that the harness exists and runs before a prompt change ships.