# Evals that keep an AI agent honest

> How the AI agent on this site is tested like code: 44 questions against the real model, retrieval gated in CI, a versioned prompt and a trace under every answer. Published 24 September 2026.

The chat on this site answers as me, from what the site says. That is a promise with a lot of ways to break: an answer can cite a page that doesn't say it, round a number, speak about me instead of as me, or promise a meeting I never agreed to. A model that is right most of the time is still wrong in public.

So the agent is tested the way code is. A prompt change ships only after it passes, and every answer carries the record of how it was made.

*Diagram:* One answer, to scale, from its own trace: the card arrives while the loop is still going; the words take a second request.

## Tests that need no model

Most of what can go wrong has nothing to do with the model's judgement, so most of the tests run without one. The stream parser is cut at every possible point. Tool calls are reassembled across chunks and across providers. Keys rotate, schemas are checked, every shape of leaked prompt is caught, and the guards hold. Retrieval is measured too: each eval question has the page its answer lives on, and CI fails if the right page is in the top five less than 90% of the time.

## Questions against the real model

The eval set is 44 questions. Each one says where the answer lives and which words a right answer has to contain, and the set includes follow-ups, persona checks and attempts to talk it out of its rules. A script runs them against the real model before a prompt change ships.

The checks are deliberately crude. A number in an answer that the site never says is a failure, whatever the model meant by it: on the last run, the one miss rounded 12,493 downloads to 12,000. An answer that talks about "Rajveer" instead of "I" is a failure. Saying it is an AI, and handing over instead of promising, are checked on their own.

**The last run**

| Check | Result |
| --- | --- |
| Tool choice | 21 of 23 |
| Says the words a right answer must | 24 of 26 |
| No invented numbers | 43 of 44 |
| First person | 40 of 44 |
| Persona | 6 of 6 |

Against the real model on 1 October 2026, spaced out to stay under the free tier's limit a minute.

## The prompt has a version

The prompt is code, so it has a version, and the version goes into every trace and into the key of every kept answer. When the prompt changes, an answer written by the old one is never served again. A regression can be traced to the exact prompt that produced it.

## Every answer shows its work

Under every answer is its trace: each step as a span, when anything first appeared on screen, tokens in and out, which provider answered, and the prompt's version. It records timings and names, never your words. Over 20 timed questions on a production build, something was on screen within 0.3 s for 14 of them, and the first words came at 1.85 s at the median.

## Measuring before building more

I built search by meaning as well as by words, and measured both on the same questions before choosing. Words alone found the right page in the top five for 30 of 32; adding meaning found 31, for about half a second more on every question. At this size that doesn't pay, so it is off. It stays written, measured and ready for when the site grows past the point where words are enough.

## What's next

A nightly run that asks a sample of the questions and messages my phone when answers drop below the bar is written and waiting on the model's keys in CI. After that: routing between models when the evals say a cheaper one is good enough, and a semantic cache for the questions people ask most.
