Blog3 min read
Evals that keep an AI agent honest
How the AI agent on this site is tested like code: 44 questions against the real model, retrieval gated in CI, a versioned prompt and a trace under every answer.
The chat on this site answers as me, from what the site says. That is a promise with a lot of ways to break: an answer can cite a page that doesn't say it, round a number, speak about me instead of as me, or promise a meeting I never agreed to. A model that is right most of the time is still wrong in public.
So the agent is tested the way code is. A prompt change ships only after it passes, and every answer carries the record of how it was made.
Tests that need no model
Most of what can go wrong has nothing to do with the model's judgement, so most of the tests run without one. The stream parser is cut at every possible point. Tool calls are reassembled across chunks and across providers. Keys rotate, schemas are checked, every shape of leaked prompt is caught, and the guards hold. Retrieval is measured too: each eval question has the page its answer lives on, and CI fails if the right page is in the top five less than 90% of the time.
Questions against the real model
The eval set is 44 questions. Each one says where the answer lives and which words a right answer has to contain, and the set includes follow-ups, persona checks and attempts to talk it out of its rules. A script runs them against the real model before a prompt change ships.
The checks are deliberately crude. A number in an answer that the site never says is a failure, whatever the model meant by it: on the last run, the one miss rounded 12,493 downloads to 12,000. An answer that talks about "Rajveer" instead of "I" is a failure. Saying it is an AI, and handing over instead of promising, are checked on their own.
| Check | Result |
|---|---|
| Tool choice | 21 of 23 |
| Says the words a right answer must | 24 of 26 |
| No invented numbers | 43 of 44 |
| First person | 40 of 44 |
| Persona | 6 of 6 |
The prompt has a version
The prompt is code, so it has a version, and the version goes into every trace and into the key of every kept answer. When the prompt changes, an answer written by the old one is never served again. A regression can be traced to the exact prompt that produced it.
Every answer shows its work
Under every answer is its trace: each step as a span, when anything first appeared on screen, tokens in and out, which provider answered, and the prompt's version. It records timings and names, never your words. Over 20 timed questions on a production build, something was on screen within 0.3 s for 14 of them, and the first words came at 1.85 s at the median.
Measuring before building more
I built search by meaning as well as by words, and measured both on the same questions before choosing. Words alone found the right page in the top five for 30 of 32; adding meaning found 31, for about half a second more on every question. At this size that doesn't pay, so it is off. It stays written, measured and ready for when the site grows past the point where words are enough.
What's next
A nightly run that asks a sample of the questions and messages my phone when answers drop below the bar is written and waiting on the model's keys in CI. After that: routing between models when the evals say a cheaper one is good enough, and a semantic cache for the questions people ask most.
Keep reading
A load balancer from scratch, A/B tested on live trafficA load balancer in Node for five moderation replicas: four strategies, a circuit breaker and health check each, failover, and a load test that proves the spread.Read
A lake that never makes the page waitHow the WebGL lake at the top of this site paints in a worker, compiles its shaders alongside, holds to a pixel budget, and costs nothing when you can't see it.Read
Shipping to one box without a staging serverCI/CD for a live EC2 box with 3,000+ members and no staging server: an SSH push with pinned keys, one deploy at a time, a 90 second health gate and a rollback.Read
Voice rooms on an SFUHow IABTM's live audio rooms moved from a peer-to-peer mesh to a mediasoup SFU: a router per room, five-speaker stages and a worker that restarts on its own.Read
A chapter site that keeps itself current from InstagramA FastAPI service that reads an IEEE student chapter's Instagram, has Gemini turn each post into a structured event, drops reposts and feeds the chapter's PWA.Read