The AI me
this site · the Ask tab- side project
- In progress
- 2026
- RAG
- Tool calling
- Generative UI
- Guardrails
- Evals
The AI me is a version of me that answers for my work while I’m away. Open the Ask tab on any page and ask: it reads what this site says, answers in my voice with its sources linked, and puts the real things between its words (the project cards, my calling card, my calendar) instead of describing them. Tell it to, and it moves around the site for you: it scrolls to a part, takes you to a page, makes it night or lets it rain, each with an undo.
It is honest about what it is. It says it’s an AI when asked, and when the site doesn’t say something it says so and offers to send your question to the real me rather than guess. Recruiters can paste a job description and get a table of where my work fits and where it doesn’t.
I built every layer from scratch, on no agent framework and no vendor’s SDK, so any chat-completions model can run it: the engine, retrieval, tools, guards, interface, tracing and evals. It’s live and still growing. Under the paragraphs that describe a feature there is a question to try: press it and the chat opens with it typed in, for you to send and watch that feature happen.



- Type
- side project
- Role
- Everything: the engine, retrieval, tools, guards, interface and evals
- When
- Sep 2026 to now
- Stack
- Next.js, React 19, TypeScript, Vercel Functions, Upstash Redis, MongoDB Atlas, Server-sent events, NDJSON, JSON Schema, Bun test, Playwright
- Links
- Ask it: the tab on the right, or press aIs it up?
- tools it can call, each checked
- 31
- fetched only when you reach for it
- 25 KB
- questions find their page in the top five
- 31 of 32
- unit tests on the engine
- 284
What it does
Answers from the site, with sources
Every answer is built from passages of these pages, and a small number after a sentence links to the part it came from.
How it worksThe site’s own components in the chat
The same cards and buttons the pages use, not descriptions of them, shown when the answer is one.
How it worksActs on the page, with an undo
Scrolls to a part, rings it, takes you to a page, switches day and night or the weather, opens every way to reach me.
How it worksA fit check for recruiters
Paste a job description; see each requirement against my work, strong, some or a gap, linked to where it is shown.
How it worksHonest by construction
Grounded in what the site says, cited, corrected when it claims what didn’t happen, and handing over instead of guessing.
How it worksFour providers, then the pages
Keys in turn, backups each with its own quota, answers from the pages alone, then the real me.
How it worksGuarded against abuse
A session per visitor, limits in six windows, a daily token budget, a breaker per provider.
How it worksPrivate by default
Counts, not transcripts; 30 days of redacted questions for my review; your conversation stays on your device.
How it worksFree until you reach for it
The chat’s code arrives when a hand nears the tab, and a command is done without a model at all.
How it worksTraced, counted and reported
Every answer carries its steps, timings and tokens; my phone hears about fit checks, trouble, and the day at 8 pm.
How it worksEvals that gate the prompt
44 questions scored against the real model, and retrieval checked in CI.
How it works
How it fits together
What it’s for, and who
It is for when I’m away, which is most of the time someone reads this site. A portfolio gets the same few questions, about the work, how it was made and whether we should talk. The AI me answers them in a couple of seconds, shows the proof, and hands over to the real me when it should.
A recruiter pastes a job description and gets a fit check: the role’s requirements, the evidence from my work beside each, how strong a match it is, and the gaps said plainly and sorted last, then a way to talk.
A founder says what they’re building and gets the closest thing I’ve shipped. Someone in hiring asks about notice periods or rates and is told, politely, that it can’t promise anything for me, with my calendar one tap away.
An engineer asks how it works and can open the steps, timings and tokens of their own answer, which is what the rest of this page is about.
The fit check is one request for JSON only, and what comes back is checked before it is shown: its shape against a schema, every link against the pages of this site, and a gap never gets a link. If the reply doesn’t hold together, it says so and offers to have me look at the description myself.
One question, end to end
Take “How do you stop double charges in billing?”, asked on the home page. This is one real run, timed.
The chat’s code was fetched when the pointer neared the tab, and it asked for a session when it opened. The question goes out with that session, the page it was asked on, the labels of that page’s parts, the time and weather showing, and the last few turns with notes of what they showed and did.
The guards count it in six windows, look for a kept answer, and check both providers’ breakers and the day’s tokens, all in one round trip to the store.
Then the turn. The five best passages are found in 8 ms. Nothing in this question asks to see a card, so the model is asked, with 17 of the 31 tools on offer. Its first request ends by calling for the IABTM card, which is on screen at 1.28 s.
That request said nothing, so the loop goes again, and the second writes the answer, citing passage 5, the billing section of the IABTM write-up. First words at 2.35 s, done at 2.73 s: two requests, 10,260 tokens in, 95 out.
Under the answer, one line folds everything it did: the pages it read, the tools it ran, the time it took. Open it for each step, then each request’s milliseconds and tokens.
An engine that doesn’t care what carries it
The same brain will answer by voice, and later on other people’s sites, so it is a core with nothing about the chat in it.
One function, runTurn, takes a question, the turns before it and the page it was asked on, and yields events: the passages it read, words, a component to show, an action to take on the page, what it’s busy with, a trace, and done.
The chat’s route writes them one JSON object a line (NDJSON) as they come, so a card can arrive before the model has said a word. A voice transport will turn the same events into speech and captions. The eval script runs the core with no HTTP at all.
The route waits for the first word or card before it answers. A model that never begins is then said plainly with a status, instead of a stream that opens and breaks. Every turn has a deadline, and a stream that goes quiet partway is given up on rather than left hanging.
lib/agent/ui.ts · what a turn streams
/** A turn as it streams to the chat: one JSON object a line. */export type TurnEvent = | { t: 'sources'; sources: Source[]; kept?: boolean; /** Answered from the pages alone. */ pages?: boolean; } | { t: 'text'; v: string } /** 'thinking', 'writing', or a tool's name. */ | { t: 'doing'; what: string; detail?: string } | ({ t: 'show'; id: string } & Shown) | ({ t: 'act'; id: string } & Act) | { t: 'trace'; trace: Trace } | { t: 'error'; m: string } | { t: 'done' };/** Something done on the page rather than shown in the chat. */export type Act = | { a: 'scroll'; mark: string } | { a: 'highlight'; mark: string } | { a: 'navigate'; href: string; label: string } | { a: 'time'; time: 'day' | 'night' } | { a: 'season'; season: string } | { a: 'contact' };Any model, and a line of backups
The model is an address, a name and keys in the environment: any endpoint that speaks chat completions and streams server-sent events. Nothing in the code names a provider.
The stream is read by hand: words as they come; each tool call assembled from its pieces, by index or by id, and handed over the moment it is whole; arguments whether sent as text or as an object; any thinking an endpoint sends beside the words, never shown; and the tokens used, estimated when an endpoint doesn’t report them.
Keys are taken in turn. One the endpoint refuses rests a minute (a rate limit) or an hour (spent, revoked, suspended, or a 400 whose words name the key, which is how one provider says a key is bad), and the next is tried within the same request. Keys in one account or project share its quota, so rotation covers a failing key, not more free use.
Behind the first provider is a line of backups, three today, each wholly apart: its own endpoint, model, keys, quota and breaker. The next is asked when every key before it is resting or refused, when one answers with any error or nothing, while its breaker is open, or when the first has written nothing after 4 seconds (a backup gets 8, so two slow ones still fit inside the turn’s 25). A turn a backup starts, it finishes, and a question counts as one failure against a provider however many of its keys were tried.
The backups are offered a slimmer set of tools, the ones most answers need, because their allowance of tokens a minute is often the smaller one. Only when every provider is down does it answer from the pages alone, and my phone hears about it at once.
| Try | Who | Why it moves on |
|---|---|---|
| 1 | The first provider, each key in turn | A key refused (429, 401, 403, or a 400 naming the key) rests; the next is tried at once. |
| 2 | The same endpoint’s spare model, if one is set | The first model fails or is slow to start. |
| 3 | Each backup in order, each of its keys | No word yet after 4 s (8 s for a backup), any error, or nothing at all. |
| 4 | The pages alone | Every provider failing or its breaker open, or the day’s tokens spent: links to the best passages, and the obvious card. |
| 5 | The real me | Anything that still fails ends in a note to my phone, on the visitor’s press. |
A short loop, and tools as data
The loop is short on purpose: at most three requests and 20 seconds a question. The model writes, or calls tools.
A show tool’s component goes to the visitor as soon as its call is whole. A data tool’s result goes back to the model for another pass. The last pass is offered no tools, so it has to answer, and once it has shown something and said something, the turn ends. Across 20 timed questions it took 1.5 requests on average.
Every capability is one entry in a registry: a name, what the model is told, a JSON schema for its arguments, and what it does. The schema the model is shown is the one its arguments are checked against. Properties it doesn’t name are dropped, and a bad call (an unknown tool, broken JSON, an argument out of range) comes back to the model as an error to read, never a throw.
Only 17 of the 31 are offered by default, the rest when the question’s words ask for them, because every tool offered is read before the first word. The page actions are offered only for a command.
lib/agent/turn.ts · the loop
export const MAX_STEPS = 3;export const BUDGET_MS = 20_000;for (let step = 0; step < MAX_STEPS; step++) { const last = step === MAX_STEPS - 1 || tracer.elapsed() > BUDGET_MS * 0.6; for await (const event of streamStep(config, messages, { tools: last ? undefined : specs })) { if (event.type === 'text') yield* words(event.text); else if (event.type === 'call') { calls.push(event.call); const tool = registry.get(event.call.function.name); // What shows, shows now; what looks up waits // for the words. if (tool && tool.kind !== 'data') results.push(yield* call(/* the call */)); } } if (!calls.length) break; for (const c of data) results.push(yield* call(/* the lookup */)); // Shown, and something said: nothing more to ask for. if (!data.length && text.trim()) break; if (last) break; messages.push(/* the calls, then each result */);}tools/page.ts and tools/registry.ts
{ name: 'navigate', kind: 'page', description: 'Take the visitor to another page of the site now, ' + 'when they ask to go there.', parameters: { type: 'object', properties: { path: { type: 'string', enum: sitePages().map((p) => p.path) }, }, required: ['path'], }, run(args) { const page = sitePages().find((p) => p.path === args.path)!; return { result: `Going to ${page.path}.`, act: { a: 'navigate', href: page.path, label: page.what }, }; },},/* ... and every call, whatever the model wrote: */const tool = this.get(name);if (!tool) return { ok: false, result: { error: `There is no tool called ${name}.` } };const parsed = parseArguments(rawArguments);if (!parsed.ok) return { ok: false, result: { error: parsed.error } };const args = check(tool.parameters, parsed.value);if (!args.ok) return { ok: false, result: { error: args.error } };Cards before the model, and only when asked
Some questions don’t need a model to know what to show. A page of patterns plans the obvious calls and runs them before the model is asked; it is told they are done and not offered them again. That is why a card can be on screen in milliseconds.
A plain command (“make it night”) goes further: it is done and said in a line with no passages and no model at all.
The patterns learned restraint the hard way. They used to fire on questions: a question about summer changed the weather, one about a day job made it day. Now an action needs a command, a card comes before the model only for an explicit ask to see it, and the model’s own cards are one to an answer unless a list was asked for, never the one shown just before.
The components are the site’s own: the same project cards, calling card, résumé and calendar as the pages, drawn in the chat from props the server builds out of the site’s content. The model names a project by its slug, from a list; it can’t describe one into being.
lib/agent/turn.ts · the obvious calls
// The obvious calls first, before the model has// started (intent.ts); it is told they are done,// and not offered them again.const planned = plan(ask);const done: string[] = [];for (const p of planned) { const out = yield* call(p.name, JSON.stringify(p.args)); done.push(`${p.name}: ${JSON.parse(out.content)}`);}if (done.length) { const names = new Set(planned.map((p) => p.name)); specs = specs.filter((s) => !names.has(s.function.name)); messages[0].content += `\n\nALREADY SHOWN to the visitor for this question` + ` (don't call these again; write your answer` + ` around them):\n${done.join('\n')}`;}Memory, and what a follow-up carries
A conversation is only useful if “tell me more about that” means something. Each question carries the last six turns, and with each answer a note of what it showed and did, so the model knows the card on screen and the night it made.
The last four turns are kept whole, and the facts and passages are fitted into what is left. It used to be the other way round: the prompt’s budget went to the facts and passages first, and a long answer or two starved the conversation out of it, so a follow-up lost its thread. Keeping memory first, and a prompt of up to 28,000 characters, fixed that.
Retrieval reads a follow-up with the turn before it, without a second model call: a short question that points back borrows the last question’s words and the projects the last answer named.
Retrieval from the site’s own pages
Every answer is built from what this site says, and says where.
The corpus is the site itself. Every page already has a markdown copy for agents (this one is at /projects/ai-me.md). Those are cut at their headings, then into passages of 400 to 900 characters at paragraph and sentence ends, each with a title and the anchor of its own section. With two more about the site itself (its day, night and seasons, and this chat), that is 173 passages from 14 pages, this one included, made once, in memory.
The fact sheet, about 7,300 characters of who I am, every job and project in a line, and the record, goes into every prompt whole, so a question about dates never depends on retrieval.
Passages are found by their words with BM25: titles counted twice, a light stemmer, stop words out, a small table of a visitor’s words for the words the pages use, a few phrases read for what they mean, and a lift of 1.6 for the passages of a project the question names. It takes under a fifth of a millisecond a question.
They can also be found by meaning, where the words are unsure. The passages’ vectors are made once, when the site is built, only for what changed, and kept in MongoDB Atlas; a question embeds itself, Atlas Vector Search finds its neighbours, and the two ranked lists are merged by reciprocal rank, with a 700 ms cap. Measured on the 32 eval questions it didn’t pay: 30 in the top five against the words’ 31, and about 0.4 s a question. So it is off, ready for a better embedding model or a larger corpus, and the words answer alone.
On a write-up, a question that names no project is about that one: asked here, “this” means the AI me.
Citations are checked, not trusted. The model may cite only the passages it was given, by number. A citation to anything else is dropped as the words stream, one written loosely, as [2], is made a real one, and a search mid-turn adds its passages to the list so they can be cited too.
| Retrieval | Words only | Words and meaning |
|---|---|---|
| hit@1 | 21 of 32 | 26 of 32 |
| hit@3 | 29 of 32 | 30 of 32 |
| hit@5 | 30 of 32 | 31 of 32 |
| MRR@5 | 0.78 | 0.87 |
| Time a question | Under 0.2 ms | 0.52 s median, 1.14 s at p90 |
lib/agent/retrieve.ts · BM25
for (const doc of this.docs) { let score = 0; for (const [w, weight] of q) { const f = doc.tf.get(w); if (!f) continue; const n = this.df.get(w) ?? 0; const idf = Math.log(1 + (N - n + 0.5) / (n + 0.5)); score += weight * idf * ((f * (k1 + 1)) / (f + k1 * (1 - b + (b * doc.length) / this.avg))); } if (score > 0 && doc.chunk.slug && named.has(doc.chunk.slug)) score *= 1.6; if (score > 0) hits.push({ chunk: doc.chunk, score });}Tools, and acting for you
It can do 31 things, each on a list it can’t add to. Five look something up for the model. Eighteen put one of the site’s components in the chat: my timeline, what I’m reading, my local time, a drafted note and the rest. Eight act on the page, a tour of it and undoing the last change among them.
Told to, and only when told to, it scrolls to a part of this page, rings a part for a moment, walks you through the page’s parts one at a time, takes you to another page, makes it day or night, changes the weather, or opens every way to reach me. It knows when it’s already night, and says so instead.
The server only proposes. The chat carries each action out from its own short list, checks the arguments again, finds its target by the label the page gives it (never a selector, a script or an address from the model), and leaves a receipt with an Undo where there is something to undo. On a phone, where the chat covers the page, it steps aside so you see what it did.
Commands still work misspelled (“sprng”, “drak mode”), mended by edit distance, and never fire on a question: “what’s summer like?” changes nothing. Asking it to tell me something drafts the note, and a part named by its label is scrolled to or ringed, all without a model. A test matrix of more than 100 ways of asking, from each state of the page, holds it to that.
Nothing it does sends anything on your behalf. A drafted note opens for you to read and send; a question for the real me reaches my phone only when you press Send; booking opens my calendar for you to pick a time. The model has no tool that sends, pays or writes.
| Kind | Tools | What they do |
|---|---|---|
| Look up (5) | search_site, get_project and three more | Go back to the model: another search, or one thing in full. |
| Show (18) | show_project, compare_projects, open_page, draft_note, ask_real_me, suggest and twelve more | Put one of the site’s components in the chat the moment the call is whole. |
| Act (8) | scroll_to, highlight, tour, navigate, set_time, set_season and two more | Do something on the page, carried out by the chat from its own list, with a receipt and an undo. |
components/ask/actions.ts
case 'navigate': { if (!/^\/(?!\/)[a-z0-9\-/]*$/.test(act.href)) return { ok: false, label: "Couldn't go there" }; hooks.stepAside(); hooks.go(act.href); return { ok: true, label: `Took you to ${act.label}` };}case 'time': { if (!isTimeChoice(act.time)) return { ok: false, label: "Couldn't change the time" }; const choice = getTimeChoice(); setTimeChoice(act.time); return { ok: true, label: `Switched to ${act.time}`, undo: { a: 'time', choice }, };}Keeping it honest
It is meant to be trusted by someone making a decision about me, so it would rather say less.
It speaks as me, but says plainly that it’s an AI version of me when asked, and never commits me to a date, a rate or a meeting. When the site doesn’t cover something it says “I haven’t written about that here, so I’d rather not guess” and offers a card that sends the question to the real me.
Grounding is layered. The brief forbids numbers, dates, employers, links and opinions that aren’t in what it was given; the facts and passages are its only sources; citations must point at a passage it read; the fit check’s links must be pages of the site, and a gap never gets one; and project cards come from content, so a project that isn’t there can’t be shown.
What it did is checked too. If an answer says it did something on the page that nothing did, a line after it sets that straight, and “nothing happened” is answered from what the last answer actually ran, with how to do it by hand.
The eval checks every number in every answer against the site’s own text, which is how one rounding of 12,493 downloads to 12,000 was caught.
Prompt injection, and what comes out
Prompt injection is handled by capability more than by wording.
The visitor’s message, the passages, tool results, the page’s parts and a pasted job description are all data. The brief says so, the job description is fenced and described as text to read, never instructions, and the page’s parts ride with the question as data rather than as part of the brief.
Obvious tries (“ignore your previous instructions”, “you are now”) and jobs it isn’t for (write me code, the weather) get a line back in my voice without reaching the model at all, in under 0.1 s.
The real defence is that there is nothing worth taking over: no tool sends, pays or writes, arguments are enums and patterns checked on the server and again in the browser, and page actions find things only by their visible labels.
What comes out is filtered as it streams. No HTML is ever drawn: answers are parsed into plain data (paragraphs, lists, bold, code and links) and drawn as elements. A link goes somewhere only if it is https or a page of this site, and there are no images, so an answer can’t carry data away in an image’s address. Long dashes become commas, and answers stop at 2,400 characters.
A tool call the model writes out as words instead of making it comes in many shapes: a name and its arguments, a name in bold, a bullet, tags, a JSON code fence, a bracketed marker. From where one could begin, the words are held until it is clear what they are; a call is dropped, its questions become follow-ups, and a card it asked for is shown for real.
lib/agent/leaks.ts
/* * A tool call the model wrote as words instead of * making it, in many shapes: `suggest: [...]`, * `show_project({...})`, `functions.show_project(...)`, * a bold or bulleted name, `<tool_call>` tags, a JSON * code fence, `[TOOL_CALLS]`. From where such a thing * could begin, the words are held until it is clear * what they are; a call is dropped, and what it asked * for is kept. */constructor(private readonly names: readonly string[]) { const alt = names .map((n) => n.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')) .join('|'); this.trigger = new RegExp( '<tool_call>|```|\\[TOOL_CALLS\\]|\\bfunctions\\.' + `|(?:\\*\\*)?\\b(?:${alt})\\b`, 'i', );}The guards, and why each value
A chat that calls a paid model from a public page is a bill anyone can run up, so every limit is counted in the store, in fixed windows, and everything fails closed: if the limits can’t be counted, the model isn’t asked.
A question has to come from the site’s own pages, from something that says it is a browser, and with a session token the chat fetched when it opened: an HMAC over when it was issued and the address it was issued to, good for two hours, compared in constant time and never stored.
The limits count a visitor under a hash of the address with a secret and the day, so no address is kept and today’s can’t be tied to yesterday’s. The token is sealed to the address itself, so a session doesn’t die at midnight when that hash changes, which it used to.
Answers in flight are a set of requests stamped with when they began, not a counter. A counter drifted up whenever a function was killed mid-answer and never counted itself out; a stamped entry simply ages out after 90 seconds.
| Guard | Value | Why |
|---|---|---|
| Origin | The site’s own pages | A question has to come from here. |
| Browser | No script or crawler user agents | Scrapers and curl stop before anything is counted. |
| Session | HMAC token, 2 hours, sealed to the address | A script that never opened the chat has none, and one visitor’s can’t be reused by another. |
| Size | 28,000 bytes a request; 400 characters a question; 4,000 a job description | The largest honest request, and not a byte more. |
| Burst | 4 in 10 s from one address | A held-down Enter or a loop, stopped without bothering a person. |
| Per visitor | 8 a minute, 30 an hour, 80 a day | More than any real conversation; one visitor can’t spend the day. |
| Everyone | 60 a minute, 1,500 a day | A ceiling on the bill, however many come. |
| In flight | 24 answers at once, each aging out after 90 s | Within the model’s per-minute limits and the function’s concurrency. |
| Tokens | 1,500,000 a day, prompt and answer; stopped turns charged too | Past it, the pages answer until tomorrow. |
| Breaker | 5 failed questions in 2 minutes, open 3 minutes, per provider | A failing model isn’t hammered, and one provider’s trouble doesn’t close the other. |
| Keys | A refused key rests 1 minute (rate limit) or 1 hour (spent, revoked) | The next is tried in the same request, before a word is said. |
| Turn | 3 requests, 20 s, 600 tokens an answer, a deadline, 8 s to the first byte | No question can run away with the budget. |
lib/agent/guard.ts
export const LIMITS = { burst: { count: 4, seconds: 10 }, minute: { count: 8, seconds: 60 }, hour: { count: 30, seconds: 60 * 60 }, day: { count: 80, seconds: 60 * 60 * 24 }, allMinute: { count: 60, seconds: 60 }, allDay: { count: 1500, seconds: 60 * 60 * 24 }, inFlight: 24,} as const;/** Tokens a day, prompt and answer together. */export const dailyTokens = () => Number(process.env.ASK_DAILY_TOKENS) || 1_500_000;const BREAK = { failures: 5, windowSeconds: 120, openSeconds: 180 };const SESSION_MS = 2 * 60 * 60 * 1000;/** A token for this address, good for two hours. */export function issueSession(address: string) { const issued = String(Date.now()); return `${issued}.${seal(issued, address)}`;}When it breaks
Everything upstream fails sometimes, so each failure has an answer a visitor can still use.
Five failed questions in two minutes open a provider’s breaker for three. A question counts once against a provider, however many of its keys were tried; counting every try used to let a single question with three tired keys trip it.
While both providers are resting, or the day’s tokens are spent, it answers from the pages alone: what the question obviously wants shown, and links to the three best passages, under a line saying so.
When an answer doesn’t come (it broke, it’s resting, the limits were hit, it didn’t know, or a guard stopped it), the question is handed to the real me: a note with the question and where it was asked already in, a name and a way back if they like, sent to my phone only on their press, once a question.
A stream cut off before its end is a failure with that note, not a half answer left looking finished. A health endpoint tells an uptime monitor whether each provider can be asked and whether the store answers, in counts and nothing secret.
What it keeps, and for how long
Your conversation lives on your device. It is kept in your browser for a week (the last 15 turns, and five earlier chats) so a reload or another page picks up where you were, and saved when the browser has a spare moment rather than while an answer streams.
On my side there are counts, and the questions. The counts (turns, tools, how long each part took, tokens) say nothing about who asked or what.
Each question is kept 30 days, in MongoDB with an index that deletes it on the day, with the first 240 characters of its answer, the page it was asked on, the tools that ran and a city and country as the edge placed them, so I can read what people ask and write the pages they’re missing. Email addresses and phone numbers are blacked out first. No address, no browser, no device, and the chat says so under its field.
| What | Where | How long |
|---|---|---|
| The conversation | Your browser | 7 days, 15 turns, 5 earlier chats |
| Counts: turns, tools, timings, tokens | The store | 35 days |
| A question, 240 characters of its answer, the page, a city | The store, redacted | 30 days, the newest 3,000 |
| A fresh question’s whole answer | The store, keyed by its words and page | 1 day |
| A thumbs up or down, with its question | The store, redacted | 30 days |
| A question for the real me, or a wish to talk | The store, and my phone | 30 days |
| Rate-limit counts | The store, under a daily hash of the address | 10 s to 1 day, per window |
| Anything that was said | Server logs | Never |
conversations.ts and guestbook/ticket.ts
/** Email addresses and phone numbers, blacked out. */export function redact(text: string) { return text .replace(/[\w.+-]+@[\w-]+(\.[\w-]+)+/g, '[email]') .replace(/(?:\+?\d[\s().-]*){8,}\d/g, '[phone]');}/* ... and who is asking, for the limits only: */export function visitorKey(ip: string) { const day = Math.floor(Date.now() / 86_400_000); return createHmac('sha256', KEY) .update(`${day}:${ip}`) .digest('base64url') .slice(0, 16);}Kept answers, keyed with care
An answer to a fresh question is kept a day, so the next person to ask it gets it in about 25 ms, replayed a few words at a time, for no tokens.
The key is the question as normalised, the page it was asked on, that page’s parts as the visitor’s page sent them, the time and weather showing, and a hash of the prompt and everything the agent knows. A change to the content or the prompt begins afresh.
Two mistakes shaped that key. Kept by the words alone, an answer from one page was replayed on another. And because a page’s parts come from the visitor, a made-up list could have planted an answer for everyone; with the parts in the key, a made-up list is only ever answered to itself.
An answer that acted on the page, or failed, is never kept. A kept “I’ve made it night” used to replay the words without the night. The vectors for retrieval are kept as well (30 days a passage, a day a question), under the embedding model’s name and the text, so neither is paid for twice.
Next: a semantic cache, so “what’s your stack” and “what do you build with” share an answer, behind the same interface the exact cache uses, and precomputed answers for the suggested questions (in progress).
Fast, and free until it’s reached for
The chat costs a page nothing until someone reaches for it. The tab is a button with my portrait; the chat itself, 25 KB compressed in two files, is fetched when a pointer nears the tab or it takes focus, so the press finds it ready. Once opened it stays mounted, and opening it again is instant.
While an answer streams, events are folded in once a frame however fast they arrive, so there is at most one render a frame; a finished answer is never drawn again while the next one streams; and the conversation is saved to the device when the browser is idle.
On the server, the passages and their index are built when the function boots rather than inside the first question, every count and check is one round trip, and the obvious card goes out before the model is asked. The questions offered to tap were answered once, ahead of time, and ship with the site: tapping one asks no model, and an answer is served only while the passages and brief it came from are unchanged.
What a visitor shouldn’t wait for is done when the site is built. A step before the build makes the passages’ vectors, only for the passages whose words changed (a hash of each says which), and keeps them in Atlas, removing the ones the site no longer has; a full index of 179 passages took 42 seconds, and an ordinary deploy embeds a handful. Without the database or the keys the step says so and the build goes on, and it can never fail a deploy. The chips’ answers are written the same way, by a script run between deploys, a few at a time within the free quota.
Measured on 1 October 2026 on a production build on my laptop, against the real hosted model on its free tier, with a stand-in store on the same machine: 20 questions visitors ask, each asked once, before meaning was added to retrieval (it adds about half a second to a new question). The trip to the host and to the store adds to these, and isn’t measured here.
| Step | Measured | Notes |
|---|---|---|
| Find passages by words | 1 to 11 ms | In memory; under 0.2 ms to score a question once warm. |
| Find passages by meaning | 0.52 s median | A new question; a repeated one is kept a day. |
| An obvious card | 36 ms median | Inside the turn; 14 of 20 questions had something on screen within 0.3 s. |
| One model request | 1.26 s median, 1.57 s at p90 | 30 requests; 1.5 a question. |
| First words | 1.85 s median, 2.3 s at p90 | In the client, through the route. |
| Whole answer | 2.15 s median, 2.85 s at p90 | In the client; 1.1 to 3.6 s inside the turn. |
| A kept answer | About 25 ms to the first word | The whole answer replayed in 0.1 to 0.3 s. |
| A command, off topic, or a try at its rules | Under 0.1 s | No model request. |
| A fit check | 2.5 s | One request, 3,525 tokens in, 464 out. |
| The chat’s own code | 25 KB compressed, 2 files | Fetched on the way to the tab, not with the page. |
components/ask/Ask.tsx and AskPanel.tsx
/* The chat itself is its own chunk, fetched as a hand nears the tab. */const loadPanel = () => import('./AskPanel').then((m) => m.AskPanel);const AskPanel = dynamic(loadPanel, { ssr: false });/* ... and while an answer streams: */// Events are folded in once a frame, however fast// they arrive: at most one render a frame while the// words stream.let queue: TurnEvent[] = [];let frame = 0;const flush = () => { frame = 0; const batch = queue; queue = []; if (batch.length) update((t) => batch.reduce(fold, t));};await ask(body, abort.signal, (event) => { queue.push(event); frame ||= requestAnimationFrame(flush);});Seeing what it does
Every answer carries its own trace, and anyone can open it under the answer: each part as a span (rewrite, retrieve, each model request, each tool), when anything first showed, the total, the tokens in and out, which tools ran, which provider answered and the prompt’s version. It is made of timings and names, never words.
The same trace is folded into counts a day: questions, kept answers, limited, failed, stopped, fallbacks, answers from the pages, questions it couldn’t answer, off-topic replies, fit checks and leads; each part’s latency in buckets (under 250 ms, 500 ms, 1 s, 1.5 s, 2.5 s, 5 s, and slower); tokens; each tool’s runs and bad calls; which provider answered; and the prompt’s version. They’re kept 35 days.
Logged: nothing that was said. The counts are numbers and the traces are timings. The questions live apart, redacted, for 30 days of review, and nowhere else.
lib/agent/ui.ts and telemetry.ts
/** The visible trace of a turn: what ran, how long, nothing that was said. */export type Trace = { version: string; model: 'primary' | 'fallback' | 'kept' | 'pages' | 'local'; /** Which provider answered. */ provider?: 'primary' | 'fallback'; steps: number; spans: { name: string; ms: number; ok: boolean }[]; firstMs: number; totalMs: number; tokens: { prompt: number; completion: number }; tools: string[];};/** Latency buckets, by upper bound in milliseconds. */export const BUCKETS = [250, 500, 1000, 1500, 2500, 5000, Infinity] as const;The review desk, and my phone
The questions are for reading, so they have a desk: /ask/review, opened only with my key, holding nothing without it.
It shows the conversations of the last 30 days, newest first, filterable by a lead, a fit check, a page or a thumbs down; the people who want me to get back to them, with a reply link; the questions it couldn’t answer, which are the pages I should write next; and the numbers: questions a day, what was kept, fit checks, leads, how fast the first words came, which tools ran and the tokens spent against the day’s budget.
My phone hears the rest. A fit check the moment someone runs one (the role, how strong a fit, the verdict and where they are, linked to that conversation on the desk), once a conversation and role, thirty a day at most. Trouble, each once and a dozen a day at most: tokens past 80% or spent; every provider failing a question, urgently, and the chat back again; every key of a provider resting; the main provider failing while a backup takes the questions, naming which.
And one push each evening at 8 pm, asked for by a scheduled workflow with its own key: the day’s conversations, fit checks, leads, unanswered and most-asked questions, thumbs down, résumé views, visitors, and how the budget and the providers held up, or a single line on a quiet day.
Evals, and how they gate a change
A prompt change is a code change, so it is tested like one. The prompt has a version (agent-12 today) that goes into every trace, splits the counts, and is part of the key kept answers are stored under, so a new prompt never serves an old answer and its numbers can be set beside the last one’s.
284 unit tests cover the engine with no model at all: the stream parser cut at every point, tool calls assembled across chunks and providers, the loop’s step cap and fallbacks, key rotation, the schema checks, markers and citations, every leak shape, the guards, the route’s refusals, the fit check’s validation, and retrieval, which fails CI if hit@5 on the eval set drops under 90%. The cases from an audit of the engine are among them.
The eval set is 44 questions a visitor might ask, each with where the answer is on the site and words a right answer has to say, plus follow-ups, persona checks (does it say it’s an AI; does it hand over rather than promise) and tries at its rules. A script runs them against the real model and scores retrieval, tool choice, the answer’s facts, invented numbers, first person and persona. It costs tokens, so it runs by hand before a prompt change ships.
The last run, on 1 October 2026: the tool a case expects for 21 of 23; the words a right answer needs in 24 of 26; no invented number in 43 of 44 answers (the one rounded 12,493 downloads to 12,000); first person in 40 of 44 (four led with a project rather than with I); and all 6 persona checks.
Retrieval’s one miss is “Does he know AI?”, which now finds this page first: a fair answer the set doesn’t expect yet.
The actions have their own matrix: more than 100 ways of asking, from each state of the page, each command against what should run and each look-alike question against nothing running, so an action firing on a question is a failing case, not a surprise.
The providers are checked for real before a deploy, by hand. One script asks every key once for a token and says which is revoked, suspended or out of quota, by its provider, its place in the list and its last four characters. A fire drill breaks the main provider’s keys, then the second’s, then every one, and watches the next provider answer, a refused key rest and not be asked again, and the end of the line fail in under a second. Its first run found two keys suspended, a backup whose model was no longer free, a bad key answered with a 400 that rested nothing and stopped the turn, and a stalled backup that held a question for 25 seconds; each is fixed, with a test.
What I’ll add next: a model as judge of faithfulness (does each sentence follow from the passage it cites), checked against a small set I score by hand before it is trusted with a large one; regression cases drawn from the review desk, where every question it couldn’t answer or that got a thumbs down becomes a case; each tool’s success rate from the counts; a latency objective (a card within 0.3 s and the whole answer within 3 s at p90) checked on every deploy; and the evals in CI on a small model, gating any change to the brief.
| Check | Result | How |
|---|---|---|
| Retrieval, hit@1 | 26 of 32 | The expected page is the first passage (words and meaning). |
| Retrieval, hit@5 | 31 of 32 | In the top five, the offline gate. |
| Retrieval, MRR@5 | 0.87 | Mean reciprocal rank of the expected page. |
| Tool choice | 21 of 23 | Ran a tool the case expects; off-topic cases answered without the model. |
| Answer checks | 24 of 26 | Says the words a right answer must. |
| No invented numbers | 43 of 44 | Every number is one the site says. |
| First person | 40 of 44 | Speaks as me, not about me. |
| Persona | 6 of 6 | Says it’s an AI; hands over instead of promising. |
scripts/agent/eval.ts · invented numbers
const said = ( knowledge() + '\n' + corpus().map((c) => c.text).join('\n')).toLowerCase();/** Numbers in an answer that the site never says. */const invented = (text: string, also = '') => ( text .replace(/\[\[source:\d+\]\]/g, '') .match(/\d[\d,.]*\d%?|\d%?/g) ?? [] ).filter( (n) => !said.includes(n.toLowerCase()) && !also.includes(n.replace(/,/g, '')), );What broke, and how it was fixed
Most of what makes it dependable came from something going wrong. Each of these is now a test.
Memory starved by the prompt budget: facts and passages were fitted first and the conversation got what was left, so follow-ups forgot. Now the last four turns come first.
Actions claimed but not run: the model said “I’ve switched to night” without calling the tool. Now a claim is checked against what ran, and corrected in a line.
Intent patterns firing on questions: summer changed the weather, a day job made it day. Now an action needs a command, and a card before the model needs an explicit ask.
A kept answer poisoned by the page’s parts, and one that replayed an action’s words without the action. Now the parts and state are in the key, and answers that act are never kept.
The breaker counting tries rather than questions, so one question with tired keys could trip it; the in-flight counter drifting when a function died; tokens not counted for stopped turns; sessions dying at midnight. Each is fixed above, where it lives.
Tool calls leaking into the words, first as “suggest: [...]” lines, then as bold names, tags, fences and bracketed markers. The filter now holds every shape from where it could begin.
And on a phone: the notes of what an answer showed and did weren’t going with the next question, so a follow-up about what it had just done had nothing to go on; a stream cut short looked finished; a conversation kept in an old shape broke the chat on the next visit; and the chat followed the words even when the reader had scrolled up. All four are fixed, and the chat now takes the whole height between the phone’s safe areas.
Trade-offs I considered and rejected
An agent framework, or a vendor’s SDK. Either would have named a vendor in the code, added weight, and put the stream’s shape in someone else’s hands. The loop is a few hundred lines, and owning it is what lets a card arrive mid-answer and a second provider take over mid-turn.
A vector database. The whole corpus is about 150 passages; words in memory score one in under a fifth of a millisecond, and vectors live in the same store as everything else. A database becomes worth it at thousands of pages or many sites, which is the module, not this.
A model to rewrite follow-ups into standalone questions. It would cost a request before every follow-up; borrowing the last question’s words and projects costs nothing and is right often enough, with the model’s own memory of the turns behind it.
Server-sent events to the browser. The question is a POST with a session header, which EventSource can’t send; NDJSON over a streamed fetch keeps the request honest and the parsing one line at a time.
Answers as JSON. Streaming words with cards as tool calls lets text start at once; JSON is kept for the one place its shape matters, the fit check.
Many agents. See routing, below: not until the evals say one agent has stopped being enough.
Running on free quotas, gracefully
It runs on free tiers today, so the strategy is to spend tokens only where a model adds something, and to degrade in steps a visitor can still use.
Zero tokens first. Commands (“make it night”), off-topic questions, tries at its rules and kept answers never reach a model. Precomputed answers for the suggested questions are in progress.
Then fewer tokens. 17 tools on offer rather than 31; the second provider offered only the tools most answers need; the facts written once, compactly; passages capped at 3,600 characters.
Then limits. Six windows of rate limits and a daily token budget of 1,500,000, set against the free allowance, not the paid one.
Then another quota. Keys in turn cover a failing key, not a bigger allowance, since keys in one project share one quota. The second provider has a quota of its own, and the order of who is asked is fixed and ordered.
Then no model. Past every limit, or with both breakers open, the pages answer: the obvious card, links to the best passages, and a note to the real me.
The capacity maths is short. An answered question averaged about 8,000 tokens over 1.5 requests, and more now that memory comes first. Back to back, about 13 questions a minute, the first provider refused after eight; spaced to five a minute, 25 in a row went through. So a free tier carries a few questions a minute, and at the budget’s 1,500,000 tokens, about 185 answered questions a day, before kept answers and commands.
On a paid tier the budget, not the provider, becomes the limit, and that is a choice: at an assumed $0.10 a million tokens in and $0.40 out, 1,000 conversations of three questions cost about $2.50.
What breaks first at scale, and the plan
Beyond a portfolio’s traffic, here is what gives first and what I’d do about each.
The model’s tokens a minute go first; the fix is a paid tier on the first provider, with the second on its own quota taking over when the first can’t answer.
Then the store: an answered question is 46 to 53 store commands (the limits alone are 12), so a plan is sized in commands, not memory. One script per question that counts every window and returns the verdict, and one write of the day’s counts per turn, would cut that by more than half.
Concurrency comes last: 24 answers in flight at about 2 seconds each is several hundred a minute, far above the 60 a minute allowed, while cold starts matter on a quiet site, where the index is built as a function boots.
Under real load: rate limits and a firewall rule at the edge for the chat’s route, so a flood never wakes a function; the semantic cache; the prompt’s fixed part first and the per-question parts after it, so a provider’s prefix cache can reuse the roughly 4,000 tokens every question shares; and a short queue with a place in line instead of a refusal, and past it, answers from the pages alone.
Cost is tokens, and tokens are measured. A conversation of three questions is about 24,000 tokens in and 320 out, before the history a follow-up carries. At an assumed $0.10 a million in and $0.40 out, the list price of a small hosted model, that is about $2.50 per 1,000 conversations; at $0.30 and $2.50, about $8.
| What | Number | How it was found |
|---|---|---|
| Tokens an answered question | 8,044 in, 107 out | Average of 20 live questions, before memory came first. |
| Model requests a question | 1.5 | Same 20; one or two each. |
| Store commands an answered question | 46 to 53 | Counted on a stand-in store. |
| Answered questions a day on the budget | About 185 | 1,500,000 ÷ about 8,150. |
| Questions a day, everyone | 1,500 | The guard; kept answers count, but cost no tokens. |
| Answers in flight | 24 | About 11 a second at 2.15 s each; the 60-a-minute limit binds long before. |
| Cost per 1,000 conversations | About $2.50 to $8 | Three questions each, at the two assumed prices above. |
What’s next: routing
Today a question is routed in two layers: a deterministic layer of patterns that answers commands and off-topic questions with no model and plans the obvious cards, then one model with a subset of tools chosen for that question. It is cheap, fast and easy to test, and it is the part most likely to be outgrown.
A learned router. An embedding-similarity or small-classifier router choosing among five paths: the command path, a precomputed answer, retrieval-only answers, the full agent, and the fit-check pipeline. It would catch the phrasings the patterns miss, at the cost of a vector lookup or a small model call on every question.
Model routing. A small, fast model for the simple questions and a larger one for a fit check or a long job description, chosen by expected cost and latency. The second provider already shows the shape of it: a different model, fewer tools.
Many agents: a planner with specialist workers (retrieval QA, a fit analyst, a site operator) and a verifier over them. Deliberately not used yet: it adds latency to every answer, more calls against free quotas, more ways to fail, and evals that are much harder to trust. What would justify it is the evals showing a single agent plateauing on multi-step tasks, which is most likely when the voice agent starts operating the site for someone.
However it is routed, it is judged the same way: precision and recall per route on real questions from the review desk, labelled by hand; each route’s place on a latency and cost Pareto curve; and a regression gate, so no routing change ships that sends a known question down a worse path.
Voice, next
Next, it speaks. A pill grows out of the tab with a soft orb in it that moves with the real audio, captions of both sides above it, the same suggested questions, and a way to open it into the text chat with the conversation carried over.
Voice is a transport, not a second brain. A realtime speech-to-speech model runs over a socket the browser opens itself, with a short-lived, single-use token my server mints and locks: the brief, the tools and the voice are sealed into it, so a modified page can’t change them. Its tool calls run on my server through the same registry, with the same session, guards and events, folded into the same conversation. The voice is a stock one, never a clone, and the pill says it’s an AI.
A spike put numbers on it: first audio 1.2 to 1.5 s after the visitor stops talking, and a site action visible 1.1 to 1.9 s after asking for it. It also showed what not to do: with the text brief and all 29 tools, the first audio took 2.8 to 5.5 s, because the model waited for its own card; a brief for speech (answer first, out loud) and 16 tools brought it back to about 1.5 s.
The guards are the text chat’s plus minutes: three minutes a session, a few sessions a visitor a day, a cap on sessions at once under the provider’s (the tier I tested took six and refused the seventh), a daily budget of minutes, and a breaker. When voice is full it falls back to the browser’s own speech recognition with the same text turn read aloud, then to text. Anything that leaves the page waits for a spoken or tapped yes.
As a module
The goal is for any site to have one that knows its own pages, shows its own components and moves around it for its visitors. 10xAnswers, my chat component on npm, was the first step; this will ship as a package of its own.
Most of it is already shaped for that. The core yields events and knows nothing of HTTP; the model is an address in the environment; tools are entries in a registry; the store is a handful of commands behind one file; the chat draws components by name from props; and anything on a page can open it with a question, through one event.
What has to be made general first: the corpus, which reads this site’s markdown today and needs a content adapter (a folder of markdown, a sitemap to crawl, a CMS); the persona, since the brief is written as me; the intent patterns, which know my projects; the store, which speaks one provider’s REST API; the page actions, which rely on this site’s labels, night and weather; and the chat’s styles. The sketch below is where it’s heading, not a published API.
Planned · not published
// Planned: the shape it's heading for, not a// published API.import { createAgent, ndjson } from './agent';const agent = createAgent({ persona: { name: 'Ada', away: true }, content: markdownFolder('./content'), tools: [...defaultTools, showProduct, bookDemo], store: redisStore(process.env.REDIS_URL), model: fromEnv('ASK_MODEL'), limits: { visitorPerDay: 80, tokensPerDay: 1_500_000 },});// The chat's route: one JSON event a line.export const POST = ndjson(agent);// A voice transport over the same brain, later.export const voice = realtime(agent, { secondsCap: 180 });// The chat, drawing your own components by name.<Assistant endpoint='/api/ask' blocks={{ product: ProductCard }} />What I learned
Most of the speed comes from not asking the model. A pattern that knows what a question obviously wants puts it on screen in milliseconds, and the model writes around it.
A small model is unreliable in small ways: a tool call written out as words, a citation as [2], a long dash, a claim of something it didn’t do. So what it writes is a stream to be parsed and checked, not text to be trusted, and each of those has a test.
The limits are the product. An assistant on a public page is only as good as what happens when it’s abused, rate-limited, out of tokens or talking to a broken endpoint; here each of those has an answer a visitor can still use.
Next. Precomputed answers for the suggested questions, the action eval matrix, voice on the same brain, a judge in the evals, routing when the evals ask for it, and the module.
