# Rajveer Singh > Rajveer Singh (rajveeerr on GitHub, X and LinkedIn) is a software engineer in New Delhi, India, who builds AI products end to end: the interface, the backend and the infrastructure that keeps them running, across full-stack, product and design engineering. His work includes live voice rooms and Stripe billing for IABTM, a US self-improvement platform, where he is a software engineer; HyperPersona, an agentic recommender on AWS Bedrock that his team took to the top 0.1% of 22,000+ at Cognizant Technoverse 2.0; Safire, a harassment shield for LinkedIn DMs that took first place at HackWIE 3.0 and Code Kshetra 2.0; 10xAnswers, a React AI chat component on npm; most of the web app for Wii Hop, a US local-deals startup; and an AI version of himself on his site that answers for his work, shows it and checks a job description against it. He is a 7× hackathon winner or finalist. I build the whole product: the interface people touch, the backend and realtime underneath, and the infrastructure that keeps it up. I’ve shipped live voice rooms and billing at IABTM, an AI engine that placed in the top 0.1% of Cognizant’s Technoverse hackathon, a harassment shield that took first place at two hackathons, an AI chat on npm that strangers install, and the lake at the top of the home page, which you can fish. Right now I’m building a voice agent that speaks for my work. Shipping since August 2024. ## Now - Building: [Talk to my agent](https://rajveers.com/projects/ai-me). An AI version of me for this site. It answers for my work from the pages here, with sources, shows the real projects, checks a job description against what I have built, and acts on the page when asked. Next, it speaks. - Learning: Retrieval and evals. Hybrid search, reranking, and a set of questions scored on every change. - Learning: LLM tracing. Latency, tokens and cost for every stage of a turn. - Learning: React Native. Two small apps. They show up here once they’re on a phone. ## Experience - **IABTM**, Software Engineer (internship), Sep 2025 to now. Building the core of a US self-improvement platform. Rebuilt its live voice rooms on an SFU, rebuilt its Stripe billing so no one is charged twice, and built the daily AI pipeline (Gemini) that curates what each member reads. Also built the email and the team’s admin tools, and runs the AWS. - [Inside the voice rooms and billing](https://rajveers.com/projects/iabtm) - **Wii Hop**, Software Engineer (contract), Jul 2025 to May 2026. Built most of the web app for a US local-deals startup: deals on a live map, table booking, events, loyalty streaks, and heists that turn a night out into a game (React, TypeScript, Leaflet). On the backend, wrote the parts where money and stock change hands, from the heist item shop to catering orders priced on the server. Also built screens for the Expo app. - [Inside the deals, bookings and loyalty](https://rajveers.com/projects/wiihop) - **Modulus**, Frontend Engineer (internship), Apr 2025 to Aug 2025. Built the product’s interface from the first screen: cohort pages, booking and payment on Razorpay, scheduling through Cal.com, and live sessions with chat running beside the stream (React, TypeScript, hls.js). Wired PostHog funnels so the team could see where people dropped off before they enrolled. - [Inside booking and live sessions](https://rajveers.com/projects/modulus) ## Projects - [HyperPersona](https://rajveers.com/projects/hyperpersona) (hackathon, 2026): An agentic recommender on AWS Bedrock that four of us built in a week: top 0.1% of 22,000+ at Cognizant Technoverse. I built the storefront and event SDK. - [Safire](https://rajveers.com/projects/safire) (hackathon, 2025): A harassment shield for LinkedIn DMs that won HackWIE 3.0 and Code Kshetra 2.0. I built the backend: the API, moderation failover and evidence reports. - [10xAnswers](https://rajveers.com/projects/10xanswers) (side project, 2024 to 2025): A drop-in React chat widget for AI answers, built and published alone: 12,493 downloads across 55 releases on npm. - [Kernel](https://rajveers.com/projects/kernel) (take-home, 2025): Anonymous chat rooms with peer-to-peer voice, where the server only relays the handshake. The take-home for the role I have now. - [The lake](https://rajveers.com/projects/lake) (side project, 2026): This site's hero: a lake painted in code, with water simulated on the GPU, creatures that make ripples, and fishing with your phone as the rod. - [The AI me](https://rajveers.com/projects/ai-me) (side project, 2026, in progress): An AI version of me that answers for my work while I’m away: it reads this site, replies in my voice with its sources, shows the real projects instead of describing them, checks a job description against what I’ve built, and moves around the site when you ask. ## Track record ### Hackathons - [HackWIE 3.0](https://dorahacks.io/hackathon/826/winner), 1st place, IEEE MSIT · 2025 - [Code Kshetra 2.0](https://x.com/RajveeerrSingh/status/1894034885855138202), 1st place, out of 15,000+ registrations, Geek Room · 2025 - [Technoverse 2.0](/projects/hyperpersona), Top 0.1% of 22,000+, Cognizant · 2026 - UPSTART, Second runner-up, E-Summit ’25, E-Cell MSIT · 2025 - SheBuilds × OpsTree, Award, 2025 - MSC Hack-It-Up, Special mention, IGDTUW · 2025 - Excalibur, Techspardha, Final round, NIT Kurukshetra Winner or finalist: 7× ### Open source - [10xanswers](https://www.npmjs.com/package/10xanswers), An AI chat for React sites, 55 releases, npm - [axios-docs #217](https://github.com/axios/axios-docs/pull/217), A fix in the official docs, merged, axios - [100xDevs CMS](https://github.com/code100x/cms/pulls?q=is%3Apr+author%3Arajveeerr+is%3Amerged), 4 pull requests, merged, code100x Downloads on npm: 12,493 ### Stack - Interface, React, Next.js, TypeScript, Tailwind, Expo, ×5 - Server, Node, Express, Socket.IO, Prisma, ×4 - Realtime, WebRTC, mediasoup, hls.js, ×3 - Data, MongoDB, Redis, Firebase, ×3 - Money and AI, Stripe, Razorpay, Shopify, Gemini, ×4 - Cloud and graphics, AWS, Cloudflare Workers, WebGL2, GLSL, ×4 Tools in production: 23 ## Contact - Email: rajveergreets@gmail.com - Book a call: https://cal.com/rajveer - Résumé: https://rajveers.com/resume.pdf - Elsewhere: [github](https://github.com/rajveeerr), [x](https://x.com/rajveeerrsingh), [linkedin](https://www.linkedin.com/in/rajveeerr), [peerlist](https://peerlist.io/rajveeerr), [npm](https://www.npmjs.com/~rajveeerr), [devfolio](https://devfolio.co/@rajveeerr), [producthunt](https://www.producthunt.com/@rajveeerr), [holopin](https://holopin.io/@rajveeerr), [wakatime](https://wakatime.com/@rajveeerr) - Open to conversations. --- # About Rajveer Singh > Rajveer Singh, a software engineer who builds AI products end to end. How he got here, how he works, and what he does off-screen. I’m a software engineer in New Delhi. What I’ve shipped is on the projects page; this one is the rest: how I got here, how I work, and what I do when I’m not working. I do my best work on small teams building AI products, where an engineer carries a feature from the first screen to the server it runs on, then watches it run. I also like working fast: in one week, our team took an AI engine from a plan to the top 0.1% of 22,000. ## Story I started in July 2024 with Harkirat Singh’s 100xDevs cohort and shipped my first project two weeks later. The first thing strangers used was 10xanswers, an AI chat component I published to npm that December. In 2025, Safire, a harassment shield I built with friends, won HackWIE 3.0 and Code Kshetra 2.0. Then I built Modulus’s product UI and most of Wii Hop’s frontend, and a take-home called Kernel led to my job at IABTM. In 2026, the same team built HyperPersona for Cognizant’s Technoverse hackathon. Before all that, I was learning. I wrote Python on a phone through lockdown because I didn’t have a laptop. The one thing from then that mattered was a bill calculator for my father, who logged our electricity meter in a paper diary. It ran on 22 months of his readings, and it was the first time something I wrote did a real job for someone. I like the two ends of the stack most: the interface, where every detail is a decision, and the infra, where things either stay up or they don’t. The first end has been pulling harder lately: I’ve been exploring design engineering, designing in code down to the motion and the sound, and this site is where I practise it. AI goes in the middle, treated like any dependency that can fail. ## Timeline - Jul 2024: Joined the 100xDevs cohort - Dec 2024: [Published 10xanswers to npm](https://rajveers.com/projects/10xanswers) - Jan 2025: [First place at HackWIE 3.0](https://rajveers.com/projects/safire) - Feb 2025: [First place at Code Kshetra 2.0](https://rajveers.com/projects/safire) - Apr 2025: [Frontend engineer at Modulus](https://rajveers.com/projects/modulus) - Jul 2025: [Software engineer at Wii Hop](https://rajveers.com/projects/wiihop) - Sep 2025: [Software engineer at IABTM](https://rajveers.com/projects/iabtm) - May 2026: [Top 0.1% of 22,000+ at Cognizant Technoverse 2.0](https://rajveers.com/projects/hyperpersona) ## What I build - **Interfaces.** Fast, accessible front ends in React and Next.js, with keyboard paths, focus handling and reduced motion built in from the start. I did most of the accessibility work in IABTM’s app. - **Backends and data.** APIs in Node and Express over MongoDB, Postgres and Redis, with indexes checked against real query plans, pages that scroll by cursor, and webhooks that verify their sender and never count an event twice. - **Realtime.** Voice rooms on a mediasoup SFU, chat, presence and live notifications over WebSockets, and WebRTC calls where the server only makes the introduction. - **AI in the product.** Models treated like any dependency that can fail: they choose from a list rather than invent, fall back to plain code, stop at the first rate limit, and never write a member’s words into telemetry. - **Scale.** Ready for more servers before a product needs them: caches and sockets shared through Redis, scheduled jobs that claim their turn so each runs once, and duplicates stopped by the database, not by memory. - **Infrastructure.** AWS from the box up: EC2 behind nginx and PM2, media on S3 behind CloudFront, uploads of up to 2 GB sent straight from the browser to S3 in parallel parts that retry on their own, and a release tag to roll back to. - **Speed.** Pages that answer before their heaviest part arrives: work moved off the main thread, feeds served from a cache, code and images sent only where they show. The lake on this site paints in a worker and costs nothing off screen. - **Search and sharing.** SEO as a system: one metadata builder behind every route, structured data, a share image drawn for each page, and a sitemap that leaves things out on purpose. This site scores 100 for SEO in Lighthouse. - **Keeping it running.** Uptime checks, error tracking, model telemetry and alerts that reach the team within the hour, and failures fixed as a class, so the same one cannot happen quietly again. ## What I bring Everyone has the same AI tools now, and I use them every day. What they can’t give you is knowing what’s worth making, seeing when it isn’t good yet, and caring enough to finish it. That part is mine. I see the thing before it exists, and I keep going until the real one matches it. The lake on the home page is a live simulation running on your GPU, and it still lets the page load first. ## How I work - **Idea to production.** Design it, build it, ship it, then watch it run and keep it up. - **Read before changing.** I read a codebase before I touch it, and write down what I find: audits, runbooks, decision records. - **Leave a way back.** Release tags, kill switches, and migrations that start as dry runs. - **Fix the class of bug.** Not the instance. If something failed silently once, it shouldn’t be able to again. ## Off-screen - **Chess.** Chess, when I’m away from the screen. Same habit as debugging: read the whole board before you move. - **Matiks.** Timed mental-maths duels on Matiks, against strangers. No calculator, and no hiding. - **Investing.** I invest a little. Compounding is the whole idea, in code and out of it. - **Sherlock Holmes.** The books, the films and the series. For the method, not the mystery: notice what everyone skips, and stay calm on the clock. ## Contact - Email: rajveergreets@gmail.com - Book a call: https://cal.com/rajveer - Résumé: https://rajveers.com/resume.pdf - Elsewhere: [github](https://github.com/rajveeerr), [x](https://x.com/rajveeerrsingh), [linkedin](https://www.linkedin.com/in/rajveeerr), [peerlist](https://peerlist.io/rajveeerr), [npm](https://www.npmjs.com/~rajveeerr), [devfolio](https://devfolio.co/@rajveeerr), [producthunt](https://www.producthunt.com/@rajveeerr), [holopin](https://holopin.io/@rajveeerr), [wakatime](https://wakatime.com/@rajveeerr) - Open to conversations. --- # Projects > Projects by Rajveer Singh: live voice rooms, Stripe billing, an agentic recommender on AWS Bedrock, a harassment shield, an npm package and an AI agent for this site. ## Internships and contracts - [IABTM](https://rajveers.com/projects/iabtm) (internship, 2025 to now): The core of a US self-improvement platform: live voice rooms, Stripe subscriptions, a daily AI pipeline and the AWS it runs on. - [Wii Hop](https://rajveers.com/projects/wiihop) (contract, 2025 to 2026): Most of the web app for a US local-deals startup, from deals on a live map to heists that make a night out a game, plus the backend where money moves. - [Modulus](https://rajveers.com/projects/modulus) (internship, 2025): Interview prep for consulting, product and finance roles. I built its product from the first screen: programs, booking, payment and live sessions. ## Projects - [HyperPersona](https://rajveers.com/projects/hyperpersona) (hackathon, 2026): An agentic recommender on AWS Bedrock that four of us built in a week: top 0.1% of 22,000+ at Cognizant Technoverse. I built the storefront and event SDK. - [Safire](https://rajveers.com/projects/safire) (hackathon, 2025): A harassment shield for LinkedIn DMs that won HackWIE 3.0 and Code Kshetra 2.0. I built the backend: the API, moderation failover and evidence reports. - [10xAnswers](https://rajveers.com/projects/10xanswers) (side project, 2024 to 2025): A drop-in React chat widget for AI answers, built and published alone: 12,493 downloads across 55 releases on npm. - [Kernel](https://rajveers.com/projects/kernel) (take-home, 2025): Anonymous chat rooms with peer-to-peer voice, where the server only relays the handshake. The take-home for the role I have now. - [The lake](https://rajveers.com/projects/lake) (side project, 2026): This site's hero: a lake painted in code, with water simulated on the GPU, creatures that make ripples, and fishing with your phone as the rod. - [The AI me](https://rajveers.com/projects/ai-me) (side project, 2026, in progress): An AI version of me that answers for my work while I’m away: it reads this site, replies in my voice with its sources, shows the real projects instead of describing them, checks a job description against what I’ve built, and moves around the site when you ask. ## Smaller things - [destroyShorts](https://github.com/rajveeerr/destroyShorts) (2025): A Chrome extension that clears Shorts from your YouTube history by driving YouTube's own menus, and picks up where it left off after a reload. --- # IABTM > The core of a US self-improvement platform: live voice rooms, Stripe subscriptions, a daily AI pipeline and the AWS it runs on. - Kind: internship - Role: Software Engineer, internship, remote - When: Sep 2025 to now - Stack: Next.js, React, Express, MongoDB, Redis, Socket.IO, mediasoup, Stripe, Shopify, Gemini, PostHog, AWS - Live: https://iambetterthanme.com - iambetterthanme.com: https://iambetterthanme.com IABTM is a US self-improvement platform: personalised growth paths, curated film, music and writing, a social feed, live voice stages, paid tiers and two Shopify stores. I joined in September 2025 and picked up a 55,000-line codebase. Since then I’ve rebuilt its voice rooms and its billing, built the AI pipeline, the email platform and the tools the team runs the product from, and I look after the servers it runs on. ## In numbers - 6: publishers, curated daily - 7: jobs that run themselves - 25: admin tools for the team ## What I’ve shipped - **Live Stripe billing.** Three tiers and four plans, rebuilt after I found four holes in the old one. Prices are set on the server, and a payment event that arrives twice only counts once. - **Voice rooms on an SFU.** Live audio rooms where five members hold the stage and the rest raise a hand, rebuilt from a peer-to-peer mesh onto mediasoup. - **A daily AI pipeline.** Every morning, new articles from six major publishers land in the right members’ growth paths. Gemini picks the fit from the platform’s own list, never its own words. - **Seven jobs that run themselves.** Email campaigns, daily nudges in each member’s own time zone, store syncs and health digests, each run once on schedule however many servers are up. - **Premium, enforced once.** Free members can browse the whole catalogue and get a taste of every premium section. One check on the server decides what plays. - **Failures that make noise.** If a new member’s growth path fails to build, the account is marked and the team hears about it, instead of signup quietly saying yes. - **Two Shopify stores.** Two stores behind one headless backend. Orders arrive verified, a repeat can’t duplicate one, and new products publish themselves overnight. - **Email end to end.** Sign-in codes, billing, lifecycle and campaign email moved onto Resend. Every send is logged, bounces are kept off the list, and failures reach the team within the hour. - **The tools the team works in.** 25 admin screens: a support desk, a campaign CRM, analytics and A/B panels, and controls for subscriptions and the shops. - **SEO as a system.** Every page gets its own title, share image and structured data from one builder, across about 30 layouts. - **Monitoring and alerts.** Uptime checks, server errors and model telemetry that never records what members wrote, plus a daily health digest. - **The server it runs on.** EC2 with nginx and PM2, media on S3 behind CloudFront, and a release tag to roll back to for every launch. - **Uploads straight to S3.** Posts and banners of up to 2 GB go from the browser to S3 in parallel parts that retry on their own, so a dropped connection costs one part, not the whole upload. - **A feed that stays fast.** One aggregation, cursor pagination checked against the query plan, and a short Redis cache, so page fifty costs the database what page one does. - **Ready for more servers.** Sockets and caches shared through Redis, jobs that claim their turn, and duplicates stopped in the database, so the web side can run on several servers. - **Accessible by default.** I built most of the app's accessibility: keyboard paths, focus handling, screen-reader labels and reduced motion. ## Billing that counts every payment once The Stripe integration I inherited let the browser choose the price, let anyone cancel anyone’s subscription, never actually switched a paying member on, and handled every retried event all over again. I rebuilt it around three rules: prices live on the server, one service decides who has access, and every Stripe event is written down before it is acted on, so a retry can’t do anything twice. It went live in August 2026 with existing members carried over, covered by unit, end-to-end and live-Stripe tests. ```js // billingWebhookController.js · handleStripeWebhook try { await StripeWebhookEvent.create({ eventId: event.id, type: event.type, }); } catch (error) { if (error?.code === 11000) { return res.status(200).send('Already processed'); } return res.status(500).send('Could not record event'); } try { await dispatch(event); await StripeWebhookEvent.updateOne( { eventId: event.id }, { $set: { completed: true } }, ); return res.status(200).send('OK'); } catch (error) { // Deliberately a 200: a 5xx makes Stripe retry a // request that will fail the same way. The row // above records the failure so it can be replayed. await StripeWebhookEvent.updateOne( { eventId: event.id }, { $set: { completed: false, error: error.message } }, ); return res.status(200).send('Recorded with error'); } ``` Insert first, work second. A duplicate key is not an error here; it is the answer. ## Seven jobs, one run each Campaign emails, daily nudges, store syncs and the morning AI pipeline all run on a schedule. When I arrived, that schedule had never been installed on the server at all, and the app is built to run as several processes, where a plain timer fires once in each and emails every member several times. So every job now claims its slot in the database before it runs. The first claim wins and the rest skip, and if the database can’t answer, the job skips rather than risk sending twice. ```js // helpers/runWithLease.js export const acquireRunLease = async (name, bucket) => { try { await CronLock.create({ key: `${name}:${bucket}` }); return true; } catch (err) { // another worker won this window if (err?.code === 11000) return false; throw err; } }; export const runWithLease = async (name, bucket, fn) => { let owns = false; try { owns = await acquireRunLease(name, bucket); } catch (err) { // Fail closed: skip this window rather than // risk a double send. The job fires again on // its next schedule. return; } if (!owns) return; // another worker has it await fn(); }; ``` Every scheduled job goes through this, with a kill switch of its own. ## The model only chooses from a list Every morning, new articles from six publishers, among them The Atlantic, Wired and MIT Technology Review, go to the members they suit. Gemini reads each one and picks who it is for from the platform’s own list of growth attributes, up to three of each kind. Anything it makes up is thrown away. If the model answers with nonsense, answers with nothing or isn’t there at all, a plain keyword tagger does the job instead. Articles it has already seen are skipped before a single model call is paid for. ```js // scripts/automated-editorial-pipeline.js function normalizeAndValidateClassification( parsed, allowedCurrent, allowedImagine, ) { const cur = Array.isArray(parsed?.currentSelf) ? parsed.currentSelf : []; const img = Array.isArray(parsed?.imagineSelf) ? parsed.imagineSelf : []; const currentSelf = uniq( cur.map(normalizeAttributeTitle) .filter(a => allowedCurrent.has(a)), ).slice(0, 3); const imagineSelf = uniq( img.map(normalizeAttributeTitle) .filter(a => allowedImagine.has(a)), ).slice(0, 3); return { currentSelf, imagineSelf, via: 'gemini' }; } // In classifyArticle, on any model error: if (looksLikeGeminiQuotaError(err)) { // Rate limited: off for the rest of this run. geminiDisabledForRun = true; } return keywordFallbackTagger({ title, description, currentSelfAttrs, imagineSelfAttrs, }); ``` What the model says is filtered against the list, deduplicated and capped; the first rate limit turns it off for the run. ## Voice rooms on an SFU Members meet in live audio rooms and on stages, where five people speak and everyone else listens and raises a hand. The rooms were a peer-to-peer mesh, where every speaker sends their audio to every listener, which falls apart past a handful of people. I rebuilt them on an SFU (mediasoup): each speaker sends once and the server forwards, the ring lights up on whoever is talking, and a crashed media worker restarts on its own instead of taking the app down. Then I audited the whole realtime layer, chat and notifications included, and shipped the critical fixes. ## Premium locks the tracks, not the tab Free members still see the whole catalogue; what they can play is decided on the server, by the one entitlement service, with seven days' grace after a failed renewal. ## A signup that said yes for over a week Every new member is meant to leave signup with a growth path of their own. For over a week none did, and signup kept saying yes: a change had broken the path builder, and the way it was called swallowed the error before anyone could see it. The fix was one line. The real work was making sure it can’t happen quietly again: signup now checks the path exists, flags the account if it doesn’t, and tells the team. ## Next Refunds and dispute handling, and a load test that puts a number on how many listeners one voice box carries. --- # HyperPersona > An agentic recommender on AWS Bedrock that four of us built in a week: top 0.1% of 22,000+ at Cognizant Technoverse. I built the storefront and event SDK. - Kind: hackathon - Role: The storefront and everything the browser sends - When: One week, May 2026 - Team: Four of us - Stack: React 19, TypeScript, FastAPI, DynamoDB, OpenSearch Serverless, Bedrock, Redis - Live: https://hyperpersona-web.vercel.app - hyperpersona-web.vercel.app: https://hyperpersona-web.vercel.app We built HyperPersona in a week as a team of four at Cognizant's Technoverse 2.0, and it placed in the top 0.1% of 22,000+. It learns why a shopper buys, not just what. The browser streams consented events, agents on Bedrock turn them into facts, and a recommender ranks those facts, drafts an offer and checks every claim against its sources before showing it. I built the storefront and everything the browser sends, and shaped the architecture with the team. ## My part - **The storefront.** The React 19 shop the engine personalises, from the catalogue to the product pages. - **An event SDK.** Nothing lost to a crash or a closed tab: events stored first, sent in batches, flushed on the way out. - **Consent on every event.** Each event carries what the shopper agreed to share, checked in the browser and again on the server. Only what both allow is kept. - **Search that knows when to stop.** A relevance cutoff, so a search for “towel” stopped returning 48 shirts, on OpenSearch Serverless, which I moved us to. - **A noise gate.** Six low-signal event types never reach the model, so it learns from intent, not page views. - **The architecture, with the team.** How events travel from the browser to Bedrock and back, decided together. The agents, the fact ranking and the chain-of-verification are Divyansh Sharma's; the shop backend and CI are Anamika Aggarwal's. ## Events that survive the tab closing Recommendations are only as good as the signals that reach them, and a browser loses signals all the time: tabs close, pages reload, phones drop off the network. So every view and click is saved in the browser first and sent in small batches, and whatever is left leaves as the page closes. Nothing is lost to a reload, and nothing arrives twice. Underneath: events wait in IndexedDB with an idempotency key, go every 3 s or at 50, and the last batch rides out with keepalive, trimmed under the browser’s ~64 KB budget. --- # Safire > A harassment shield for LinkedIn DMs that won HackWIE 3.0 and Code Kshetra 2.0. I built the backend: the API, moderation failover and evidence reports. - Kind: hackathon - Role: The backend, and the extension's popup - When: Jan to Feb 2025, polished through July - Team: Four of us - Stack: Plasmo (MV3), React, Express, MongoDB, Redis, Gemini, Puppeteer, Next.js - Live: https://safire-five.vercel.app - safire-five.vercel.app: https://safire-five.vercel.app - Demo: https://vimeo.com/1059208124 - Code: https://github.com/rajveeerr/Safire Harassment in LinkedIn DMs usually gets read before anyone can report it. Safire screens every incoming message, hides abuse behind a warning, flags a sender once five different people have hidden them, and turns their messages into a PDF evidence report. Four of us built it, and it took first place at HackWIE 3.0 and Code Kshetra 2.0. It is a Chrome extension backed by an Express API and a Gemini moderation service, and it screens cheapest first: a list of abusive words, then verdicts already cached, and only then the model. My part was the backend (the API, the failover across five moderation deployments, the verdict cache and the evidence reports) and, in the extension, the popup and the login hand-off. ## My part - **The backend.** The Express and MongoDB API: auth, hiding senders and messages, the archive with filters and paging, and stats. - **Moderation that fails over.** Moved into its own service on five deployments, with health checks and a circuit breaker for each. - **A verdict cache.** Redis, keyed by message and platform, so the same message is never paid for twice. - **Evidence reports.** Gemini's findings rendered by Puppeteer into an A4 PDF, safe even when the model's answer isn't. - **The crowd rule.** A sender is flagged once five different people have hidden them. - **The extension's popup.** Preferences, blocked words, the hidden-message archive and reports, with login handed over from the dashboard. ## Cheapest first Most messages never reach the model. A word list in the page catches the obvious, a shared Redis cache answers anything judged in the last hour, and only what's left costs a call. ## Five deployments, one moderation service The model ran on free tiers, and free tiers run out, usually in the middle of a demo. So moderation moved into its own service on five deployments, and the API watches all of them: one that keeps failing is rested for a while, and each message goes to the first healthy one. ```js // api/moderation/detect-harassment.js router.post('/detect-harassment', async (req, res) => { const errors = []; const cacheKey = JSON.stringify({ message: req.body.message.toLowerCase(), platform: req.body.platform.toLowerCase() }); const cachedData = await redisClient.get(cacheKey); if (cachedData) { return res.json(JSON.parse(cachedData)); } const availableServers = servers.filter(server => { const state = getCircuitBreakerState(server.url); return server.healthy || state.status === 'HALF_OPEN'; }); for (const server of availableServers) { try { return res.json(await handleRequest(server, req)); } catch (error) { errors.push(`${server.url}: ${error.message}`); } } res.status(503).json({ error: 'All servers failed to process request', retryAfter: 5 }); }); ``` The cache answers first; then each deployment whose breaker is closed or half open gets its turn. ## Evidence that survives bad model output Gemini turns a sender's hidden messages into findings, patterns and a risk level, and Puppeteer renders them into an A4 PDF from a Handlebars template. The model's answer is never trusted as it comes: parse, strip the formatting and parse again, fall back to what can be read from the text, then type-check every field, so the renderer never gets a broken shape. ```js // api/user/report.js let parsedResponse; try { parsedResponse = JSON.parse(text); } catch (parseError) { const cleanedText = text .replace(/```json\n?|\n?```/g, '') .replace(/\\n/g, ' ') .trim(); try { parsedResponse = JSON.parse(cleanedText); } catch (secondParseError) { parsedResponse = { severityAssessment: text.includes('HIGH') ? 'HIGH' : text.includes('MEDIUM') ? 'MEDIUM' : 'LOW', recommendedActions: ['Review evidence manually'], }; } } const sanitizedResponse = { severityAssessment: parsedResponse.severityAssessment || 'MEDIUM', keyFindings: Array.isArray(parsedResponse.keyFindings) ? parsedResponse.keyFindings .filter(item => item && typeof item === 'string') : ['Manual review recommended'], }; ``` Three tries at the model's answer, then a safe default for every field the report needs. ## Next A labelled message set to measure detection accuracy, and a Chrome Web Store listing. --- # Wii Hop > Most of the web app for a US local-deals startup, from deals on a live map to heists that make a night out a game, plus the backend where money moves. - Kind: contract - Role: Software Engineer, contract - When: Jul 2025 to May 2026 - Stack: React, TypeScript, Leaflet, Node, Prisma, Expo - Live: https://beta.yohop.com - beta.yohop.com: https://beta.yohop.com - The web app: https://geolocation-mvp.vercel.app Wii Hop, formerly Yohop, is a US app for local deals, events, table booking and loyalty, with heists that turn a night out into a game. On a contract from July 2025 to May 2026, I built most of the web app customers and merchants use, the backend features where coins, stock and prices change hands, and screens for its Expo app. ## What I built - **Most of the web app.** Deals on a live map, merchant onboarding, table booking, events, heists, loyalty and streaks: most of what a customer or merchant sees. - **An eight-area admin console.** Where the people running the platform manage it, across eight areas. - **AI tools for merchants.** The interface for seven of them, switched on together behind one feature flag. - **The heist item shop.** Players spend coins on items mid-heist, and the purchase lands inside the heist itself, so a coin is never half spent. - **Events from SeatGeek.** Nearby concerts and games pulled in by place, date and keyword, shown like the app’s own events. - **Catering and bulk orders.** Quotes priced on the server with stock checked line by line, and a restaurant can edit up to 200 menu items in one go. - **The Expo app.** About 40% of its interface: merchant tabs, onboarding, and the deal, event and menu editors. --- # Modulus > Interview prep for consulting, product and finance roles. I built its product from the first screen: programs, booking, payment and live sessions. - Kind: internship - Role: Frontend Engineer Intern, remote - When: Apr to Aug 2025 - Stack: React, TypeScript, Vite, React Router, TanStack Query, Tailwind, hls.js, Firebase, PostHog - Live: https://www.trymodulus.com - trymodulus.com: https://www.trymodulus.com Modulus prepares people for interviews in consulting, product, finance and other non-tech roles, with cohorts, mentorship, live sessions and peer practice. As its frontend engineer I built the product from the first screen: the pages that sell a program, booking and payment, scheduling with a mentor, and live sessions with chat beside the stream. ## In numbers - 47: routes - #8: of the week on Peerlist ## What I built - **The product, from the first screen.** 47 routes, from the first page a visitor sees to the session they booked. - **Cohorts and programs.** The pages that sell and run them, with every filter kept in the URL. - **Razorpay booking.** Enrollment paid with credits or card, with verification left to the server's webhook by design. - **Scheduling.** Sessions booked through Cal.com without leaving the product. - **Live sessions.** HLS streams with Firebase chat running beside them. - **Fast, and measured.** Code split by route, and PostHog funnels so the team could see where people dropped off. --- # 10xAnswers > A drop-in React chat widget for AI answers, built and published alone: 12,493 downloads across 55 releases on npm. - Kind: side project - Role: Solo - When: Dec 2024 to May 2025, maintained since - Stack: React, Recoil, Vite, Express, Gemini - Live: https://10x-answers.vercel.app - 10x-answers.vercel.app: https://10x-answers.vercel.app - npm: https://www.npmjs.com/package/10xanswers 10xAnswers is a React chat widget you drop into a site: pass it a persona prompt and a backend URL, and it answers questions about you or your product. I built it alone and published it to npm, where it has 12,493 downloads across 55 releases. It renders markdown with highlighted code, lets you edit a question, and is themed through about 20 CSS variables, while a hosted Express proxy keeps the model key on the server. ## What I built - **The component.** A draggable React chat widget: markdown, highlighted code, editable questions, themed with about 20 CSS variables. - **Published on npm.** 12,493 downloads across 55 releases, as ES and UMD builds behind an exports map. - **The hosted proxy.** The model key stays on the server; input validated, size capped, clean 400, 413, 429 and 502 answers. - **A playground that writes the code.** A 15-field configurator that turns your choices into the snippet to paste. - **A TypeScript port.** For Plura AI as plura-chatbot, where I'm the largest contributor. ## Kept alive In September 2026 I found that a junk dependency had made the published package uninstallable for everyone, and shipped 1.1.17 to fix it. The hosted backend was calling a retired model. It now uses an always-current alias and answers like an API: 400 for a malformed request, 413 over the prompt cap, 429 when the model is rate limited, 502 when it fails. ```js // 10xAnswers-BE · api/index.js app.post("/", async (req, res) => { const prompt = req.body?.contents?.[0]?.parts?.[0]?.text if (typeof prompt !== "string" || !prompt.trim()) { return res.status(400).json({ message: /* shape */ }) } if (prompt.length > MAX_PROMPT_CHARS) { return res.status(413).json({ message: /* limit */ }) } try { // History arrives inside the prompt from the client, // so this function stays stateless and conversations // from different sites never mix. const result = await model.generateContent(prompt) const text = result.response.text() res.json({ candidates: [{ content: { parts: [{ text }] } }], }) } catch (err) { const status = err?.status === 429 ? 429 : 502 res.status(status).json({ message: /* try again */ }) } }) ``` Stateless on purpose: history travels with each question, so two sites using the same backend never share a conversation. ## On npm 12,493 downloads so far, across 55 releases. ## Next Version 2: streaming, React 19, provider adapters and a voice mode, as the interface of an agent on this site. --- # Kernel > Anonymous chat rooms with peer-to-peer voice, where the server only relays the handshake. The take-home for the role I have now. - Kind: take-home - Role: Solo - When: Two days, Sep 2025 - Stack: React, Socket.IO, WebRTC, MongoDB - Live: https://kernelchat.vercel.app - kernelchat.vercel.app: https://kernelchat.vercel.app - Code: https://github.com/rajveeerr/Kernel Kernel was the take-home for the role I have now. The brief was anonymous, room-based chat; I added peer-to-peer voice calls and delivered it twelve hours before the deadline, built, styled and deployed. The server's job is to introduce two browsers, then get out of the way. One Socket.IO server handles chat and signalling; for a call it only relays offers, answers and network candidates between two people, and never sees the audio. ## What I built - **Anonymous rooms.** Room-based chat with the history kept per room, and no accounts at all. - **Peer-to-peer voice.** Calls I added beyond the brief, with the audio going straight from browser to browser. - **A signalling server.** One Socket.IO server relays offers, answers and candidates, and never touches the audio. - **A handshake that can't race.** Candidates that arrive early wait until the other side's description lands. - **Delivered early.** Built, styled and deployed in two days, twelve hours before the deadline. ## Only the handshake For a call the server does three things, and each is a relay: the offer goes to the callee, the answer comes back, and network candidates pass between them. ```js // server/src/index.js socket.on("call-user", (payload) => { const { to, from, offer } = payload; io.to(to).emit("call-made", { offer, from }); }); socket.on("make-answer", (data) => { const { to, answer } = data; activeCalls[socket.roomId] = true; io.in(socket.roomId).emit('call_in_progress'); io.to(to).emit("answer-made", { answer }); }); socket.on("ice-candidate", (data) => { const { to, candidate } = data; socket.to(to).emit("ice-candidate", { candidate }); }); ``` Nothing here reads a stream. Once the peers connect, the server is out of the call. ## Candidates that arrive early A network candidate can arrive before the other side's description is set, and adding it then fails. So early ones wait in a queue and are applied once the answer lands, and a fast network never races the handshake. ```js // client/src/pages/ChatRoom.jsx const iceCandidateQueue = []; const iceCandidateListener = (data) => { if (peerConnectionRef.current) { if ( peerConnectionRef.current.remoteDescription && peerConnectionRef.current.remoteDescription.type ) { peerConnectionRef.current.addIceCandidate( new RTCIceCandidate(data.candidate) ); } else { iceCandidateQueue.push(data.candidate); } } }; // Process queued ICE candidates after setting the // remote description const processIceCandidateQueue = () => { while (iceCandidateQueue.length > 0) { const candidate = iceCandidateQueue.shift(); peerConnectionRef.current.addIceCandidate( new RTCIceCandidate(candidate) ); } }; ``` Queued, then drained in order the moment the remote description is set. ## What I'd change It's one-to-one by design. At IABTM I later moved group audio to an SFU, for exactly the reason a mesh doesn't scale. --- # The lake > This site's hero: a lake painted in code, with water simulated on the GPU, creatures that make ripples, and fishing with your phone as the rod. - Kind: side project - Role: Solo - When: 2026 - Stack: WebGL2, GLSL, Web Workers, OffscreenCanvas, WebRTC, Cloudflare Workers, Durable Objects, Device sensors, TypeScript, React - Live: https://rajveers.com/pond - Full screen: https://rajveers.com/pond - The rod: https://rajveers.com/rod The lake at the top of this site is about 5,400 lines of WebGL2 and TypeScript. The water is a simulation, the light is physics done cheaply, the land is painted in code, and the fish, ducks, frogs and heron read the same terrain, so nothing swims onto the beach. It is also the most expensive thing on the page, so it goes last: the text paints first, it never blocks input, and it costs nothing when you can't see it. Full screen you can sit down and fish it, and the rod can be your phone: scan the code on the lake, point the phone, and lift it when a fish bites. That half is WebRTC, two independent pairing servers and a little orientation maths. ## What I built - **Water that is simulated.** A damped wave equation on the GPU, stepped at a fixed 60 Hz, stopped by the shore. - **Light, cheaply.** Caustics from how refracted triangles bunch up, absorption by depth, and glints of the sun. - **Terrain painted off the main thread.** A worker paints it on an OffscreenCanvas and hands it back without copying. - **A pond that is alive.** Fish, ducks, frogs and a six-state heron that read the terrain and push waves into the water. - **Your phone as the rod.** A QR code pairs a phone over WebRTC, and the way it points steers the hook. - **Never blank, never in the way.** A blurred picture first, paused off screen, one still frame under reduced motion. The sound synthesis and two of the toys elsewhere on the site are third-party packages. ## Water that is simulated, not drawn Each frame, the height of the water steps forward on the GPU: a damped wave equation on half-float textures that swap every step, at a fixed 60 Hz whatever the display runs at. Sampling eight neighbours rather than four keeps a ring round as it spreads, and the shore stops it dead. ```glsl // hero/lake/shaders.ts · the simulation step float l = texture(u_state, uv - vec2(u_px.x, 0.0)).r; float r = texture(u_state, uv + vec2(u_px.x, 0.0)).r; float d = texture(u_state, uv - vec2(0.0, u_px.y)).r; float u = texture(u_state, uv + vec2(0.0, u_px.y)).r; float a = texture(u_state, uv + vec2(-u_px.x, -u_px.y)).r; float b = texture(u_state, uv + vec2(u_px.x, -u_px.y)).r; float c = texture(u_state, uv + vec2(-u_px.x, u_px.y)).r; float e = texture(u_state, uv + vec2(u_px.x, u_px.y)).r; float lap = (4.0 * (l + r + d + u) + (a + b + c + e) - 20.0 * h) / 6.0; v = (v + u_c2 * lap) * u_damp; h += v; // Land holds no water, so waves stop dead at the shore. float keep = smoothstep(0.2, 0.65, texture(u_info, st).r); h *= keep; v *= keep; ``` The nine-point stencil: a five-point one spreads a drop as a diamond. ## Painted off the main thread Painting the terrain takes the better part of a second on a laptop and several on a phone. A worker does it on an OffscreenCanvas, and the layers come back as bitmaps whose memory is handed over rather than copied. ```ts // hero/lake/paint.worker.ts const bitmap = (c: Picture) => (c as OffscreenCanvas).transferToImageBitmap(); self.onmessage = (event: MessageEvent) => { const { id, height } = event.data; const painted = paint(height); const layers: Layers = { ...painted, bed: bitmap(painted.bed), ground: bitmap(painted.ground), canopy: bitmap(painted.canopy), edge: bitmap(painted.edge), }; const { field, info } = layers; self.postMessage({ id, layers }, { transfer: [ layers.bed, layers.ground, layers.canopy, layers.edge, info.data.buffer, field.water.buffer, field.depth.buffer, field.canopy.buffer, ], }); }; ``` Transferred, not copied: the worker gives up the memory, so the hand-off costs nothing. ## Your phone as the rod Sit down to fish and the lake shows a QR code and six letters. The phone that scans them becomes the rod: the way it points steers the hook, a bite tugs in your hand, and lifting it sharply brings the fish out. The page is the game; the phone only sends where it points and when it was lifted, sixty times a second. Pairing is the part that can fail, so there are two servers that can do it and they share nothing: our own Cloudflare Worker, one hibernating Durable Object per room, and PeerJS's free public one. The page opens the room on both. The phone tries ours, and brings PeerJS in as well if ours is slow, unreachable or has no page in that room. Whichever answers carries the game, and if it drops the other takes over. Once either has introduced them the messages go straight between the devices over a WebRTC data channel, usually a few milliseconds apart, with no server in the path: the aim on an unordered channel that never resends, since a late sample is worthless, and everything else in order. Where a network will not allow a direct path the relay keeps carrying them itself, a little slower but never stuck. A socket can die without closing, so both sides send a heartbeat Cloudflare answers without waking the room, and one phone holds a room at a time. ```ts // cast/link/link.ts · whichever answers /* The link carrying the game: the one already doing so while it can, else the first with the other device on it. */ const pick = () => { if (active?.state.peer) return active; active = links.find((l) => l.state.peer) ?? null; return active; }; /* ... published on every change: */ // The phone: a relay that cannot help brings // PeerJS in at once. if ( RELAY_URL && role === 'rod' && !a && links.length === 1 && (links[0].state.failed || links[0].state.missing) ) { withPeer(); } /* ... and at the start: */ if (RELAY_URL) { links.push(createRelayLink(RELAY_URL, transport())); if (role === 'host') withPeer(); else fallbackTimer = window.setTimeout( () => !pick() && withPeer(), HEAD_START, ); } else { links.push(createPeerLink(transport())); } ``` No preference and no failover logic to get wrong: the link carrying the game is simply the one with a device on it. ## Where a phone is pointing A phone does not report where it points, and the one angle that looks like it (beta) stops meaning anything once the phone is rolled. So the aim is read from the whole orientation: the three angles are turned into the direction of the phone's top edge in the world, or of its back for a phone held up like a screen, and the yaw and pitch of that direction become how far out and how far to the side the hook goes. Rolling the phone then does not move the aim, and which way it is held is decided once, when you point at the screen to set the centre. A lift is read from the same direction rising fast, rather than from the rotation rates, because Safari and Chrome do not name those axes the same way. And while the phone is swinging the aim holds where it was a moment before, so the lift itself does not drag the hook off the fish it was lifted for. What is left is smoothed with a one euro filter: steady when the hand is still, quick when it moves. None of it is required. With no sensors to read, refused, not on https, or simply absent, the lake on the phone becomes a trackpad: the float follows your finger and a tap lifts. ```ts // cast/sense.ts · the aim, from the whole orientation export function pose( alpha: number, beta: number, gamma: number, hold: Hold, ): Pose { const ca = Math.cos(alpha * RAD); const sa = Math.sin(alpha * RAD); const cb = Math.cos(beta * RAD); const sb = Math.sin(beta * RAD); const cg = Math.cos(gamma * RAD); const sg = Math.sin(gamma * RAD); // The phone's top edge in the world (east, north, // up), and its back. const p = hold === 'top' ? [-sa * cb, ca * cb, sb] : [-(ca * sg + sa * sb * cg), -(sa * sg - ca * sb * cg), -cb * cg]; const n = Math.hypot(p[0], p[1], p[2]) || 1; return { yaw: Math.atan2(p[1], p[0]) * DEG, pitch: Math.asin(clamp(p[2] / n, -1, 1)) * DEG, }; } ``` Three angles to one direction in the world: roll drops out, and nothing depends on how a browser names its axes. ## Next Publishing LCP, INP and frame time on a mid-range phone. --- # The AI me > An AI version of me that answers for my work while I’m away: it reads this site, replies in my voice with its sources, shows the real projects instead of describing them, checks a job description against what I’ve built, and moves around the site when you ask. - Kind: side project - Status: in progress, live and still being built - Role: Everything: the engine, retrieval, tools, guards, interface and evals - When: Sep 2026 to now - Stack: Next.js, React 19, TypeScript, Vercel Functions, Upstash Redis, MongoDB Atlas, Server-sent events, NDJSON, JSON Schema, Bun test, Playwright - Live: https://rajveers.com - Ask it: the tab on the right, or press a: https://rajveers.com - Is it up?: https://rajveers.com/api/ask/health The AI me is a version of me that answers for my work while I’m away. Open the Ask tab on any page and ask: it reads what this site says, answers in my voice with its sources linked, and puts the real things between its words (the project cards, my calling card, my calendar) instead of describing them. Tell it to, and it moves around the site for you: it scrolls to a part, takes you to a page, makes it night or lets it rain, each with an undo. It is honest about what it is. It says it’s an AI when asked, and when the site doesn’t say something it says so and offers to send your question to the real me rather than guess. Recruiters can paste a job description and get a table of where my work fits and where it doesn’t. I built every layer from scratch, on no agent framework and no vendor’s SDK, so any chat-completions model can run it: the engine, retrieval, tools, guards, interface, tracing and evals. It’s live and still growing. Under the paragraphs that describe a feature there is a question to try: press it and the chat opens with it typed in, for you to send and watch that feature happen. ## In numbers - 31: tools it can call, each checked - 25 KB: fetched only when you reach for it - 31 of 32: questions find their page in the top five - 284: unit tests on the engine ## What it does - **Answers from the site, with sources.** Every answer is built from passages of these pages, and a small number after a sentence links to the part it came from. - **The site’s own components in the chat.** The same cards and buttons the pages use, not descriptions of them, shown when the answer is one. - **Acts on the page, with an undo.** Scrolls to a part, rings it, takes you to a page, switches day and night or the weather, opens every way to reach me. - **A fit check for recruiters.** Paste a job description; see each requirement against my work, strong, some or a gap, linked to where it is shown. - **Honest by construction.** Grounded in what the site says, cited, corrected when it claims what didn’t happen, and handing over instead of guessing. - **Four providers, then the pages.** Keys in turn, backups each with its own quota, answers from the pages alone, then the real me. - **Guarded against abuse.** A session per visitor, limits in six windows, a daily token budget, a breaker per provider. - **Private by default.** Counts, not transcripts; 30 days of redacted questions for my review; your conversation stays on your device. - **Free until you reach for it.** The chat’s code arrives when a hand nears the tab, and a command is done without a model at all. - **Traced, counted and reported.** Every answer carries its steps, timings and tokens; my phone hears about fit checks, trouble, and the day at 8 pm. - **Evals that gate the prompt.** 44 questions scored against the real model, and retrieval checked in CI. ## What it’s for, and who It is for when I’m away, which is most of the time someone reads this site. A portfolio gets the same few questions, about the work, how it was made and whether we should talk. The AI me answers them in a couple of seconds, shows the proof, and hands over to the real me when it should. A recruiter pastes a job description and gets a fit check: the role’s requirements, the evidence from my work beside each, how strong a match it is, and the gaps said plainly and sorted last, then a way to talk. A founder says what they’re building and gets the closest thing I’ve shipped. Someone in hiring asks about notice periods or rates and is told, politely, that it can’t promise anything for me, with my calendar one tap away. An engineer asks how it works and can open the steps, timings and tokens of their own answer, which is what the rest of this page is about. The fit check is one request for JSON only, and what comes back is checked before it is shown: its shape against a schema, every link against the pages of this site, and a gap never gets a link. If the reply doesn’t hold together, it says so and offers to have me look at the description myself. ## One question, end to end Take “How do you stop double charges in billing?”, asked on the home page. This is one real run, timed. The chat’s code was fetched when the pointer neared the tab, and it asked for a session when it opened. The question goes out with that session, the page it was asked on, the labels of that page’s parts, the time and weather showing, and the last few turns with notes of what they showed and did. The guards count it in six windows, look for a kept answer, and check both providers’ breakers and the day’s tokens, all in one round trip to the store. Then the turn. The five best passages are found in 8 ms. Nothing in this question asks to see a card, so the model is asked, with 17 of the 31 tools on offer. Its first request ends by calling for the IABTM card, which is on screen at 1.28 s. That request said nothing, so the loop goes again, and the second writes the answer, citing passage 5, the billing section of the IABTM write-up. First words at 2.35 s, done at 2.73 s: two requests, 10,260 tokens in, 95 out. Under the answer, one line folds everything it did: the pages it read, the tools it ran, the time it took. Open it for each step, then each request’s milliseconds and tokens. ## An engine that doesn’t care what carries it The same brain will answer by voice, and later on other people’s sites, so it is a core with nothing about the chat in it. One function, runTurn, takes a question, the turns before it and the page it was asked on, and yields events: the passages it read, words, a component to show, an action to take on the page, what it’s busy with, a trace, and done. The chat’s route writes them one JSON object a line (NDJSON) as they come, so a card can arrive before the model has said a word. A voice transport will turn the same events into speech and captions. The eval script runs the core with no HTTP at all. The route waits for the first word or card before it answers. A model that never begins is then said plainly with a status, instead of a stream that opens and breaks. Every turn has a deadline, and a stream that goes quiet partway is given up on rather than left hanging. ```ts // lib/agent/ui.ts · what a turn streams /** A turn as it streams to the chat: one JSON object a line. */ export type TurnEvent = | { t: 'sources'; sources: Source[]; kept?: boolean; /** Answered from the pages alone. */ pages?: boolean; } | { t: 'text'; v: string } /** 'thinking', 'writing', or a tool's name. */ | { t: 'doing'; what: string; detail?: string } | ({ t: 'show'; id: string } & Shown) | ({ t: 'act'; id: string } & Act) | { t: 'trace'; trace: Trace } | { t: 'error'; m: string } | { t: 'done' }; /** Something done on the page rather than shown in the chat. */ export type Act = | { a: 'scroll'; mark: string } | { a: 'highlight'; mark: string } | { a: 'navigate'; href: string; label: string } | { a: 'time'; time: 'day' | 'night' } | { a: 'season'; season: string } | { a: 'contact' }; ``` The whole protocol between the brain and whatever carries it. The chat folds each event into the turn; voice will speak the same ones. ## Any model, and a line of backups The model is an address, a name and keys in the environment: any endpoint that speaks chat completions and streams server-sent events. Nothing in the code names a provider. The stream is read by hand: words as they come; each tool call assembled from its pieces, by index or by id, and handed over the moment it is whole; arguments whether sent as text or as an object; any thinking an endpoint sends beside the words, never shown; and the tokens used, estimated when an endpoint doesn’t report them. Keys are taken in turn. One the endpoint refuses rests a minute (a rate limit) or an hour (spent, revoked, suspended, or a 400 whose words name the key, which is how one provider says a key is bad), and the next is tried within the same request. Keys in one account or project share its quota, so rotation covers a failing key, not more free use. Behind the first provider is a line of backups, three today, each wholly apart: its own endpoint, model, keys, quota and breaker. The next is asked when every key before it is resting or refused, when one answers with any error or nothing, while its breaker is open, or when the first has written nothing after 4 seconds (a backup gets 8, so two slow ones still fit inside the turn’s 25). A turn a backup starts, it finishes, and a question counts as one failure against a provider however many of its keys were tried. The backups are offered a slimmer set of tools, the ones most answers need, because their allowance of tokens a minute is often the smaller one. Only when every provider is down does it answer from the pages alone, and my phone hears about it at once. **Who a step asks, in order** | Try | Who | Why it moves on | | --- | --- | --- | | 1 | The first provider, each key in turn | A key refused (429, 401, 403, or a 400 naming the key) rests; the next is tried at once. | | 2 | The same endpoint’s spare model, if one is set | The first model fails or is slow to start. | | 3 | Each backup in order, each of its keys | No word yet after 4 s (8 s for a backup), any error, or nothing at all. | | 4 | The pages alone | Every provider failing or its breaker open, or the day’s tokens spent: links to the best passages, and the obvious card. | | 5 | The real me | Anything that still fails ends in a note to my phone, on the visitor’s press. | Only while nothing has been said: once words are on screen, a step never switches provider under them. ## A short loop, and tools as data The loop is short on purpose: at most three requests and 20 seconds a question. The model writes, or calls tools. A show tool’s component goes to the visitor as soon as its call is whole. A data tool’s result goes back to the model for another pass. The last pass is offered no tools, so it has to answer, and once it has shown something and said something, the turn ends. Across 20 timed questions it took 1.5 requests on average. Every capability is one entry in a registry: a name, what the model is told, a JSON schema for its arguments, and what it does. The schema the model is shown is the one its arguments are checked against. Properties it doesn’t name are dropped, and a bad call (an unknown tool, broken JSON, an argument out of range) comes back to the model as an error to read, never a throw. Only 17 of the 31 are offered by default, the rest when the question’s words ask for them, because every tool offered is read before the first word. The page actions are offered only for a command. ```ts // lib/agent/turn.ts · the loop export const MAX_STEPS = 3; export const BUDGET_MS = 20_000; for (let step = 0; step < MAX_STEPS; step++) { const last = step === MAX_STEPS - 1 || tracer.elapsed() > BUDGET_MS * 0.6; for await (const event of streamStep(config, messages, { tools: last ? undefined : specs })) { if (event.type === 'text') yield* words(event.text); else if (event.type === 'call') { calls.push(event.call); const tool = registry.get(event.call.function.name); // What shows, shows now; what looks up waits // for the words. if (tool && tool.kind !== 'data') results.push(yield* call(/* the call */)); } } if (!calls.length) break; for (const c of data) results.push(yield* call(/* the lookup */)); // Shown, and something said: nothing more to ask for. if (!data.length && text.trim()) break; if (last) break; messages.push(/* the calls, then each result */); } ``` Three requests at most, the last without tools. A card shows while the words are still coming; a lookup waits for them. ```ts // tools/page.ts and tools/registry.ts { name: 'navigate', kind: 'page', description: 'Take the visitor to another page of the site now, ' + 'when they ask to go there.', parameters: { type: 'object', properties: { path: { type: 'string', enum: sitePages().map((p) => p.path) }, }, required: ['path'], }, run(args) { const page = sitePages().find((p) => p.path === args.path)!; return { result: `Going to ${page.path}.`, act: { a: 'navigate', href: page.path, label: page.what }, }; }, }, /* ... and every call, whatever the model wrote: */ const tool = this.get(name); if (!tool) return { ok: false, result: { error: `There is no tool called ${name}.` } }; const parsed = parseArguments(rawArguments); if (!parsed.ok) return { ok: false, result: { error: parsed.error } }; const args = check(tool.parameters, parsed.value); if (!args.ok) return { ok: false, result: { error: args.error } }; ``` A tool is one object. The model can only name a page from the list, and whatever it writes, the call is checked against the schema it was shown. ## Cards before the model, and only when asked Some questions don’t need a model to know what to show. A page of patterns plans the obvious calls and runs them before the model is asked; it is told they are done and not offered them again. That is why a card can be on screen in milliseconds. A plain command (“make it night”) goes further: it is done and said in a line with no passages and no model at all. The patterns learned restraint the hard way. They used to fire on questions: a question about summer changed the weather, one about a day job made it day. Now an action needs a command, a card comes before the model only for an explicit ask to see it, and the model’s own cards are one to an answer unless a list was asked for, never the one shown just before. The components are the site’s own: the same project cards, calling card, résumé and calendar as the pages, drawn in the chat from props the server builds out of the site’s content. The model names a project by its slug, from a list; it can’t describe one into being. ```ts // lib/agent/turn.ts · the obvious calls // The obvious calls first, before the model has // started (intent.ts); it is told they are done, // and not offered them again. const planned = plan(ask); const done: string[] = []; for (const p of planned) { const out = yield* call(p.name, JSON.stringify(p.args)); done.push(`${p.name}: ${JSON.parse(out.content)}`); } if (done.length) { const names = new Set(planned.map((p) => p.name)); specs = specs.filter((s) => !names.has(s.function.name)); messages[0].content += `\n\nALREADY SHOWN to the visitor for this question` + ` (don't call these again; write your answer` + ` around them):\n${done.join('\n')}`; } ``` Why a card can be on screen in milliseconds: it was never the model’s to make. ## Memory, and what a follow-up carries A conversation is only useful if “tell me more about that” means something. Each question carries the last six turns, and with each answer a note of what it showed and did, so the model knows the card on screen and the night it made. The last four turns are kept whole, and the facts and passages are fitted into what is left. It used to be the other way round: the prompt’s budget went to the facts and passages first, and a long answer or two starved the conversation out of it, so a follow-up lost its thread. Keeping memory first, and a prompt of up to 28,000 characters, fixed that. Retrieval reads a follow-up with the turn before it, without a second model call: a short question that points back borrows the last question’s words and the projects the last answer named. ## Retrieval from the site’s own pages Every answer is built from what this site says, and says where. The corpus is the site itself. Every page already has a markdown copy for agents (this one is at /projects/ai-me.md). Those are cut at their headings, then into passages of 400 to 900 characters at paragraph and sentence ends, each with a title and the anchor of its own section. With two more about the site itself (its day, night and seasons, and this chat), that is 173 passages from 14 pages, this one included, made once, in memory. The fact sheet, about 7,300 characters of who I am, every job and project in a line, and the record, goes into every prompt whole, so a question about dates never depends on retrieval. Passages are found by their words with BM25: titles counted twice, a light stemmer, stop words out, a small table of a visitor’s words for the words the pages use, a few phrases read for what they mean, and a lift of 1.6 for the passages of a project the question names. It takes under a fifth of a millisecond a question. They can also be found by meaning, where the words are unsure. The passages’ vectors are made once, when the site is built, only for what changed, and kept in MongoDB Atlas; a question embeds itself, Atlas Vector Search finds its neighbours, and the two ranked lists are merged by reciprocal rank, with a 700 ms cap. Measured on the 32 eval questions it didn’t pay: 30 in the top five against the words’ 31, and about 0.4 s a question. So it is off, ready for a better embedding model or a larger corpus, and the words answer alone. On a write-up, a question that names no project is about that one: asked here, “this” means the AI me. Citations are checked, not trusted. The model may cite only the passages it was given, by number. A citation to anything else is dropped as the words stream, one written loosely, as [2], is made a real one, and a search mid-turn adds its passages to the list so they can be cited too. **Words alone, and words with meaning** | Retrieval | Words only | Words and meaning | | --- | --- | --- | | hit@1 | 21 of 32 | 26 of 32 | | hit@3 | 29 of 32 | 30 of 32 | | hit@5 | 30 of 32 | 31 of 32 | | MRR@5 | 0.78 | 0.87 | | Time a question | Under 0.2 ms | 0.52 s median, 1.14 s at p90 | The 32 eval questions with an expected page, on 1 October 2026, over 149 passages. Meaning puts the right page first five more times, for about half a second a new question; the slowest tenth pass the 900 ms cut and fall back to words. ```ts // lib/agent/retrieve.ts · BM25 for (const doc of this.docs) { let score = 0; for (const [w, weight] of q) { const f = doc.tf.get(w); if (!f) continue; const n = this.df.get(w) ?? 0; const idf = Math.log(1 + (N - n + 0.5) / (n + 0.5)); score += weight * idf * ((f * (k1 + 1)) / (f + k1 * (1 - b + (b * doc.length) / this.avg))); } if (score > 0 && doc.chunk.slug && named.has(doc.chunk.slug)) score *= 1.6; if (score > 0) hits.push({ chunk: doc.chunk, score }); } ``` Plain BM25 (k1 1.2, b 0.75) over synonyms weighted a half, and the named project lifted. In memory, in under a fifth of a millisecond. ## Tools, and acting for you It can do 31 things, each on a list it can’t add to. Five look something up for the model. Eighteen put one of the site’s components in the chat: my timeline, what I’m reading, my local time, a drafted note and the rest. Eight act on the page, a tour of it and undoing the last change among them. Told to, and only when told to, it scrolls to a part of this page, rings a part for a moment, walks you through the page’s parts one at a time, takes you to another page, makes it day or night, changes the weather, or opens every way to reach me. It knows when it’s already night, and says so instead. The server only proposes. The chat carries each action out from its own short list, checks the arguments again, finds its target by the label the page gives it (never a selector, a script or an address from the model), and leaves a receipt with an Undo where there is something to undo. On a phone, where the chat covers the page, it steps aside so you see what it did. Commands still work misspelled (“sprng”, “drak mode”), mended by edit distance, and never fire on a question: “what’s summer like?” changes nothing. Asking it to tell me something drafts the note, and a part named by its label is scrolled to or ringed, all without a model. A test matrix of more than 100 ways of asking, from each state of the page, holds it to that. Nothing it does sends anything on your behalf. A drafted note opens for you to read and send; a question for the real me reaches my phone only when you press Send; booking opens my calendar for you to pick a time. The model has no tool that sends, pays or writes. **The 31 tools, by kind** | Kind | Tools | What they do | | --- | --- | --- | | Look up (5) | search_site, get_project and three more | Go back to the model: another search, or one thing in full. | | Show (18) | show_project, compare_projects, open_page, draft_note, ask_real_me, suggest and twelve more | Put one of the site’s components in the chat the moment the call is whole. | | Act (8) | scroll_to, highlight, tour, navigate, set_time, set_season and two more | Do something on the page, carried out by the chat from its own list, with a receipt and an undo. | 17 are offered on every question; the rest join when the question’s words ask for them, and the actions only for a command. ```ts // components/ask/actions.ts case 'navigate': { if (!/^\/(?!\/)[a-z0-9\-/]*$/.test(act.href)) return { ok: false, label: "Couldn't go there" }; hooks.stepAside(); hooks.go(act.href); return { ok: true, label: `Took you to ${act.label}` }; } case 'time': { if (!isTimeChoice(act.time)) return { ok: false, label: "Couldn't change the time" }; const choice = getTimeChoice(); setTimeChoice(act.time); return { ok: true, label: `Switched to ${act.time}`, undo: { a: 'time', choice }, }; } ``` The browser checks again what the server already checked, and remembers what to put back. ## Keeping it honest It is meant to be trusted by someone making a decision about me, so it would rather say less. It speaks as me, but says plainly that it’s an AI version of me when asked, and never commits me to a date, a rate or a meeting. When the site doesn’t cover something it says “I haven’t written about that here, so I’d rather not guess” and offers a card that sends the question to the real me. Grounding is layered. The brief forbids numbers, dates, employers, links and opinions that aren’t in what it was given; the facts and passages are its only sources; citations must point at a passage it read; the fit check’s links must be pages of the site, and a gap never gets one; and project cards come from content, so a project that isn’t there can’t be shown. What it did is checked too. If an answer says it did something on the page that nothing did, a line after it sets that straight, and “nothing happened” is answered from what the last answer actually ran, with how to do it by hand. The eval checks every number in every answer against the site’s own text, which is how one rounding of 12,493 downloads to 12,000 was caught. ## Prompt injection, and what comes out Prompt injection is handled by capability more than by wording. The visitor’s message, the passages, tool results, the page’s parts and a pasted job description are all data. The brief says so, the job description is fenced and described as text to read, never instructions, and the page’s parts ride with the question as data rather than as part of the brief. Obvious tries (“ignore your previous instructions”, “you are now”) and jobs it isn’t for (write me code, the weather) get a line back in my voice without reaching the model at all, in under 0.1 s. The real defence is that there is nothing worth taking over: no tool sends, pays or writes, arguments are enums and patterns checked on the server and again in the browser, and page actions find things only by their visible labels. What comes out is filtered as it streams. No HTML is ever drawn: answers are parsed into plain data (paragraphs, lists, bold, code and links) and drawn as elements. A link goes somewhere only if it is https or a page of this site, and there are no images, so an answer can’t carry data away in an image’s address. Long dashes become commas, and answers stop at 2,400 characters. A tool call the model writes out as words instead of making it comes in many shapes: a name and its arguments, a name in bold, a bullet, tags, a JSON code fence, a bracketed marker. From where one could begin, the words are held until it is clear what they are; a call is dropped, its questions become follow-ups, and a card it asked for is shown for real. ```ts // lib/agent/leaks.ts /* * A tool call the model wrote as words instead of * making it, in many shapes: `suggest: [...]`, * `show_project({...})`, `functions.show_project(...)`, * a bold or bulleted name, `` tags, a JSON * code fence, `[TOOL_CALLS]`. From where such a thing * could begin, the words are held until it is clear * what they are; a call is dropped, and what it asked * for is kept. */ constructor(private readonly names: readonly string[]) { const alt = names .map((n) => n.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')) .join('|'); this.trigger = new RegExp( '|```|\\[TOOL_CALLS\\]|\\bfunctions\\.' + `|(?:\\*\\*)?\\b(?:${alt})\\b`, 'i', ); } ``` Ordinary words pass at once; only from a trigger are words held, and only until they can’t be a call. ## The guards, and why each value A chat that calls a paid model from a public page is a bill anyone can run up, so every limit is counted in the store, in fixed windows, and everything fails closed: if the limits can’t be counted, the model isn’t asked. A question has to come from the site’s own pages, from something that says it is a browser, and with a session token the chat fetched when it opened: an HMAC over when it was issued and the address it was issued to, good for two hours, compared in constant time and never stored. The limits count a visitor under a hash of the address with a secret and the day, so no address is kept and today’s can’t be tied to yesterday’s. The token is sealed to the address itself, so a session doesn’t die at midnight when that hash changes, which it used to. Answers in flight are a set of requests stamped with when they began, not a counter. A counter drifted up whenever a function was killed mid-answer and never counted itself out; a stamped entry simply ages out after 90 seconds. **Every guard, in the order a question meets it** | Guard | Value | Why | | --- | --- | --- | | Origin | The site’s own pages | A question has to come from here. | | Browser | No script or crawler user agents | Scrapers and curl stop before anything is counted. | | Session | HMAC token, 2 hours, sealed to the address | A script that never opened the chat has none, and one visitor’s can’t be reused by another. | | Size | 28,000 bytes a request; 400 characters a question; 4,000 a job description | The largest honest request, and not a byte more. | | Burst | 4 in 10 s from one address | A held-down Enter or a loop, stopped without bothering a person. | | Per visitor | 8 a minute, 30 an hour, 80 a day | More than any real conversation; one visitor can’t spend the day. | | Everyone | 60 a minute, 1,500 a day | A ceiling on the bill, however many come. | | In flight | 24 answers at once, each aging out after 90 s | Within the model’s per-minute limits and the function’s concurrency. | | Tokens | 1,500,000 a day, prompt and answer; stopped turns charged too | Past it, the pages answer until tomorrow. | | Breaker | 5 failed questions in 2 minutes, open 3 minutes, per provider | A failing model isn’t hammered, and one provider’s trouble doesn’t close the other. | | Keys | A refused key rests 1 minute (rate limit) or 1 hour (spent, revoked) | The next is tried in the same request, before a word is said. | | Turn | 3 requests, 20 s, 600 tokens an answer, a deadline, 8 s to the first byte | No question can run away with the budget. | Counted in one pipelined round trip with the kept-answer check, the breakers and the budget. A kept answer counts toward the limits too: they are against abuse, not cost. ```ts // lib/agent/guard.ts export const LIMITS = { burst: { count: 4, seconds: 10 }, minute: { count: 8, seconds: 60 }, hour: { count: 30, seconds: 60 * 60 }, day: { count: 80, seconds: 60 * 60 * 24 }, allMinute: { count: 60, seconds: 60 }, allDay: { count: 1500, seconds: 60 * 60 * 24 }, inFlight: 24, } as const; /** Tokens a day, prompt and answer together. */ export const dailyTokens = () => Number(process.env.ASK_DAILY_TOKENS) || 1_500_000; const BREAK = { failures: 5, windowSeconds: 120, openSeconds: 180 }; const SESSION_MS = 2 * 60 * 60 * 1000; /** A token for this address, good for two hours. */ export function issueSession(address: string) { const issued = String(Date.now()); return `${issued}.${seal(issued, address)}`; } ``` Every value in one place, each with a reason in the table above. ## When it breaks Everything upstream fails sometimes, so each failure has an answer a visitor can still use. Five failed questions in two minutes open a provider’s breaker for three. A question counts once against a provider, however many of its keys were tried; counting every try used to let a single question with three tired keys trip it. While both providers are resting, or the day’s tokens are spent, it answers from the pages alone: what the question obviously wants shown, and links to the three best passages, under a line saying so. When an answer doesn’t come (it broke, it’s resting, the limits were hit, it didn’t know, or a guard stopped it), the question is handed to the real me: a note with the question and where it was asked already in, a name and a way back if they like, sent to my phone only on their press, once a question. A stream cut off before its end is a failure with that note, not a half answer left looking finished. A health endpoint tells an uptime monitor whether each provider can be asked and whether the store answers, in counts and nothing secret. ## What it keeps, and for how long Your conversation lives on your device. It is kept in your browser for a week (the last 15 turns, and five earlier chats) so a reload or another page picks up where you were, and saved when the browser has a spare moment rather than while an answer streams. On my side there are counts, and the questions. The counts (turns, tools, how long each part took, tokens) say nothing about who asked or what. Each question is kept 30 days, in MongoDB with an index that deletes it on the day, with the first 240 characters of its answer, the page it was asked on, the tools that ran and a city and country as the edge placed them, so I can read what people ask and write the pages they’re missing. Email addresses and phone numbers are blacked out first. No address, no browser, no device, and the chat says so under its field. **What is kept, where, and for how long** | What | Where | How long | | --- | --- | --- | | The conversation | Your browser | 7 days, 15 turns, 5 earlier chats | | Counts: turns, tools, timings, tokens | The store | 35 days | | A question, 240 characters of its answer, the page, a city | The store, redacted | 30 days, the newest 3,000 | | A fresh question’s whole answer | The store, keyed by its words and page | 1 day | | A thumbs up or down, with its question | The store, redacted | 30 days | | A question for the real me, or a wish to talk | The store, and my phone | 30 days | | Rate-limit counts | The store, under a daily hash of the address | 10 s to 1 day, per window | | Anything that was said | Server logs | Never | Nothing is kept to know who someone is. What is kept is there to answer better, and each says when it goes. ```ts // conversations.ts and guestbook/ticket.ts /** Email addresses and phone numbers, blacked out. */ export function redact(text: string) { return text .replace(/[\w.+-]+@[\w-]+(\.[\w-]+)+/g, '[email]') .replace(/(?:\+?\d[\s().-]*){8,}\d/g, '[phone]'); } /* ... and who is asking, for the limits only: */ export function visitorKey(ip: string) { const day = Math.floor(Date.now() / 86_400_000); return createHmac('sha256', KEY) .update(`${day}:${ip}`) .digest('base64url') .slice(0, 16); } ``` Contact details are blacked out before anything is kept; the address is only ever a keyed hash that changes every day. ## Kept answers, keyed with care An answer to a fresh question is kept a day, so the next person to ask it gets it in about 25 ms, replayed a few words at a time, for no tokens. The key is the question as normalised, the page it was asked on, that page’s parts as the visitor’s page sent them, the time and weather showing, and a hash of the prompt and everything the agent knows. A change to the content or the prompt begins afresh. Two mistakes shaped that key. Kept by the words alone, an answer from one page was replayed on another. And because a page’s parts come from the visitor, a made-up list could have planted an answer for everyone; with the parts in the key, a made-up list is only ever answered to itself. An answer that acted on the page, or failed, is never kept. A kept “I’ve made it night” used to replay the words without the night. The vectors for retrieval are kept as well (30 days a passage, a day a question), under the embedding model’s name and the text, so neither is paid for twice. Next: a semantic cache, so “what’s your stack” and “what do you build with” share an answer, behind the same interface the exact cache uses, and precomputed answers for the suggested questions (in progress). ## Fast, and free until it’s reached for The chat costs a page nothing until someone reaches for it. The tab is a button with my portrait; the chat itself, 25 KB compressed in two files, is fetched when a pointer nears the tab or it takes focus, so the press finds it ready. Once opened it stays mounted, and opening it again is instant. While an answer streams, events are folded in once a frame however fast they arrive, so there is at most one render a frame; a finished answer is never drawn again while the next one streams; and the conversation is saved to the device when the browser is idle. On the server, the passages and their index are built when the function boots rather than inside the first question, every count and check is one round trip, and the obvious card goes out before the model is asked. The questions offered to tap were answered once, ahead of time, and ship with the site: tapping one asks no model, and an answer is served only while the passages and brief it came from are unchanged. What a visitor shouldn’t wait for is done when the site is built. A step before the build makes the passages’ vectors, only for the passages whose words changed (a hash of each says which), and keeps them in Atlas, removing the ones the site no longer has; a full index of 179 passages took 42 seconds, and an ordinary deploy embeds a handful. Without the database or the keys the step says so and the build goes on, and it can never fail a deploy. The chips’ answers are written the same way, by a script run between deploys, a few at a time within the free quota. Measured on 1 October 2026 on a production build on my laptop, against the real hosted model on its free tier, with a stand-in store on the same machine: 20 questions visitors ask, each asked once, before meaning was added to retrieval (it adds about half a second to a new question). The trip to the host and to the store adds to these, and isn’t measured here. **The latency budget, measured** | Step | Measured | Notes | | --- | --- | --- | | Find passages by words | 1 to 11 ms | In memory; under 0.2 ms to score a question once warm. | | Find passages by meaning | 0.52 s median | A new question; a repeated one is kept a day. | | An obvious card | 36 ms median | Inside the turn; 14 of 20 questions had something on screen within 0.3 s. | | One model request | 1.26 s median, 1.57 s at p90 | 30 requests; 1.5 a question. | | First words | 1.85 s median, 2.3 s at p90 | In the client, through the route. | | Whole answer | 2.15 s median, 2.85 s at p90 | In the client; 1.1 to 3.6 s inside the turn. | | A kept answer | About 25 ms to the first word | The whole answer replayed in 0.1 to 0.3 s. | | A command, off topic, or a try at its rules | Under 0.1 s | No model request. | | A fit check | 2.5 s | One request, 3,525 tokens in, 464 out. | | The chat’s own code | 25 KB compressed, 2 files | Fetched on the way to the tab, not with the page. | Medians and p90 over 20 live questions. The model is almost all of it; everything around it is milliseconds. ```ts // components/ask/Ask.tsx and AskPanel.tsx /* The chat itself is its own chunk, fetched as a hand nears the tab. */ const loadPanel = () => import('./AskPanel').then((m) => m.AskPanel); const AskPanel = dynamic(loadPanel, { ssr: false }); /* ... and while an answer streams: */ // Events are folded in once a frame, however fast // they arrive: at most one render a frame while the // words stream. let queue: TurnEvent[] = []; let frame = 0; const flush = () => { frame = 0; const batch = queue; queue = []; if (batch.length) update((t) => batch.reduce(fold, t)); }; await ask(body, abort.signal, (event) => { queue.push(event); frame ||= requestAnimationFrame(flush); }); ``` Its own chunk, warmed on the way to the tab; and however fast the tokens come, one render a frame. ## Seeing what it does Every answer carries its own trace, and anyone can open it under the answer: each part as a span (rewrite, retrieve, each model request, each tool), when anything first showed, the total, the tokens in and out, which tools ran, which provider answered and the prompt’s version. It is made of timings and names, never words. The same trace is folded into counts a day: questions, kept answers, limited, failed, stopped, fallbacks, answers from the pages, questions it couldn’t answer, off-topic replies, fit checks and leads; each part’s latency in buckets (under 250 ms, 500 ms, 1 s, 1.5 s, 2.5 s, 5 s, and slower); tokens; each tool’s runs and bad calls; which provider answered; and the prompt’s version. They’re kept 35 days. Logged: nothing that was said. The counts are numbers and the traces are timings. The questions live apart, redacted, for 30 days of review, and nowhere else. ```ts // lib/agent/ui.ts and telemetry.ts /** The visible trace of a turn: what ran, how long, nothing that was said. */ export type Trace = { version: string; model: 'primary' | 'fallback' | 'kept' | 'pages' | 'local'; /** Which provider answered. */ provider?: 'primary' | 'fallback'; steps: number; spans: { name: string; ms: number; ok: boolean }[]; firstMs: number; totalMs: number; tokens: { prompt: number; completion: number }; tools: string[]; }; /** Latency buckets, by upper bound in milliseconds. */ export const BUCKETS = [250, 500, 1000, 1500, 2500, 5000, Infinity] as const; ``` What a turn says about itself. The counters read these fields and nothing else. ## The review desk, and my phone The questions are for reading, so they have a desk: /ask/review, opened only with my key, holding nothing without it. It shows the conversations of the last 30 days, newest first, filterable by a lead, a fit check, a page or a thumbs down; the people who want me to get back to them, with a reply link; the questions it couldn’t answer, which are the pages I should write next; and the numbers: questions a day, what was kept, fit checks, leads, how fast the first words came, which tools ran and the tokens spent against the day’s budget. My phone hears the rest. A fit check the moment someone runs one (the role, how strong a fit, the verdict and where they are, linked to that conversation on the desk), once a conversation and role, thirty a day at most. Trouble, each once and a dozen a day at most: tokens past 80% or spent; every provider failing a question, urgently, and the chat back again; every key of a provider resting; the main provider failing while a backup takes the questions, naming which. And one push each evening at 8 pm, asked for by a scheduled workflow with its own key: the day’s conversations, fit checks, leads, unanswered and most-asked questions, thumbs down, résumé views, visitors, and how the budget and the providers held up, or a single line on a quiet day. ## Evals, and how they gate a change A prompt change is a code change, so it is tested like one. The prompt has a version (agent-12 today) that goes into every trace, splits the counts, and is part of the key kept answers are stored under, so a new prompt never serves an old answer and its numbers can be set beside the last one’s. 284 unit tests cover the engine with no model at all: the stream parser cut at every point, tool calls assembled across chunks and providers, the loop’s step cap and fallbacks, key rotation, the schema checks, markers and citations, every leak shape, the guards, the route’s refusals, the fit check’s validation, and retrieval, which fails CI if hit@5 on the eval set drops under 90%. The cases from an audit of the engine are among them. The eval set is 44 questions a visitor might ask, each with where the answer is on the site and words a right answer has to say, plus follow-ups, persona checks (does it say it’s an AI; does it hand over rather than promise) and tries at its rules. A script runs them against the real model and scores retrieval, tool choice, the answer’s facts, invented numbers, first person and persona. It costs tokens, so it runs by hand before a prompt change ships. The last run, on 1 October 2026: the tool a case expects for 21 of 23; the words a right answer needs in 24 of 26; no invented number in 43 of 44 answers (the one rounded 12,493 downloads to 12,000); first person in 40 of 44 (four led with a project rather than with I); and all 6 persona checks. Retrieval’s one miss is “Does he know AI?”, which now finds this page first: a fair answer the set doesn’t expect yet. The actions have their own matrix: more than 100 ways of asking, from each state of the page, each command against what should run and each look-alike question against nothing running, so an action firing on a question is a failing case, not a surprise. The providers are checked for real before a deploy, by hand. One script asks every key once for a token and says which is revoked, suspended or out of quota, by its provider, its place in the list and its last four characters. A fire drill breaks the main provider’s keys, then the second’s, then every one, and watches the next provider answer, a refused key rest and not be asked again, and the end of the line fail in under a second. Its first run found two keys suspended, a backup whose model was no longer free, a bad key answered with a 400 that rested nothing and stopped the turn, and a stalled backup that held a question for 25 seconds; each is fixed, with a test. What I’ll add next: a model as judge of faithfulness (does each sentence follow from the passage it cites), checked against a small set I score by hand before it is trusted with a large one; regression cases drawn from the review desk, where every question it couldn’t answer or that got a thumbs down becomes a case; each tool’s success rate from the counts; a latency objective (a card within 0.3 s and the whole answer within 3 s at p90) checked on every deploy; and the evals in CI on a small model, gating any change to the brief. **The last run** | Check | Result | How | | --- | --- | --- | | Retrieval, hit@1 | 26 of 32 | The expected page is the first passage (words and meaning). | | Retrieval, hit@5 | 31 of 32 | In the top five, the offline gate. | | Retrieval, MRR@5 | 0.87 | Mean reciprocal rank of the expected page. | | Tool choice | 21 of 23 | Ran a tool the case expects; off-topic cases answered without the model. | | Answer checks | 24 of 26 | Says the words a right answer must. | | No invented numbers | 43 of 44 | Every number is one the site says. | | First person | 40 of 44 | Speaks as me, not about me. | | Persona | 6 of 6 | Says it’s an AI; hands over instead of promising. | 44 questions against the real model, spaced out to stay under the free tier’s limit a minute; retrieval scored offline over the 32 with an expected page. ```ts // scripts/agent/eval.ts · invented numbers const said = ( knowledge() + '\n' + corpus().map((c) => c.text).join('\n') ).toLowerCase(); /** Numbers in an answer that the site never says. */ const invented = (text: string, also = '') => ( text .replace(/\[\[source:\d+\]\]/g, '') .match(/\d[\d,.]*\d%?|\d%?/g) ?? [] ).filter( (n) => !said.includes(n.toLowerCase()) && !also.includes(n.replace(/,/g, '')), ); ``` Crude on purpose: any number the site never says is a failure, whatever the model meant by it. ## What broke, and how it was fixed Most of what makes it dependable came from something going wrong. Each of these is now a test. Memory starved by the prompt budget: facts and passages were fitted first and the conversation got what was left, so follow-ups forgot. Now the last four turns come first. Actions claimed but not run: the model said “I’ve switched to night” without calling the tool. Now a claim is checked against what ran, and corrected in a line. Intent patterns firing on questions: summer changed the weather, a day job made it day. Now an action needs a command, and a card before the model needs an explicit ask. A kept answer poisoned by the page’s parts, and one that replayed an action’s words without the action. Now the parts and state are in the key, and answers that act are never kept. The breaker counting tries rather than questions, so one question with tired keys could trip it; the in-flight counter drifting when a function died; tokens not counted for stopped turns; sessions dying at midnight. Each is fixed above, where it lives. Tool calls leaking into the words, first as “suggest: [...]” lines, then as bold names, tags, fences and bracketed markers. The filter now holds every shape from where it could begin. And on a phone: the notes of what an answer showed and did weren’t going with the next question, so a follow-up about what it had just done had nothing to go on; a stream cut short looked finished; a conversation kept in an old shape broke the chat on the next visit; and the chat followed the words even when the reader had scrolled up. All four are fixed, and the chat now takes the whole height between the phone’s safe areas. ## Trade-offs I considered and rejected An agent framework, or a vendor’s SDK. Either would have named a vendor in the code, added weight, and put the stream’s shape in someone else’s hands. The loop is a few hundred lines, and owning it is what lets a card arrive mid-answer and a second provider take over mid-turn. A vector database. The whole corpus is about 150 passages; words in memory score one in under a fifth of a millisecond, and vectors live in the same store as everything else. A database becomes worth it at thousands of pages or many sites, which is the module, not this. A model to rewrite follow-ups into standalone questions. It would cost a request before every follow-up; borrowing the last question’s words and projects costs nothing and is right often enough, with the model’s own memory of the turns behind it. Server-sent events to the browser. The question is a POST with a session header, which EventSource can’t send; NDJSON over a streamed fetch keeps the request honest and the parsing one line at a time. Answers as JSON. Streaming words with cards as tool calls lets text start at once; JSON is kept for the one place its shape matters, the fit check. Many agents. See routing, below: not until the evals say one agent has stopped being enough. ## Running on free quotas, gracefully It runs on free tiers today, so the strategy is to spend tokens only where a model adds something, and to degrade in steps a visitor can still use. Zero tokens first. Commands (“make it night”), off-topic questions, tries at its rules and kept answers never reach a model. Precomputed answers for the suggested questions are in progress. Then fewer tokens. 17 tools on offer rather than 31; the second provider offered only the tools most answers need; the facts written once, compactly; passages capped at 3,600 characters. Then limits. Six windows of rate limits and a daily token budget of 1,500,000, set against the free allowance, not the paid one. Then another quota. Keys in turn cover a failing key, not a bigger allowance, since keys in one project share one quota. The second provider has a quota of its own, and the order of who is asked is fixed and ordered. Then no model. Past every limit, or with both breakers open, the pages answer: the obvious card, links to the best passages, and a note to the real me. The capacity maths is short. An answered question averaged about 8,000 tokens over 1.5 requests, and more now that memory comes first. Back to back, about 13 questions a minute, the first provider refused after eight; spaced to five a minute, 25 in a row went through. So a free tier carries a few questions a minute, and at the budget’s 1,500,000 tokens, about 185 answered questions a day, before kept answers and commands. On a paid tier the budget, not the provider, becomes the limit, and that is a choice: at an assumed $0.10 a million tokens in and $0.40 out, 1,000 conversations of three questions cost about $2.50. ## What breaks first at scale, and the plan Beyond a portfolio’s traffic, here is what gives first and what I’d do about each. The model’s tokens a minute go first; the fix is a paid tier on the first provider, with the second on its own quota taking over when the first can’t answer. Then the store: an answered question is 46 to 53 store commands (the limits alone are 12), so a plan is sized in commands, not memory. One script per question that counts every window and returns the verdict, and one write of the day’s counts per turn, would cut that by more than half. Concurrency comes last: 24 answers in flight at about 2 seconds each is several hundred a minute, far above the 60 a minute allowed, while cold starts matter on a quiet site, where the index is built as a function boots. Under real load: rate limits and a firewall rule at the edge for the chat’s route, so a flood never wakes a function; the semantic cache; the prompt’s fixed part first and the per-question parts after it, so a provider’s prefix cache can reuse the roughly 4,000 tokens every question shares; and a short queue with a place in line instead of a refusal, and past it, answers from the pages alone. Cost is tokens, and tokens are measured. A conversation of three questions is about 24,000 tokens in and 320 out, before the history a follow-up carries. At an assumed $0.10 a million in and $0.40 out, the list price of a small hosted model, that is about $2.50 per 1,000 conversations; at $0.30 and $2.50, about $8. **Capacity at today’s limits** | What | Number | How it was found | | --- | --- | --- | | Tokens an answered question | 8,044 in, 107 out | Average of 20 live questions, before memory came first. | | Model requests a question | 1.5 | Same 20; one or two each. | | Store commands an answered question | 46 to 53 | Counted on a stand-in store. | | Answered questions a day on the budget | About 185 | 1,500,000 ÷ about 8,150. | | Questions a day, everyone | 1,500 | The guard; kept answers count, but cost no tokens. | | Answers in flight | 24 | About 11 a second at 2.15 s each; the 60-a-minute limit binds long before. | | Cost per 1,000 conversations | About $2.50 to $8 | Three questions each, at the two assumed prices above. | The first thing to fill is the day’s token budget, which is a choice; after it, the provider’s tokens a minute, which isn’t. ## What’s next: routing Today a question is routed in two layers: a deterministic layer of patterns that answers commands and off-topic questions with no model and plans the obvious cards, then one model with a subset of tools chosen for that question. It is cheap, fast and easy to test, and it is the part most likely to be outgrown. A learned router. An embedding-similarity or small-classifier router choosing among five paths: the command path, a precomputed answer, retrieval-only answers, the full agent, and the fit-check pipeline. It would catch the phrasings the patterns miss, at the cost of a vector lookup or a small model call on every question. Model routing. A small, fast model for the simple questions and a larger one for a fit check or a long job description, chosen by expected cost and latency. The second provider already shows the shape of it: a different model, fewer tools. Many agents: a planner with specialist workers (retrieval QA, a fit analyst, a site operator) and a verifier over them. Deliberately not used yet: it adds latency to every answer, more calls against free quotas, more ways to fail, and evals that are much harder to trust. What would justify it is the evals showing a single agent plateauing on multi-step tasks, which is most likely when the voice agent starts operating the site for someone. However it is routed, it is judged the same way: precision and recall per route on real questions from the review desk, labelled by hand; each route’s place on a latency and cost Pareto curve; and a regression gate, so no routing change ships that sends a known question down a worse path. ## Voice, next Next, it speaks. A pill grows out of the tab with a soft orb in it that moves with the real audio, captions of both sides above it, the same suggested questions, and a way to open it into the text chat with the conversation carried over. Voice is a transport, not a second brain. A realtime speech-to-speech model runs over a socket the browser opens itself, with a short-lived, single-use token my server mints and locks: the brief, the tools and the voice are sealed into it, so a modified page can’t change them. Its tool calls run on my server through the same registry, with the same session, guards and events, folded into the same conversation. The voice is a stock one, never a clone, and the pill says it’s an AI. A spike put numbers on it: first audio 1.2 to 1.5 s after the visitor stops talking, and a site action visible 1.1 to 1.9 s after asking for it. It also showed what not to do: with the text brief and all 29 tools, the first audio took 2.8 to 5.5 s, because the model waited for its own card; a brief for speech (answer first, out loud) and 16 tools brought it back to about 1.5 s. The guards are the text chat’s plus minutes: three minutes a session, a few sessions a visitor a day, a cap on sessions at once under the provider’s (the tier I tested took six and refused the seventh), a daily budget of minutes, and a breaker. When voice is full it falls back to the browser’s own speech recognition with the same text turn read aloud, then to text. Anything that leaves the page waits for a spoken or tapped yes. ## As a module The goal is for any site to have one that knows its own pages, shows its own components and moves around it for its visitors. 10xAnswers, my chat component on npm, was the first step; this will ship as a package of its own. Most of it is already shaped for that. The core yields events and knows nothing of HTTP; the model is an address in the environment; tools are entries in a registry; the store is a handful of commands behind one file; the chat draws components by name from props; and anything on a page can open it with a question, through one event. What has to be made general first: the corpus, which reads this site’s markdown today and needs a content adapter (a folder of markdown, a sitemap to crawl, a CMS); the persona, since the brief is written as me; the intent patterns, which know my projects; the store, which speaks one provider’s REST API; the page actions, which rely on this site’s labels, night and weather; and the chat’s styles. The sketch below is where it’s heading, not a published API. ```ts // Planned · not published // Planned: the shape it's heading for, not a // published API. import { createAgent, ndjson } from './agent'; const agent = createAgent({ persona: { name: 'Ada', away: true }, content: markdownFolder('./content'), tools: [...defaultTools, showProduct, bookDemo], store: redisStore(process.env.REDIS_URL), model: fromEnv('ASK_MODEL'), limits: { visitorPerDay: 80, tokensPerDay: 1_500_000 }, }); // The chat's route: one JSON event a line. export const POST = ndjson(agent); // A voice transport over the same brain, later. export const voice = realtime(agent, { secondsCap: 180 }); // The chat, drawing your own components by name. ``` A sketch, labelled as one: each piece maps onto what exists today (the registry, the store, the NDJSON transport), made general for someone else’s site. ## What I learned Most of the speed comes from not asking the model. A pattern that knows what a question obviously wants puts it on screen in milliseconds, and the model writes around it. A small model is unreliable in small ways: a tool call written out as words, a citation as [2], a long dash, a claim of something it didn’t do. So what it writes is a stream to be parsed and checked, not text to be trusted, and each of those has a test. The limits are the product. An assistant on a public page is only as good as what happens when it’s abused, rate-limited, out of tokens or talking to a broken endpoint; here each of those has an answer a visitor can still use. ## Next Precomputed answers for the suggested questions, the action eval matrix, voice on the same brain, a judge in the evals, routing when the evals ask for it, and the module. --- # Rajveer Singh, résumé (CV) > I build AI products end to end: the interface, the backend, and the infrastructure that keeps them running. The PDF: https://rajveers.com/resume.pdf ## Work Experience ### IABTM · Full Stack Engineer Intern Remote · Sep 2025 – Present Re-architected production media infrastructure to S3 + CloudFront. Containerized Node.js services on EC2 and engineered a CI/CD pipeline via GitHub Actions, reducing release cycles to under 15 minutes and cutting storage costs by 40%. Engineered a real-time group chat/voice layer (Socket.IO, Mediasoup SFU) with sub-200ms latency, and designed an autonomous GenAI editorial pipeline publishing 30+ curated stories daily with zero manual review. Built a parallel chunked S3 upload service, lifting success rates to 99.4%, and integrated Shopify Storefront GraphQL & Resend for seamless webhooks, checkout flows, and lifecycle email campaigns. ### Modulus · Full Stack Developer Intern Remote · Apr 2025 – Aug 2025 Designed and built the product UI from scratch using React/TypeScript, creating a reusable design system with audience-aware UX decisions that powers 4+ surfaces; supported user-growth trajectory to 1,275+ MAUs. Integrated 30+ REST endpoints and reduced critical page load by 66% (3.2s to 1.1s) via route-based code splitting and list virtualization, achieving 95+ Lighthouse scores and instrumenting full-funnel PostHog analytics. Shipped Razorpay (payments) and Cal.com (scheduling) integrations end-to-end with idempotent webhook handlers, successfully processing Rs. 10.8K in platform transactions at 99.9% uptime. ## Technical Skills Language&Frontend: TypeScript, JavaScript, Python, SQL, React, Next.js, Recoil, Tailwind CSS, Framer Motion, Plasmo Backend: Node.js, Express.js, GraphQL, REST APIs, WebSockets (Socket.IO), Mediasoup, Microservices, BullMQ, Jest Databases & AI: PostgreSQL, MongoDB, Redis, RAG, Qdrant Cloud & DevOps: AWS (EC2, S3, CloudFront, ECS, Bedrock), Docker, GitHub Actions (CI/CD), PM2, Nginx, Puppeteer ## Projects ### Safire (Safe DM) · MERN, TypeScript, Plasmo, RAG, Gemini, Redis, Docker, Puppeteer GitHub · Live Demo Architected a real-time LLM harassment detector with RAG (Gemini API), fronted by a React/TypeScript Chrome extension leveraging session-aware state management for inline content alerts. Built a Dockerized Node.js/Express microservice backend with a custom load balancer and Redis read-through cache, drastically cutting average detection latency by 60% (480ms → 190ms). Automated tamper-proof evidence reporting via Puppeteer, generating cryptographically verifiable PDFs of harassment instances stored securely on Cloudinary with 100% data integrity. ### HyperPersona · Node.js, TypeScript, Drizzle, PostgreSQL, Qdrant, Redis, BullMQ, Gemini, AWS ECS/Bedrock GitHub Architected an agentic recommendation engine as service in a 24-hour sprint with a 4-person team; built a 3-plane microservice backend in Node.js + TypeScript with PostgreSQL as system of record, Qdrant for 3,072-dim vector memory, and Redis-backed BullMQ for async embedding workers. Implemented an Agentic Cognitive Architecture (ACE) with Chain-of-Verification RAG over Gemini 2.5 Flash, deployed on AWS ECS/ECR with Amazon Bedrock + Strands Agents; selected as Top 0.1% Finalist at Technoverse 2.0. ### 10xAnswers · React, Recoil, TypeScript, Node.js, Express, Gemini GitHub · NPM Shipped a reusable React AI chatbot SDK on NPM with 3,200+ downloads, engineering zero-config integration via Recoil to deliver 100% context retention across multi-turn conversations. Built an Express inference proxy bridging the Gemini API with retry logic and graceful timeouts, sustaining 99.5% uptime and 120ms p50 response latency under high platform load. ## Achievements & Leadership Impact: Shipped 5+ full-stack products across internships and OSS, serving 5,000+ end users and 3,200+ NPM downloads; architected GenAI features and AWS infrastructure in production. Hackathons: Winner of 4 national-level hackathons including Code Kshetra 2.0 (15,000+ participants); Finalist at 4 others including Technoverse 2.0 (22,000+ participants). Open Source & Leadership: Top 35 mentor in GSSoC; 16+ PRs to high-impact OSS (cms/100xDevs); Vice Chairperson, Web Dev at IEEE MSIT - branch of IEEE, the world's largest technical society (400K+ members, 160+ countries). ## Education Maharaja Surajmal Institute of Technology (GGSIPU) Delhi, India Bachelor of Technology in Electronics & Communication Engineering 2023 – 2027 CGPA: 8.68 / 10.0 --- # The lake, live in WebGL > The home page's lake, full screen: painted in code and simulated on the GPU. Tap the water, zoom in, or sit on the jetty and fish, with your phone as the rod over WebRTC. Try it at https://rajveers.com/pond. ## How it works Steer the hook to a fish and wait for a bite, then yank it out: move and click with a mouse, drag and tap with a finger, or point your phone and lift it sharply. Five kinds swim here. Arrows and the space bar work too. - **One lake, a real camera.** This is the same painted lake, not a new scene. The camera drops to the fisherman, turns the way he faces and tips up into his view: every pixel’s ray is followed down to the water on the GPU, and the creatures are drawn through the same perspective, each as large as its distance makes it. - **The lake’s own fish.** The fish you watched from above are the ones you catch, with a few more let into the water in front of the jetty. They come to look at the lure, and one takes it for a second: that is the bite, and a yank then always lands it. A yank at nothing sends them off. - **Something else on the hook.** Now and then a fish comes up with something about me: what I do off-screen, what I made before any of this was work, the work I am proudest of, the character I like and why, what I am listening to right now, and once in a long while how I got my job. Each goes in your creel, which this browser keeps for next time. - **Your phone as the rod.** A QR code and a six letter code open a room on two independent servers, PeerJS’s public one and our own Cloudflare relay, and the phone takes whichever answers. From then on its aim and its lift go straight to the page over WebRTC, usually a few milliseconds apart. The aim comes from the phone’s whole orientation, so it works held like a remote or upright, and a lift is the aim rising fast. One phone to a lake at a time. ## How it's made The lake is not a picture or a video. It is painted in code and runs live in WebGL2, right in your browser. - **Painted in a worker.** Every tree, reed and lily pad is drawn by a script when the page opens, off the main thread, so the page still scrolls while it paints. Until then you see a blurred preview smaller than a kilobyte, inlined in the HTML. - **Real water.** A wave simulation on the GPU spreads the ripples when you tap, and the light on the lake floor is worked out for whatever part you are looking at, so it stays sharp as you zoom. - **Sharper up close.** Zoom in and the painting is redone at a higher resolution, so the details hold instead of blurring. - **Made once a visit.** The painting is kept, so coming back here, or to the home page, is instant. - **Asleep when unseen.** The loop stops when the lake is off screen or the tab is hidden, and a canvas this big caps its pixel density at 1.5x to spare the GPU. - **Sounds made on the spot.** Every splash, croak and quack, and the radio's music, is synthesised with Web Audio. There are no sound files. --- # The shelf > The books on my desk: what I'm reading now, what I've read, and what I'd hand to you. ## Reading now - The Design of Everyday Things, Don Norman - Designing Data-Intensive Applications, Martin Kleppmann ## Read - The Pragmatic Programmer, David Thomas - Shoe Dog, Phil Knight - Steal Like an Artist, Austin Kleon - Sapiens, Yuval Noah Harari ## Would hand to you - Refactoring UI, Adam Wathan - Atomic Habits, James Clear - Deep Work, Cal Newport --- # Rajveer Singh, the plain version > Everything on the site as facts: work, projects, record, stack, contact. ## Experience - **IABTM**, Software Engineer (internship), Sep 2025 to now. Building the core of a US self-improvement platform. Rebuilt its live voice rooms on an SFU, rebuilt its Stripe billing so no one is charged twice, and built the daily AI pipeline (Gemini) that curates what each member reads. Also built the email and the team’s admin tools, and runs the AWS. - [Inside the voice rooms and billing](https://rajveers.com/projects/iabtm) - **Wii Hop**, Software Engineer (contract), Jul 2025 to May 2026. Built most of the web app for a US local-deals startup: deals on a live map, table booking, events, loyalty streaks, and heists that turn a night out into a game (React, TypeScript, Leaflet). On the backend, wrote the parts where money and stock change hands, from the heist item shop to catering orders priced on the server. Also built screens for the Expo app. - [Inside the deals, bookings and loyalty](https://rajveers.com/projects/wiihop) - **Modulus**, Frontend Engineer (internship), Apr 2025 to Aug 2025. Built the product’s interface from the first screen: cohort pages, booking and payment on Razorpay, scheduling through Cal.com, and live sessions with chat running beside the stream (React, TypeScript, hls.js). Wired PostHog funnels so the team could see where people dropped off before they enrolled. - [Inside booking and live sessions](https://rajveers.com/projects/modulus) ## Every project ### [IABTM](https://rajveers.com/projects/iabtm) The core of a US self-improvement platform: live voice rooms, Stripe subscriptions, a daily AI pipeline and the AWS it runs on. - internship, Sep 2025 to now. Software Engineer, internship, remote - Stack: Next.js, React, Express, MongoDB, Redis, Socket.IO, mediasoup, Stripe, Shopify, Gemini, PostHog, AWS - Live: https://iambetterthanme.com - 6 publishers, curated daily; 7 jobs that run themselves; 25 admin tools for the team ### [HyperPersona](https://rajveers.com/projects/hyperpersona) An agentic recommender on AWS Bedrock that four of us built in a week: top 0.1% of 22,000+ at Cognizant Technoverse. I built the storefront and event SDK. - hackathon, One week, May 2026. The storefront and everything the browser sends. Four of us - Stack: React 19, TypeScript, FastAPI, DynamoDB, OpenSearch Serverless, Bedrock, Redis - Live: https://hyperpersona-web.vercel.app ### [Safire](https://rajveers.com/projects/safire) A harassment shield for LinkedIn DMs that won HackWIE 3.0 and Code Kshetra 2.0. I built the backend: the API, moderation failover and evidence reports. - hackathon, Jan to Feb 2025, polished through July. The backend, and the extension's popup. Four of us - Stack: Plasmo (MV3), React, Express, MongoDB, Redis, Gemini, Puppeteer, Next.js - Live: https://safire-five.vercel.app ### [Wii Hop](https://rajveers.com/projects/wiihop) Most of the web app for a US local-deals startup, from deals on a live map to heists that make a night out a game, plus the backend where money moves. - contract, Jul 2025 to May 2026. Software Engineer, contract - Stack: React, TypeScript, Leaflet, Node, Prisma, Expo - Live: https://beta.yohop.com ### [Modulus](https://rajveers.com/projects/modulus) Interview prep for consulting, product and finance roles. I built its product from the first screen: programs, booking, payment and live sessions. - internship, Apr to Aug 2025. Frontend Engineer Intern, remote - Stack: React, TypeScript, Vite, React Router, TanStack Query, Tailwind, hls.js, Firebase, PostHog - Live: https://www.trymodulus.com - 47 routes; #8 of the week on Peerlist ### [10xAnswers](https://rajveers.com/projects/10xanswers) A drop-in React chat widget for AI answers, built and published alone: 12,493 downloads across 55 releases on npm. - side project, Dec 2024 to May 2025, maintained since. Solo - Stack: React, Recoil, Vite, Express, Gemini - Live: https://10x-answers.vercel.app ### [Kernel](https://rajveers.com/projects/kernel) Anonymous chat rooms with peer-to-peer voice, where the server only relays the handshake. The take-home for the role I have now. - take-home, Two days, Sep 2025. Solo - Stack: React, Socket.IO, WebRTC, MongoDB - Live: https://kernelchat.vercel.app ### [The lake](https://rajveers.com/projects/lake) This site's hero: a lake painted in code, with water simulated on the GPU, creatures that make ripples, and fishing with your phone as the rod. - side project, 2026. Solo - Stack: WebGL2, GLSL, Web Workers, OffscreenCanvas, WebRTC, Cloudflare Workers, Durable Objects, Device sensors, TypeScript, React - Live: https://rajveers.com/pond In progress, live and still being built: [The AI me](https://rajveers.com/projects/ai-me). ## Track record ### Hackathons - [HackWIE 3.0](https://dorahacks.io/hackathon/826/winner), 1st place, IEEE MSIT · 2025 - [Code Kshetra 2.0](https://x.com/RajveeerrSingh/status/1894034885855138202), 1st place, out of 15,000+ registrations, Geek Room · 2025 - [Technoverse 2.0](/projects/hyperpersona), Top 0.1% of 22,000+, Cognizant · 2026 - UPSTART, Second runner-up, E-Summit ’25, E-Cell MSIT · 2025 - SheBuilds × OpsTree, Award, 2025 - MSC Hack-It-Up, Special mention, IGDTUW · 2025 - Excalibur, Techspardha, Final round, NIT Kurukshetra Winner or finalist: 7× ### Open source - [10xanswers](https://www.npmjs.com/package/10xanswers), An AI chat for React sites, 55 releases, npm - [axios-docs #217](https://github.com/axios/axios-docs/pull/217), A fix in the official docs, merged, axios - [100xDevs CMS](https://github.com/code100x/cms/pulls?q=is%3Apr+author%3Arajveeerr+is%3Amerged), 4 pull requests, merged, code100x Downloads on npm: 12,493 ### Stack - Interface, React, Next.js, TypeScript, Tailwind, Expo, ×5 - Server, Node, Express, Socket.IO, Prisma, ×4 - Realtime, WebRTC, mediasoup, hls.js, ×3 - Data, MongoDB, Redis, Firebase, ×3 - Money and AI, Stripe, Razorpay, Shopify, Gemini, ×4 - Cloud and graphics, AWS, Cloudflare Workers, WebGL2, GLSL, ×4 Tools in production: 23 ## Contact - Email: rajveergreets@gmail.com - Book a call: https://cal.com/rajveer - Résumé: https://rajveers.com/resume.pdf - Elsewhere: [github](https://github.com/rajveeerr), [x](https://x.com/rajveeerrsingh), [linkedin](https://www.linkedin.com/in/rajveeerr), [peerlist](https://peerlist.io/rajveeerr), [npm](https://www.npmjs.com/~rajveeerr), [devfolio](https://devfolio.co/@rajveeerr), [producthunt](https://www.producthunt.com/@rajveeerr), [holopin](https://holopin.io/@rajveeerr), [wakatime](https://wakatime.com/@rajveeerr) - Open to conversations.