Blog2 min read
Shipping to one box without a staging server
CI/CD for a live EC2 box with 3,000+ members and no staging server: an SSH push with pinned keys, one deploy at a time, a 90 second health gate and a rollback.
IABTM, the platform I build and run, lives on one EC2 t3.medium: nginx in front terminating TLS and limiting each client to 30 requests a second, the Node API and the Next.js app under PM2, coturn beside them for voice behind strict networks, and media on S3 behind CloudFront. 3,000+ members use it, and there is no staging box.
So the deploy has to be the safety net. Every pull request runs CI; every merge to main ships itself. This is how it gets from a merge to live without a bad build ever staying up.
No credentials on the server
The CI runner pushes the merged commit over SSH to a branch on the server, ci-deploy, with the host's keys pinned. The server never holds a GitHub token, so a compromised box can't reach the repository.
One deploy at a time
The API and the web app deploy from separate repositories, and two merges can land a minute apart. Each deploy takes a lock with flock and waits up to half an hour for it, so two can never run over each other.
Only what changed
Dependencies are installed only when the lockfile changed. Most deploys are code, and skipping a clean install is the difference between seconds and minutes on a 2-vCPU box that is also serving traffic.
Build beside the live one
The Next.js app builds into its own folder while the live one keeps serving. When the build is done the two are swapped in one move and the previous build is kept, so there is never a moment where the site is half built.
The health gate
The API restarts under PM2 and then has 90 seconds to report healthy, checked every 3 seconds, and healthy means the database answered, not only that the process is up. The web app has to answer with a 200 in the same window.
Rolling back
If any step fails, a trap rolls back the code, the dependencies and the process to the commit that was live. Every commit that was ever live is kept under its own ref with the time it shipped, and any change made by hand on the server is saved as a patch before a deploy overwrites it, so nothing is lost and there is always a known place to go back to.
What breaks first at 10× the traffic
One box is one place to fail. The next step is two behind a load balancer, deployed one at a time, each taken out of rotation until it passes the same health gate; and voice moves to its own machine, because the SFU holds each room's state in its process, which is why the API runs as one process today.
Keep reading
A load balancer from scratch, A/B tested on live trafficA load balancer in Node for five moderation replicas: four strategies, a circuit breaker and health check each, failover, and a load test that proves the spread.Read
A lake that never makes the page waitHow the WebGL lake at the top of this site paints in a worker, compiles its shaders alongside, holds to a pixel budget, and costs nothing when you can't see it.Read
Evals that keep an AI agent honestHow the AI agent on this site is tested like code: 44 questions against the real model, retrieval gated in CI, a versioned prompt and a trace under every answer.Read
Voice rooms on an SFUHow IABTM's live audio rooms moved from a peer-to-peer mesh to a mediasoup SFU: a router per room, five-speaker stages and a worker that restarts on its own.Read
A chapter site that keeps itself current from InstagramA FastAPI service that reads an IEEE student chapter's Instagram, has Gemini turn each post into a structured event, drops reposts and feeds the chapter's PWA.Read