Blog3 min read
Voice rooms on an SFU
How IABTM's live audio rooms moved from a peer-to-peer mesh to a mediasoup SFU: a router per room, five-speaker stages and a worker that restarts on its own.
IABTM's members meet in live audio rooms and on stages, where five people speak and everyone else listens and raises a hand. The first version was a peer-to-peer mesh: every speaker sent their audio straight to every listener. That is simple and it works for four people. It does not work for forty, because a speaker's upload grows with every listener in the room.
So I rebuilt the rooms on an SFU, a selective forwarding unit, with mediasoup. Each speaker sends their audio once, to the server, and the server forwards it to everyone listening. A speaker's upload stays the same whether two people are listening or two hundred.
One worker, a router per room
mediasoup runs its media in a separate C++ worker process, and Node drives it. IABTM runs one worker, inside the API's own process under PM2, with UDP ports 40000 to 49999 open for media. Each room gets its own router, which knows the room's codec (Opus at 48 kHz) and forwards between the people in it.
Every member gets two transports: one to send on, one to receive on. They prefer UDP and fall back to TCP, and the server announces its public address so a phone behind a home router can still reach it. Where a network won't let UDP through at all, coturn on the same box relays the media over TURN.
Who is speaking
A stage has room for five speakers. The sixth gets told the stage is full and keeps their hand raised, and the host brings people up as others step down. If the host leaves, the role passes to the next person on stage, so a room never ends up with no one able to run it.
The ring that lights up around whoever is talking comes from the server, not the browser: each room has an audio level observer that checks every speaker's volume every 800 ms and names the loudest one above -80 dB. One source of truth means every listener sees the same person lit.
Joining is checked, every time
Media is expensive to send to the wrong person, so a socket is checked before it can do anything. The token is read from the member's cookie when the socket connects, before any event fires, and a socket whose token expires later is disconnected rather than left open. Joining a room, sending audio and receiving it each check that this member belongs in this room. Tokens are signed with a secret that can be rotated without logging anyone out: the current and the previous secret are both accepted until the old one is retired.
When the worker dies
A media worker is native code, and native code can crash. When it does, the rooms on it are closed cleanly, the people in them are told, and a new worker is started on its own. Chat, notifications and the rest of the API never notice, because the worker is a separate process and its death is an event, not an exception.
What breaks first at 10× the traffic
The SFU lives in the API's process, which is why the API runs as one PM2 worker today: the rooms hold their state in memory. At ten times the listeners, realtime moves to its own machine, rooms are spread across workers by a hash of the room's id, and the API scales on its own. Chat already fans out across processes over Redis, so that part is ready. The open question is how many listeners one box carries, which is the next load test to run.
Keep reading
A load balancer from scratch, A/B tested on live trafficA load balancer in Node for five moderation replicas: four strategies, a circuit breaker and health check each, failover, and a load test that proves the spread.Read
A lake that never makes the page waitHow the WebGL lake at the top of this site paints in a worker, compiles its shaders alongside, holds to a pixel budget, and costs nothing when you can't see it.Read
Evals that keep an AI agent honestHow the AI agent on this site is tested like code: 44 questions against the real model, retrieval gated in CI, a versioned prompt and a trace under every answer.Read
Shipping to one box without a staging serverCI/CD for a live EC2 box with 3,000+ members and no staging server: an SSH push with pinned keys, one deploy at a time, a 90 second health gate and a rollback.Read
A chapter site that keeps itself current from InstagramA FastAPI service that reads an IEEE student chapter's Instagram, has Gemini turn each post into a structured event, drops reposts and feeds the chapter's PWA.Read