# A load balancer from scratch, A/B tested on live traffic

> A load balancer in Node for five moderation replicas: four strategies, a circuit breaker and health check each, failover, and a load test that proves the spread. Published 4 October 2026.

Safire screens LinkedIn DMs for harassment, and the model behind it ran on free tiers. Free tiers run out, usually in the middle of a demo. So moderation moved into its own stateless service on five replicas, and I wrote the load balancer in front of them in Node.

*Diagram:* Where the balancer sits: behind a verdict cache, in front of five moderation replicas.

## Four ways to pick a server

- **Least connections:** the replica with the fewest requests in flight.
- **Round robin:** each in turn.
- **Weighted:** a weighted draw, where a replica's weight rises with each success and falls with each failure.
- **Fastest response:** the one whose recent requests came back quickest, smoothed so one slow answer doesn't swing it.

Whichever one picks first, the rest of the pool is the failover, each tried once, so a request only fails when every replica has.

From `server/src/api/moderation/detect-harassment.js`:

```js
const strategies = {
    'weighted': getWeightedServer,
    'least-connections': getLeastConnectionsServer,
    'fastest-response': getFastestResponseServer,
    'round-robin': getRoundRobinServer
};

// …

const pickStrategy = () => {
    const { strategy, abStrategies, abSplit } = config.loadBalancer;
    const name = strategy === 'ab' ?
        (Math.random() < abSplit ? abStrategies[0] : abStrategies[1]) :
        strategy;
    return strategies[name] ? name : 'least-connections';
};
```

## Breakers and health checks

A replica that is failing should stop getting traffic before it drags the others down. Each one sits behind its own circuit breaker: ten failures open it for thirty seconds, and then a single half-open request decides whether it closes again. Separately, a health check every 50 seconds marks a replica unhealthy after five failed checks and healthy again after two passes in a row, so one blip in either direction doesn't flip it.

Every response says which replica served it and which strategy chose it, in two headers. That's what makes the rest measurable.

## Proving it spreads the load

The balancer carries its own load test: `/test-servers` fires requests for 30 seconds, 20 at a time, and reports throughput, p95 and p99, and how many each replica served, read from those headers.

Against five stubbed replicas, 200 requests landed 40, 40, 40, 40 and 40. With two replicas failing, all 200 still succeeded: the two breakers opened after 10 to 12 calls each, and the other three took the load.

Then under real load, on my laptop, the balancer's code unchanged, in front of five local stand-ins: two answering in 40 ms, one in 80, one in 160, and one failing every call while still passing its health checks. With 50 connections for 30 seconds, least connections carried 788 requests a second at a p50 of 46 ms and a p95 of 166 ms. Round robin managed 402 at a p95 of 358 ms, because it keeps handing the slowest replica a full quarter of the traffic. Across 30 runs, not one request failed. The balancer itself adds under a millisecond at p50.

## A/B testing on live traffic

Which strategy is best depends on the traffic, so the balancer can answer that itself. With `LB_STRATEGY=ab` each request is given one of two strategies, and each one's share, success rate and p50 and p95 latency are kept side by side at `/metrics/strategies`.

## What breaks first at 10× the traffic

Every DM is its own request today. The plan is to stream moderation messages through Kafka, settle most of them in a split second with jev before the model ever sees them, and batch messages in the extension. The verdict cache, one Redis today, shards by a hash of the message.
