# Safire

> A harassment shield for LinkedIn DMs that won HackWIE 3.0 and Code Kshetra 2.0. I built the backend: the API, moderation failover and evidence reports.

- Kind: hackathon
- Role: The backend, and the extension's popup
- When: Jan to Feb 2025, polished through July
- Team: Four of us
- Stack: Plasmo (MV3), React, Express, MongoDB, Redis, Gemini, Puppeteer, Next.js
- Live: https://safire-five.vercel.app
- safire-five.vercel.app: https://safire-five.vercel.app
- Demo: https://vimeo.com/1059208124
- Code: https://github.com/rajveeerr/Safire

Harassment in LinkedIn DMs usually gets read before anyone can report it. Safire screens every incoming message, hides abuse behind a warning, flags a sender once five different people have hidden them, and turns their messages into a PDF evidence report. Four of us built it, and it took first place at HackWIE 3.0 and Code Kshetra 2.0.

It is a Chrome extension backed by an Express API and a Gemini moderation service, and it screens cheapest first: a list of abusive words, then verdicts already cached, and only then the model. My part was the backend (the API, the failover across five moderation deployments, the verdict cache and the evidence reports) and, in the extension, the popup and the login hand-off.

## My part

- **The backend.** The Express and MongoDB API: auth, hiding senders and messages, the archive with filters and paging, and stats.
- **Moderation that fails over.** Moved into its own service on five deployments, with health checks and a circuit breaker for each.
- **A verdict cache.** Redis, keyed by message and platform, so the same message is never paid for twice.
- **Evidence reports.** Gemini's findings rendered by Puppeteer into an A4 PDF, safe even when the model's answer isn't.
- **The crowd rule.** A sender is flagged once five different people have hidden them.
- **The extension's popup.** Preferences, blocked words, the hidden-message archive and reports, with login handed over from the dashboard.

## Cheapest first

Most messages never reach the model. A word list in the page catches the obvious, a shared Redis cache answers anything judged in the last hour, and only what's left costs a call.

## Five deployments, one moderation service

The model ran on free tiers, and free tiers run out, usually in the middle of a demo. So moderation moved into its own service on five deployments, and the API watches all of them: one that keeps failing is rested for a while, and each message goes to the first healthy one.

```js
// api/moderation/detect-harassment.js
router.post('/detect-harassment', async (req, res) => {
  const errors = [];
  const cacheKey = JSON.stringify({
    message: req.body.message.toLowerCase(),
    platform: req.body.platform.toLowerCase()
  });
  const cachedData = await redisClient.get(cacheKey);
  if (cachedData) {
    return res.json(JSON.parse(cachedData));
  }

  const availableServers = servers.filter(server => {
    const state = getCircuitBreakerState(server.url);
    return server.healthy || state.status === 'HALF_OPEN';
  });
  for (const server of availableServers) {
    try {
      return res.json(await handleRequest(server, req));
    } catch (error) {
      errors.push(`${server.url}: ${error.message}`);
    }
  }
  res.status(503).json({
    error: 'All servers failed to process request',
    retryAfter: 5
  });
});
```

The cache answers first; then each deployment whose breaker is closed or half open gets its turn.

## Evidence that survives bad model output

Gemini turns a sender's hidden messages into findings, patterns and a risk level, and Puppeteer renders them into an A4 PDF from a Handlebars template.

The model's answer is never trusted as it comes: parse, strip the formatting and parse again, fall back to what can be read from the text, then type-check every field, so the renderer never gets a broken shape.

```js
// api/user/report.js
let parsedResponse;
try {
  parsedResponse = JSON.parse(text);
} catch (parseError) {
  const cleanedText = text
    .replace(/```json\n?|\n?```/g, '')
    .replace(/\\n/g, ' ')
    .trim();
  try {
    parsedResponse = JSON.parse(cleanedText);
  } catch (secondParseError) {
    parsedResponse = {
      severityAssessment: text.includes('HIGH') ? 'HIGH' :
        text.includes('MEDIUM') ? 'MEDIUM' : 'LOW',
      recommendedActions: ['Review evidence manually'],
    };
  }
}

const sanitizedResponse = {
  severityAssessment:
    parsedResponse.severityAssessment || 'MEDIUM',
  keyFindings: Array.isArray(parsedResponse.keyFindings) ?
    parsedResponse.keyFindings
      .filter(item => item && typeof item === 'string') :
    ['Manual review recommended'],
};
```

Three tries at the model's answer, then a safe default for every field the report needs.

## Next

A labelled message set to measure detection accuracy, and a Chrome Web Store listing.
