- hackathon
- 2025
- Chrome extension
- Moderation
- Two first places
Harassment in LinkedIn DMs usually gets read before anyone can report it. Safire screens every incoming message, hides abuse behind a warning, flags a sender once five different people have hidden them, and turns their messages into a PDF evidence report. Four of us built it, and it took first place at HackWIE 3.0 and Code Kshetra 2.0.
It is a Chrome extension backed by an Express API and a Gemini moderation service, and it screens cheapest first: a list of abusive words, then verdicts already cached, and only then the model. My part was the backend (the API, the failover across five moderation deployments, the verdict cache and the evidence reports) and, in the extension, the popup and the login hand-off.



- Type
- hackathon
- Role
- The backend, and the extension's popup
- When
- Jan to Feb 2025, polished through July
- Team
- Four of us
- Stack
- Plasmo (MV3), React, Express, MongoDB, Redis, Gemini, Puppeteer, Next.js
- Links
- safire-five.vercel.appDemoCode
My part
The backend
The Express and MongoDB API: auth, hiding senders and messages, the archive with filters and paging, and stats.
Moderation that fails over
Moved into its own service on five deployments, with health checks and a circuit breaker for each.
How it worksA verdict cache
Redis, keyed by message and platform, so the same message is never paid for twice.
How it worksEvidence reports
Gemini's findings rendered by Puppeteer into an A4 PDF, safe even when the model's answer isn't.
How it worksThe crowd rule
A sender is flagged once five different people have hidden them.
The extension's popup
Preferences, blocked words, the hidden-message archive and reports, with login handed over from the dashboard.
How it fits together
Cheapest first
Most messages never reach the model. A word list in the page catches the obvious, a shared Redis cache answers anything judged in the last hour, and only what's left costs a call.
the model answers roughly one message in six
Five deployments, one moderation service
The model ran on free tiers, and free tiers run out, usually in the middle of a demo. So moderation moved into its own service on five deployments, and the API watches all of them: one that keeps failing is rested for a while, and each message goes to the first healthy one.
api/moderation/detect-harassment.js
router.post('/detect-harassment', async (req, res) => { const errors = []; const cacheKey = JSON.stringify({ message: req.body.message.toLowerCase(), platform: req.body.platform.toLowerCase() }); const cachedData = await redisClient.get(cacheKey); if (cachedData) { return res.json(JSON.parse(cachedData)); } const availableServers = servers.filter(server => { const state = getCircuitBreakerState(server.url); return server.healthy || state.status === 'HALF_OPEN'; }); for (const server of availableServers) { try { return res.json(await handleRequest(server, req)); } catch (error) { errors.push(`${server.url}: ${error.message}`); } } res.status(503).json({ error: 'All servers failed to process request', retryAfter: 5 });});Evidence that survives bad model output
Gemini turns a sender's hidden messages into findings, patterns and a risk level, and Puppeteer renders them into an A4 PDF from a Handlebars template.
The model's answer is never trusted as it comes: parse, strip the formatting and parse again, fall back to what can be read from the text, then type-check every field, so the renderer never gets a broken shape.
api/user/report.js
let parsedResponse;try { parsedResponse = JSON.parse(text);} catch (parseError) { const cleanedText = text .replace(/```json\n?|\n?```/g, '') .replace(/\\n/g, ' ') .trim(); try { parsedResponse = JSON.parse(cleanedText); } catch (secondParseError) { parsedResponse = { severityAssessment: text.includes('HIGH') ? 'HIGH' : text.includes('MEDIUM') ? 'MEDIUM' : 'LOW', recommendedActions: ['Review evidence manually'], }; }}const sanitizedResponse = { severityAssessment: parsedResponse.severityAssessment || 'MEDIUM', keyFindings: Array.isArray(parsedResponse.keyFindings) ? parsedResponse.keyFindings .filter(item => item && typeof item === 'string') : ['Manual review recommended'],};Next. A labelled message set to measure detection accuracy, and a Chrome Web Store listing.
