# Sixth-Finger BCI

> A brain-computer interface that learns a finger the body never had: stroke patients' EEG decoded with transfer learning, sent to a 3D hand. With two classmates, in progress.

- Kind: side project
- Status: in progress, live and still being built
- Role: The data pipeline, the decoders, the evaluation and the serving
- When: Aug 2026 to now
- Team: Three of us, with a faculty guide
- Stack: Python, PyTorch, MNE, pyRiemann, EEGNet, ONNX, FastAPI, React Three Fiber, Kaggle
- Live: https://rajveers.com/projects/sixth-finger

A stroke patient imagines a movement; a decoder reads which one from their EEG, and a 3D hand moves, including a robotic sixth finger the body doesn't have. The question behind it: can what a decoder learned from ordinary motor imagery cut the calibration a new command needs?

It is our B.Tech minor project, three of us with a faculty guide. I built the pipeline end to end: the audit of the data, the windows, the decoders and their transfer learning, the harness that evaluates them, and the server that streams it all to the browser.

## How it fits together

Audit the data before training on it, check every fit for leakage, and stamp every score with the code and data behind it; then stream the decoder's answer to a hand in the browser.

## In numbers

- 24: stroke patients' EEG
- 9,615: four-second windows
- 138: tests, run in CI

## My part

- **A dataset you can trust.** An integrity screen of 35,778 pairs of runs found 8 byte-identical recordings; 3 sessions were excluded before anything trained.
- **A leak, made visible.** A label-shuffle test shows the published baseline's leakage: filters fitted before the split score 80% or more on pure noise.
- **Questions fixed first.** Six hypotheses, the statistics and the calibration budgets written down and frozen before a single model ran.
- **Decoders and transfer learning.** CSP with LDA, Riemannian tangent space and EEGNet, with re-centring and Euclidean alignment carrying what was learned to a new user and a new command; EEGNet trained on Kaggle GPUs.
- **Every number traceable.** Each score carries the commit, config hash and data hash behind it, and re-running a finished experiment adds nothing.
- **Served to a 3D hand.** The model exported to ONNX (20.5 KB, 0.45 ms a window on a CPU) and streamed by FastAPI over WebSocket, with inference in the browser as the fallback.

## Audit before you train

The dataset is public: 24 stroke patients, three sessions each, imagining a hand movement, a sixth-finger movement, or rest. Before training anything I screened it, and 8 recordings turned out to be byte-identical copies of others, two patients shared a raw signal under different cue orders, and one run was empty. Those sessions came out.

What is left is cut into 9,615 windows of four seconds, 40 channels by 1,000 samples, rebuilt from the raw files in about 45 seconds and checked against a hash manifest, so a model never trains on a file that changed under it.

## Streaming a brain to a browser

A FastAPI server replays a recorded session over a WebSocket: 20 samples every 80 ms, the trial's phases as they happen, and one decode per four-second window with how long it took. The browser can pause, change the speed or jump; a paused stream sends a heartbeat every 2 s, and a server at its limit of 8 streams says so with a close code instead of hanging.

## Next

The decoders trained on the whole cohort, the calibration study they answer, the hand in the browser, and the paper.
