← Coursework

EEG Image Classification and Caption Retrieval

A deep learning system that reads half a second of brain signal, recorded while a person looks at a picture, and tries to name the picture's category and find its caption by mapping the signal into a pretrained image and text model's shared space.

Course
Introduction to Deep Learning
School
Carnegie Mellon University
Term
Fall 2025
Role
ML Researcher
Stack
  • PyTorch
  • Hugging Face
  • CLIP
  • Python
  • NumPy
  • pandas
  • scikit-learn
  • Jupyter
  • Matplotlib
  • Slurm
Code
EEG pathEEG trial122 channels × 500 time steps500 ms sampled at 1000 Hzz scored per trialds005589 · *_1000Hz.npy[B, 500, 122] to [B, 122, 500]Temporal convolutionskernels 7, 15, 31 in parallel85 channels each, concatenatedBatchNorm, GELUMultiScaleConvBlock[B, 255, 500]Separable convolutiondepthwise, kernel 3BatchNorm, GELUpointwise to 384 channelsMultiScaleConvBlock[B, 384, 500]Pool and normalisemax pool 2: 500 to 250 stepsLayerNorm over the 384MultiScaleConvBlock[B, 250, 384]8 Conformer blockshalf FF, 6 head self attentionconv module k 31, half FFwidth 384 over 250 steps28,057,191 parameters in all[B, 250, 384]Text pathMean over timeaverages the 250 stepsthe trial's embeddingforward_featuresCaptionone sentence per imagematched to each trialcaptions.txt[B, 384][N, 384]tokensA head per person13 × Linear 384 → 20routed by subjecteach head sees its own personsubject_heads · label smoothing 0.1Frozen exporta 384 dim vector per trialone .npy file per splitread frozen by the next stagemultihead_{split}_embeddings.npyPretrained text towerevery weight frozentext into 512 dimsopenai/clip-vit-base-patch32[B, 20] logits[B, 384][B, 512]Category scores20 logits per trialthe highest names the classEEG_BL_Model.ipynbProjectionMLP 384 → 1024 → 1024 → 512BatchNorm, GELU, dropout 0.1residual 384 → 512, L2 norm2,212,372 of 153,489,685 weights trainLow rank adapterx + B(A(x)) × alpha / rrank 32, alpha 832,768 weightsLoRAAdapter · on the text side[B, 512][B, 512]Cosine similarity[B, B] at temperature 0.07loss: InfoNCE 0.6plus cosine distillation 0.5retrieval: [8,331, 512] captions, top 5
The model. One encoder reads the trial, a head per person names the category, and a projection places the same vector beside the captions in the pretrained model's space.

Problem

A new kind of sensor is only useful once its signal can be read in the terms people already use: a category, a sentence, a search. Brain computer interfaces are the hardest case, because the signal is weak, noisy and shaped differently in every person and on every day, and there is little labelled data to learn from. The industry answer is to align the new signal to a pretrained model that already links images and language, so the sensor inherits that model's vocabulary instead of learning one from scratch. This project tests that pattern on EEG recorded while 13 people looked at images from 20 categories.

Solution

Every trial is indexed once and split by recording session, so the model is always judged on a day it never saw. A shared encoder turns each trial into one vector, and a small head per person maps that vector to a category, so what is common across people is learned once and what differs stays local. The trained encoder's vectors are frozen and handed on as files, and a small projection learns to place them in the pretrained model's text space, whose weights all stay frozen, with a light adapter on its text output; retrieval is then a nearest neighbour search over an index of captions. The same metrics run first on images, where the pretrained model is strong, so the brain signal is measured between a known ceiling and chance.

SignalsEEG and caption store26,000 trials from 13 people122 channels × 500 ms each20 classes, 9,825 captionsds005589 · *_1000Hz.npy · captions.txtrun files, trial mapsmmap trial sliceTrial index, session splitjoins each trial to its caption3 / 1 / 1 sessions per person15,600 / 5,200 / 5,200 trialsdataset_builder.py · seed 42Trial loadermemory maps each recordingz scores every trialbatches of 128 × 500 × 122dataloader.py · mmap_mode='r'index, splitClassify[128, 500, 122]caption per trialShared EEG encoderkernels of 7, 15 and 31 steps8 Conformer blocks, width 38428,057,191 parameterseeg_conformer_multiscale.py[N, 384] per split[128, 384]Embedding storea 384 dim vector per trialfrozen once training endsthe seam between the stagesmultihead_{split}_embeddings.npyA head per person13 linear heads, 384 → 20only present people learnpicked on validation accuracytrain_utils.py · label smoothing 0.1Align and retrieve15,600 × 384logits [128, 20]Projection, text adapterMLP 384 → 1024 → 1024 → 512rank 32 adapter on the text2.2M of 153.5M weights trainedFINETUNE_ER_CLIP.ipynb · InfoNCE + KDtext vectors, 512EEG and caption vectorsFrozen image, text modelone 512 dim space for bothevery weight frozenencodes captions and imagesopenai/clip-vit-base-patch32Caption index8,331 unique captionsadapted and normalised, 512 dimranked by cosine, top 5FINETUNE_ER_CLIP.ipynb · 8,331 × 512Evaluateimage, text vectorstop 5 per trialImage to caption referencefrozen model, image to caption9,825 pairs, the same metricsR@1 0.198, class aware 0.970task2a.ipynb · images, captions.txtreference scoresMetrics and reportsaccuracy, confusion matrix, per classRecall@1/3/5, MAP, CLIPScore, BERTScoreEEG 6.1% vs 5% chance; class R@1 0.067eval_results.py · eeg_retrieval_results.csv
One encoder, frozen once trained, feeds both questions: which category, and which caption. Images run the same metrics first, as the ceiling the brain signal is measured against.

Learnings

  • Learning

    A shared encoder with a head per person

    Signals from people differ, but most of what matters is common. Learning one encoder for everyone and a small head per person, where each head learns only from its own person's data, keeps the shared part large and the personal part cheap. It is the pattern behind any model that serves many users or devices: one backbone to maintain, and a new user costs one head, not one model.

  • Learning

    Borrowing a foundation model's space

    Training a vision and language model from brain data is out of reach; teaching a small projection to land in an existing model's space is not. Freezing the large model and training 2.2 million new weights around it, out of 153 million in the whole system, makes alignment cheap, keeps the text side a stable index, and lets a new sensor be searched with words. This is how a new modality joins a system without retraining the system.

  • Learning

    Evaluate the way it will be used

    A model that will meet a new session has to be judged on a session it never saw, or its score is leakage; here the encoder fit its training sessions at about two thirds accuracy and held out sessions at about 6%. Running the same metrics on images first gives a ceiling, and chance gives a floor, so a weak result can be told apart from a broken pipeline. Without both, no number from a noisy sensor can be trusted enough to ship.

← Back to coursework