EEG Image Classification and Caption Retrieval
A deep learning system that reads half a second of brain signal, recorded while a person looks at a picture, and tries to name the picture's category and find its caption by mapping the signal into a pretrained image and text model's shared space.
- Course
- Introduction to Deep Learning
- School
- Carnegie Mellon University
- Term
- Fall 2025
- Role
- ML Researcher
- Stack
- PyTorch
- Hugging Face
- CLIP
- Python
- NumPy
- pandas
- scikit-learn
- Jupyter
- Matplotlib
- Slurm
Problem
A new kind of sensor is only useful once its signal can be read in the terms people already use: a category, a sentence, a search. Brain computer interfaces are the hardest case, because the signal is weak, noisy and shaped differently in every person and on every day, and there is little labelled data to learn from. The industry answer is to align the new signal to a pretrained model that already links images and language, so the sensor inherits that model's vocabulary instead of learning one from scratch. This project tests that pattern on EEG recorded while 13 people looked at images from 20 categories.
Solution
Every trial is indexed once and split by recording session, so the model is always judged on a day it never saw. A shared encoder turns each trial into one vector, and a small head per person maps that vector to a category, so what is common across people is learned once and what differs stays local. The trained encoder's vectors are frozen and handed on as files, and a small projection learns to place them in the pretrained model's text space, whose weights all stay frozen, with a light adapter on its text output; retrieval is then a nearest neighbour search over an index of captions. The same metrics run first on images, where the pretrained model is strong, so the brain signal is measured between a known ceiling and chance.
- moves data
- holds the data every stage reads
- learns and judges
- shows
- provided or pretrained, used not written
- data
- frozen embeddings and adapted vectors
- scored results
Learnings
- Learning
A shared encoder with a head per person
Signals from people differ, but most of what matters is common. Learning one encoder for everyone and a small head per person, where each head learns only from its own person's data, keeps the shared part large and the personal part cheap. It is the pattern behind any model that serves many users or devices: one backbone to maintain, and a new user costs one head, not one model.
- Learning
Borrowing a foundation model's space
Training a vision and language model from brain data is out of reach; teaching a small projection to land in an existing model's space is not. Freezing the large model and training 2.2 million new weights around it, out of 153 million in the whole system, makes alignment cheap, keeps the text side a stable index, and lets a new sensor be searched with words. This is how a new modality joins a system without retraining the system.
- Learning
Evaluate the way it will be used
A model that will meet a new session has to be judged on a session it never saw, or its score is leakage; here the encoder fit its training sessions at about two thirds accuracy and held out sessions at about 6%. Running the same metrics on images first gives a ceiling, and chance gives a floor, so a weak result can be told apart from a broken pipeline. Without both, no number from a noisy sensor can be trusted enough to ship.