Movie Recommendation System
A movie recommender run as a production service: viewing events arrive as a stream, are checked and stored once, feed a loop that trains and registers candidate models, and the live traffic itself decides which one serves.
- Course
- Machine Learning in Production
- School
- Carnegie Mellon University
- Term
- Spring 2025
- Role
- ML Engineer
- Stack
- Python
- Apache Kafka
- PostgreSQL
- Surprise
- MLflow
- Flask
- Prometheus
- Grafana
- Docker Compose
- Jenkins
- pytest
Problem
Any service with a large catalogue and many users, from video streaming to retail and news, has to pick a short list for each person the moment they ask, and learn from what they do next. Error on held out ratings says little about whether people watch what they were shown, so the real test happens on live traffic, and it never ends. Tastes drift, new users arrive with no history, and every list served changes the data the next model learns from. This project builds that loop for a simulated streaming service with tens of thousands of users, from the event stream to a model chosen by an online experiment.
Solution
An event stream carries every watch, rating and served list, and an ingest job checks each event and keeps one durable store of ratings and viewing sessions, the dataset every model trains on. Training splits that history in time order, fits candidate models, and registers each one with its parameters, its error, its training data and its code version, so any model that serves can be traced back to the run that made it. A stateless service loads the latest registered version and answers with twenty movie ids per user, falling back to the movies the model scores highest overall for people it has never seen. A separate evaluator reads the same stream to score what was served against what people watched next, and an A/B split by user puts two candidates live side by side and compares their scores to pick the one to redeploy.
- moves data
- holds the governed data
- learns and judges
- shows
- platform component, configured not written
- events and rows
- governed data
- registered or winning model
- on demand
Learnings
- Learning
Evaluation that happens in production
Offline error picks a candidate, but only what people do after a list is served says whether it works. Scoring that from the event log, in a process of its own, keeps the serving path fast and gives one number that dashboards and experiments share. This is what lets a team change the model often and know whether each change helped.
- Learning
A registry as the contract between training and serving
When a model is a numbered version carrying its data, parameters and commit, the service never needs to know how it was built, and any answer can be traced to the run behind it. Swapping or rolling back becomes a name and a version instead of a file copied by hand. This is what makes retraining routine rather than a release.
- Learning
A recommender shapes its own data
Every list served changes what people watch, and so what the next model learns: the team measured that titles with divided ratings were recommended more often than the rest, and that a tenth of the catalogue drew most of the ratings. New users arrive with no history at all. Planning for both from the start, with a default for strangers and a watch on what gets recommended, is what keeps a recommender useful after launch instead of narrowing over time.