← Coursework

Air Quality Prediction and MLOps Deployment

A forecasting platform built to production standards: sensor readings arrive as a stream, are cleaned and governed once, feed a loop that trains and promotes models, and are served and monitored as containerised services.

Course
Fundamentals of Operationalizing AI
School
Carnegie Mellon University
Term
Fall 2025
Role
ML Engineer
Stack
  • Python
  • Redpanda
  • Apache Kafka
  • scikit-learn
  • statsmodels
  • XGBoost
  • MLflow
  • FastAPI
  • Evidently
  • Streamlit
  • Docker Compose
Code
Air Quality Streaming Analytics · localhost:8501liverefresh 5 s
Dashboard ControlsTime range
Last 30 days
Refresh every 5 sView
Next 6 Hours Forecasts (NOx)
Linear · test RMSE59.1 ppb
XGBoost · test RMSE52.5 ppb
SARIMA · test RMSE157.9 ppb
The operations dashboard, as it runs. Three of its seven views: live levels, anomaly detection and the forecast six hours ahead. Pick one in the sidebar.

Problem

Any operation that runs on sensors, from city air monitoring to plants and utilities, faces the same question: readings arrive every hour, and a forecast is only worth having if it is ready before the next one, keeps working as conditions shift, and can be trusted enough to act on. The hard part is rarely the model. It is the system around it: taking a live stream in, cleaning it once for everyone, retraining as data lands, promoting a better model safely, and noticing when the world has changed. This project builds that system end to end, on a public air quality dataset.

Solution

An event stream decouples the sensors from everything that consumes them, so either side can change alone. One stream processor cleans and scores each reading once and writes a single governed dataset, the contract every other stage reads. A training loop fits candidate models on that dataset, logs every run, and promotes the best on measured error while keeping earlier versions for rollback. A prediction service serves the promoted model behind a stable API, a drift monitor compares recent data with the recent past, and a dashboard shows both. Everything runs as declared services that start only when what they need is ready, so the whole platform comes up in order from one command.

IngestSensor feedHourly readings, replayed at10,000× with the real gapsbetween them keptproducer.py · 9,357 hourly rowsEvent streamDecouples the feed from itsconsumers; one topic, orderedby event timeRedpanda, Kafka API · 1 nodehourly readingsStream processorCleans and scores eachbatch once, for everyonedownstream; bad rows set asideconsumer.py · range rules · row scoresmall batches, in orderGovernclean, scored rowsGoverned dataset, the contract between every stagesingle source of truthOne table that only ever grows, read by every later stage; nothing downstream cleans the data again.air_quality_clean.csv · 25 columns · a full run keeps 9,279 rows and sets 78 aside with a reason, all 9,357 source rows accounted forLearn · watch · seetraining settwo windows of 7 dayslive rows, 5 sCandidate modelsLinear (Ridge, Lasso), SARIMAand boosted trees,each forecasting NOx 6 h ahead12 lag features · 70/15/15 time splitDrift monitorCompares the last 7 dayswith the 7 before them;report on demandEvidently · inside the API processOperations dashboardLive levels, patterns,anomalies and the 6 hforecast, in 7 viewsStreamlit :8501 · reads the datasetPromote · servelogged runsreport on demandRun log and model registryEvery run logged with itsmetrics; the lowest test errorwins, older versions archivedMLflow · train.py · 21 runs, 13 versionsPrediction serviceServes the promoted modelbehind a stable API; cached,with a local fallbackFastAPI :8000 · /api/v1/predict · 300 spromoted model
One governed dataset sits between the stream and everything that learns from it. Models compete on measured error, and the winner is served and watched.

Learnings

  • Learning

    An event stream as the backbone

    A queue between the sensors and everything that consumes them turns a script into a system. Either side can fail, restart or change on its own, order is kept, and every later stage reads the same feed. This is what lets ingestion, training and serving evolve at different speeds without breaking each other.

  • Learning

    Forecasting on a time series

    A forecast is a bet on the next few hours, so the model has to be built and judged on a clock: lag features, a chronological split, a fixed horizon. A model that looks good on a random split can be worthless in time order, and its quality depends on how fresh and clean the last few hours of data are, which ties the model to the pipeline that feeds it.

  • Learning

    Continuous retraining and promotion

    A model is a moving part, not a deliverable. Logging every run, comparing candidates on the same window, promoting on evidence and keeping the previous version, with a drift check watching the input, is the loop that makes a forecast safe to leave running unattended.

← Back to coursework