Skip to content

Hyderabad Air-Quality Data Pipeline and Next-Day PM2.5 Forecasting with AQI Dashboard

  • 12 slides
  • 15 viva questions
  • 6 modules
  • Code included

@hyderabad-air-quality-pm25-forecastingUpdated Oct 2026

CPCB station data plus weather, a scheduled ETL, time-series-validated models and a Streamlit AQI dashboard

MCA, Data Science & Analytics · Sem 4 · Intermediate · 14 weeks · Solo

More info
Level
Intermediate · 14 weeks · Solo
Relevant for
Telangana
Common at
Osmania University, JNTUH, Anna University
Syllabus
Osmania University 2-year MCA 2022-23 · Proj401 Project Work · Semester 4
Tech stack
  • Python 3.11
  • pandas
  • NumPy
  • scikit-learn
  • Matplotlib
  • SQLite / PostgreSQL
  • SQLAlchemy
  • Streamlit
  • APScheduler
  • pytest
  • TensorFlow-Keras (optional LSTM)
For educational purposes only

Unlock this project

Full PPT + speaker notes, source code and setup steps, READMEFIRST, instructions and all 15 viva answers.

One-time. No subscription, no auto-renew, no drama.

Project packs

Credits never expire and work on any project. Use one here, save the rest for your friend who “will pay you back”.

  1. Pinned

    1 min

    Overview

    AirWatch Hyderabad is a data-science project that builds a complete, reproducible pipeline from raw air-quality readings to a next-day PM2.5 forecast and a public-friendly dashboard. Hyderabad's air is monitored by continuous ambient air-quality monitoring stations whose data is published through the Central Pollution Control Board (CPCB). The data is valuable but messy: readings are missing for hours or days, units and column names vary between downloads, and a citizen looking at a raw CSV cannot tell whether tomorrow will be a "Poor" day.

    The project ingests station CSV downloads for several Hyderabad stations, joins them with daily weather features (temperature, humidity, wind speed, rainfall), cleans and imputes gaps, and engineers time-series features such as lags, rolling means, day-of-week and festival flags for Diwali, Bhogi and New Year. It then compares a persistence baseline, linear regression, random forest and gradient boosting (with an optional LSTM from the Deep Learning elective) using time-series cross-validation, MAE and RMSE.

    A scheduled ETL job loads new files into SQLite (PostgreSQL optional), retrains weekly and writes forecasts. A Streamlit dashboard shows history, the next-day forecast and its CPCB National AQI category. The work applies Python (PCC204), Machine Learning (PCC205), Data Science (PCC303) and DBMS (PCC202).

    Syllabus alignment

    Osmania University · 2-year MCA 2022-23

    Proj401 · Project Work · Semester 4 · 12 credits · CIE 50 + SEE 100

    Subjects this project applies
    • PCC204 Python
    • PCC205 Machine Learning
    • PCC303 Data Science
    • PCC202 DBMS
    • Deep Learning elective (optional LSTM)
    How it is evaluated

    See your department's project guidelines.

    Also fits: JNTUH MCA R22, Anna University MCA Regulation 2021.

    1 min read · 15 viva questions

  2. 2 min

    Synopsis

    Abstract

    Fine particulate matter (PM2.5) is the pollutant most closely linked to respiratory and cardiovascular harm, and Indian cities regularly exceed the 24-hour national standard of 60 µg/m³ in winter. This project builds an end-to-end pipeline for Hyderabad that ingests CPCB continuous-monitoring data and weather data, cleans and stores it, engineers time-series features, and forecasts next-day mean PM2.5 per station. Forecasts are mapped to the CPCB National AQI categories and shown on a Streamlit dashboard. Models are evaluated with forward-chaining time-series cross-validation against a persistence baseline.

    Introduction

    Residents, schools and a fictional citizen group, the Deccan Clean Air Forum, want a simple answer: will tomorrow be Satisfactory, Moderate or Poor? Current AQI apps show the present value, not a forecast, and public data must be downloaded, cleaned and merged by hand before anyone can model it. PM2.5 in Hyderabad follows clear patterns — a winter peak, a monsoon trough, weekday traffic effects and festival spikes — which makes it a good candidate for supervised learning on lagged features.

    Existing System

    • Raw station data is available as downloads, but with gaps, outliers and inconsistent headers.
    • Public AQI displays report current or past values only.
    • Student projects often use a random train-test split on time-series data, which leaks the future and overstates accuracy.

    Proposed System

    • Repeatable ETL: validated ingestion, unit harmonisation, outlier rules, gap imputation and a tidy daily table in a database.
    • Feature engineering: lag-1/2/3/7 PM2.5, 3- and 7-day rolling means, weather of the previous day, calendar and festival flags.
    • Models: persistence, linear regression, random forest, gradient boosting, optional LSTM, all compared with TimeSeriesSplit cross-validation.
    • Streamlit dashboard with history, forecast, AQI category and model diagnostics.

    Literature Gap

    Studies on Indian cities show that meteorology and lagged pollution explain much of the day-to-day variation, but most published work stops at a notebook. The gap addressed here is an honest, automated, re-runnable pipeline with baseline comparison and a usable interface.

    Feasibility

    • Technical: all libraries are open-source; a laptop trains the tree models in minutes.
    • Economic: zero licence cost; data is public.
    • Operational: the dashboard needs no training; the ETL runs unattended on a schedule.
  3. 1 min

    Problem statement

    Hyderabad's air quality is measured continuously at several monitoring stations, but the published data cannot be used directly for decision-making. Downloads contain missing hours, sensor spikes, changing column names and no weather context, so every analysis starts with hours of manual cleaning. Public displays show current AQI, not tomorrow's, which leaves schools, outdoor workers and people with asthma unable to plan. Where forecasting has been attempted in student work, models are often validated with random splits that leak future information and report misleadingly low errors.

    The problem is to design a reproducible data pipeline that ingests, cleans, imputes and stores station and weather data on a schedule, and to build and honestly evaluate next-day PM2.5 forecasting models against a persistence baseline using time-series cross-validation, presenting the forecast as a CPCB National AQI category on an accessible dashboard.

  4. 1 min

    Objectives & scope

    1. 01Build a scheduled ETL that ingests CPCB station CSV downloads and daily weather data into a relational store with schema validation.
    2. 02Clean the data using documented rules for outliers, unit harmonisation and duplicate timestamps, and impute gaps with time-aware methods.
    3. 03Engineer lag, rolling-window, weather, calendar and festival features for next-day PM2.5 prediction.
    4. 04Train and compare persistence, linear regression, random forest and gradient boosting models (optional LSTM) with TimeSeriesSplit cross-validation.
    5. 05Report MAE, RMSE and AQI-category accuracy per station and per season in a results table.
    6. 06Map predictions to CPCB National AQI categories using the official PM2.5 breakpoints.
    7. 07Deliver a Streamlit dashboard showing history, forecast, category and model diagnostics.

    Scope

    In scope

    • 4–6 Hyderabad monitoring stations, 2019 onwards (or whatever history the student downloads), aggregated to daily means.
    • PM2.5 as the target; PM10, NO2 and CO as optional co-features where available.
    • Daily weather features from a public historical weather API or IMD-style daily data.
    • Batch ETL on a schedule; weekly retraining; next-day (t+1) point forecast per station.
    • Streamlit dashboard for local or campus-server use.

    Out of scope

    • Real-time hourly nowcasting, satellite aerosol data and chemical transport models.
    • Health-impact modelling or medical advice.
    • Automated scraping of portals whose terms do not permit it: files are downloaded manually into a watched folder.
  5. 1 min

    Methodology

    The project follows CRISP-DM (business understanding, data understanding, data preparation, modelling, evaluation, deployment) run in short iterations.

    PhaseWeeksActivitiesDeliverable
    Understanding1–2Study AQI methodology and NAAQS, choose stations, define the forecast taskProblem definition and data plan
    Data acquisition2–3Download station CSVs, fetch weather, design schemaRaw data + data dictionary
    Preparation4–6Cleaning rules, imputation experiments, EDA (seasonality, festival spikes, correlations)Clean daily table, EDA notebook
    Features6–7Lags, rolling means, calendar and festival flags, weather of day tFeature table
    Modelling8–10Baseline, linear, random forest, gradient boosting, optional LSTM; hyper-parameter search inside CVModel comparison
    Deployment11–12ETL scheduler, model registry folder, Streamlit dashboardRunning system
    Evaluation & report13–14Final results table, error analysis, report and viva preparationReport, PPT

    Imputation: gaps of up to 3 hours are linearly interpolated before daily aggregation; a day is kept only if at least 16 valid hours exist; longer gaps are left missing and the row is excluded from training rather than invented.

    Evaluation protocol: TimeSeriesSplit(n_splits=5) on data up to the last year, with the final 12 months held out as a test set never touched during tuning. Metrics: MAE, RMSE (µg/m³) and the percentage of days whose predicted AQI category matches the observed one. A results table (station × model × metric) is filled from the student's own run; no numbers are pre-claimed.

  6. 1 min

    Architecture & tech stack

    • Python 3.11
    • pandas
    • NumPy
    • scikit-learn
    • Matplotlib
    • SQLite / PostgreSQL
    • SQLAlchemy
    • Streamlit
    • APScheduler
    • pytest
    • TensorFlow-Keras (optional LSTM)

    The system is a batch data pipeline with a presentation layer. Each stage writes to the database so that it can be re-run independently.

    flowchart TD
      A["Station CSV downloads (drop folder)"] --> I["Ingest and validate schema"]
      B["Weather API (daily)"] --> I
      I --> C["Clean: units, outliers, duplicates"]
      C --> M["Impute short gaps, aggregate to daily"]
      M --> DB[("SQLite / PostgreSQL")]
      DB --> F["Feature engineering"]
      F --> T["Train and cross-validate models"]
      T --> R["Model registry (joblib files + metrics)"]
      R --> P["Next-day forecast job"]
      P --> DB
      DB --> S["Streamlit dashboard"]
      SCH["APScheduler: daily ETL, weekly retrain"] --> I
      SCH --> T

    Data model

    erDiagram
      STATION ||--o{ READING_DAILY : "records"
      STATION ||--o{ WEATHER_DAILY : "located near"
      STATION ||--o{ FORECAST : "has"
      MODEL_RUN ||--o{ FORECAST : "produces"
      STATION {
        int id PK
        string name
        float latitude
        float longitude
      }
      READING_DAILY {
        int id PK
        int stationId FK
        date day
        float pm25
        float pm10
        int validHours
      }
      WEATHER_DAILY {
        int id PK
        int stationId FK
        date day
        float tempMean
        float humidity
        float windSpeed
        float rainMm
      }
      MODEL_RUN {
        int id PK
        string algorithm
        datetime trainedAt
        float cvMae
        float cvRmse
      }
      FORECAST {
        int id PK
        int stationId FK
        int modelRunId FK
        date targetDay
        float pm25Pred
        string aqiCategory
      }

    Three support tables complete the schema: QUARANTINE (rejected rows with the file, row number and reason), HOLIDAY (the yearly festival table) and INGEST_LOG (which files were loaded, by content hash).

    AQI mapping. The dashboard converts predicted 24-hour PM2.5 to the CPCB National AQI category using the published PM2.5 breakpoints: Good 0–30, Satisfactory 31–60, Moderate 61–90, Poor 91–120, Very Poor 121–250 and Severe above 250 µg/m³. The mapping lives in one tested function so it can be updated if CPCB revises the table.

  7. 6 modules

    Modules

    • Ingestion & Validation

      Watches a drop folder for station CSVs, normalises column names, parses timestamps to Asia/Kolkata, validates units and ranges with a schema check, and fetches daily weather for each station's coordinates. Rejected rows are logged to a quarantine table with the reason.

    • Cleaning & Imputation

      Removes duplicates and physically impossible values, flags sudden single-hour spikes, interpolates gaps of up to three hours, and aggregates to daily means only when at least 16 valid hours exist. Every rule is a named, unit-tested function.

    • Feature Engineering

      Builds lag-1, lag-2, lag-3 and lag-7 PM2.5, 3- and 7-day rolling means, previous-day weather, day-of-week, month, and festival flags for Diwali, Bhogi and New Year from a small holiday table the student maintains for each year.

    • Modelling & Evaluation

      Trains persistence, linear regression, random forest and gradient boosting (optional Keras LSTM) inside scikit-learn pipelines, tunes with TimeSeriesSplit, evaluates on a held-out final year and stores metrics and model files in a registry.

    • Scheduler & Forecast Job

      APScheduler runs the ETL daily and retraining weekly, then writes the next-day forecast and AQI category for each station into the FORECAST table with the model run that produced it.

    • Streamlit Dashboard

      Shows station history, seasonal patterns, tomorrow's forecast with its colour-coded AQI category, the cross-validation leaderboard and residual plots, so a non-technical user and an examiner both see what the model does and where it fails.

  8. Locked

    Presentation

    12 slides with speaker notes. The outline below is free; the bullets, notes and the generated .pptx unlock with the project.

    1. AirWatch Hyderabad: Next-Day PM2.5 Forecasting
    2. Why PM2.5 and Why Forecast
    3. Problem Statement
    4. Objectives
    5. Data Sources
    6. Pipeline Architecture
    7. Cleaning & Imputation
    8. Feature Engineering
    9. Models & Validation
    10. Results
    11. Dashboard Demo
    12. Conclusion & Future Scope

    Bullets, speaker notes and the .pptx download unlock with the project.

    Presentation is locked: 12 slides, Speaker notes, .pptx download.

  9. 1 min

    Future scope

    • Hourly nowcasting with gradient boosting or sequence models and a 72-hour horizon.
    • Satellite aerosol optical depth and fire-count features for crop-burning episodes.
    • Probabilistic forecasts using quantile regression to give an uncertainty band instead of one number.
    • City-wide spatial interpolation to estimate PM2.5 between stations.
    • Alert service that sends an SMS or app notification to registered schools when a Poor day is forecast.
  10. 9 sources

    References

    1. Central Pollution Control Board (CPCB), Government of India
    2. CPCB, National Air Quality Index (report and PM2.5 breakpoint table), 2014
    3. CPCB, National Ambient Air Quality Standards notification, November 2009
    4. scikit-learn: TimeSeriesSplit
    5. pandas Documentation
    6. Streamlit Documentation
    7. Open-Meteo weather API
    8. Rob J. Hyndman & George Athanasopoulos, Forecasting: Principles and Practice, 3rd ed., OTexts
    9. Aurélien Géron, Hands-On Machine Learning with Scikit-Learn, Keras and TensorFlow, 3rd ed., O'Reilly

    Cite this bundle

    OnlyProjects. (2026). Hyderabad Air-Quality Data Pipeline and Next-Day PM2.5 Forecasting with AQI Dashboard: MCA Data Science & Analytics project bundle [Educational resource]. https://onlyprojects.online/projects/mca-ds-hyderabad-air-quality-pm25-forecasting

Slides, diagrams & files

12 slides. Titles are free; bullets, speaker notes and the .pptx unlock with the project.

  1. SLIDE 1

    AirWatch Hyderabad: Next-Day PM2.5 Forecasting

  2. SLIDE 2

    Why PM2.5 and Why Forecast

  3. SLIDE 3

    Problem Statement

  4. SLIDE 4

    Objectives

  5. SLIDE 5

    Data Sources

  6. SLIDE 6

    Pipeline Architecture

  7. SLIDE 7

    Cleaning & Imputation

  8. SLIDE 8

    Feature Engineering

  9. SLIDE 9

    Models & Validation

  10. SLIDE 10

    Results

  11. SLIDE 11

    Dashboard Demo

  12. SLIDE 12

    Conclusion & Future Scope

Architecture diagrams · 2

1
flowchart TD
  A["Station CSV downloads (drop folder)"] --> I["Ingest and validate schema"]
  B["Weather API (daily)"] --> I
  I --> C["Clean: units, outliers, duplicates"]
  C --> M["Impute short gaps, aggregate to daily"]
  M --> DB[("SQLite / PostgreSQL")]
  DB --> F["Feature engineering"]
  F --> T["Train and cross-validate models"]
  T --> R["Model registry (joblib files + metrics)"]
  R --> P["Next-day forecast job"]
  P --> DB
  DB --> S["Streamlit dashboard"]
  SCH["APScheduler: daily ETL, weekly retrain"] --> I
  SCH --> T
2
erDiagram
  STATION ||--o{ READING_DAILY : "records"
  STATION ||--o{ WEATHER_DAILY : "located near"
  STATION ||--o{ FORECAST : "has"
  MODEL_RUN ||--o{ FORECAST : "produces"
  STATION {
    int id PK
    string name
    float latitude
    float longitude
  }
  READING_DAILY {
    int id PK
    int stationId FK
    date day
    float pm25
    float pm10
    int validHours
  }
  WEATHER_DAILY {
    int id PK
    int stationId FK
    date day
    float tempMean
    float humidity
    float windSpeed
    float rainMm
  }
  MODEL_RUN {
    int id PK
    string algorithm
    datetime trainedAt
    float cvMae
    float cvRmse
  }
  FORECAST {
    int id PK
    int stationId FK
    int modelRunId FK
    date targetDay
    float pm25Pred
    string aqiCategory
  }

Files

Viva questions & answers

3 of 15 questions free. Explain each answer in your own words before you move on.

  1. Concept

    Why can't you use a random train-test split for this problem?

    PM2.5 values on consecutive days are strongly correlated, so a random split puts days from the future into training and lets the model peek at neighbouring values. TimeSeriesSplit always trains on the past and validates on the following block, which matches how the forecast is used.

  2. Concept

    What is a persistence baseline and why is it important?

    Persistence predicts that tomorrow's PM2.5 equals today's. Because pollution changes slowly, it is surprisingly hard to beat. Any model I report must lower MAE compared with persistence; otherwise the added complexity is not justified, and I state that openly in the results chapter.

  3. Concept

    How is the CPCB AQI category derived from PM2.5?

    CPCB publishes breakpoints for the 24-hour PM2.5 average: 0–30 Good, 31–60 Satisfactory, 61–90 Moderate, 91–120 Poor, 121–250 Very Poor and above 250 Severe. The overall AQI uses the worst sub-index across pollutants, but my dashboard shows the PM2.5 sub-index category only.

+12 more questions

They and the answers unlock with the project. Try answering the ones above yourself first. Your examiner will.

For educational purposes only. Use this bundle to understand how the project works, then build and write your own. Submitting it verbatim is between you, your conscience and your external examiner.