Skip to content

Determinants of Child Stunting across Indian States: A Weighted Logistic-Regression Study of NFHS-5 Data

  • 12 slides
  • 16 viva questions
  • 5 modules
  • No code needed

@child-stunting-determinants-nfhs5Updated Oct 2026

Survey-weighted odds ratios for mother's education, wealth, sanitation and birth order, from unit-level NFHS-5 data in R

B.Sc, Statistics · Sem 6 · Intermediate · 14 weeks · Team of 2

More info
Level
Intermediate · 14 weeks · Team of 2
Relevant for
Delhi
Common at
Delhi University, BCU / Bangalore University, Kerala University
Syllabus
Delhi University UGCF 2022 · Project Work (Statistics DSE-4B, CBCS) · Semester 6
Tech stack
  • R
  • survey (R package)
  • tidyverse
  • ggplot2
  • haven
  • sf (state map)
  • SPSS
  • Excel
  • LaTeX or MS Word for the report
For educational purposes only

Unlock this project

Full PPT + speaker notes, the step-by-step method, READMEFIRST, instructions and all 16 viva answers.

One-time. No subscription, no auto-renew, no drama.

Project packs

Credits never expire and work on any project. Use one here, save the rest for your friend who “will pay you back”.

  1. Pinned

    1 min

    Overview

    What this study does. Stunting (low height for age) is the most widely used indicator of long-term undernutrition in young children. The NFHS-5 national fact sheet reports that roughly one in three Indian children under five (35.5 %) is stunted, with large differences between states. This project asks a focused statistical question: after adjusting for other factors, which household and maternal characteristics are associated with higher odds of a child being stunted, and how much of the state-to-state gap remains?

    We use the unit-level NFHS-5 (2019–21) children's recode file, obtained through a free registration with The DHS Program. The outcome is binary: a child is stunted when the height-for-age z-score is below −2 standard deviations of the WHO Child Growth Standards median. Covariates are mother's education, household wealth quintile, type of sanitation facility, birth order, child's age and sex, place of residence (urban/rural) and state.

    The analysis respects the survey design: sampling weights, primary sampling units (clusters) and strata are declared with the R survey package before anything is estimated. We produce weighted descriptive tables, design-adjusted chi-square tests, a survey-weighted logistic regression reporting odds ratios with 95 % confidence intervals, model diagnostics (VIF, a goodness-of-fit check and ROC/AUC) and a state-level choropleth map. Key tables are cross-checked in SPSS and Excel so that the two-member team can defend every number in the viva.

    Syllabus alignment

    Delhi University · UGCF 2022

    Project Work (Statistics DSE-4B, CBCS) · Semester 6 · 12 credits

    Subjects this project applies
    • Survey Sampling and Indian Official Statistics
    • Statistical Inference (tests of significance, chi-square)
    • Linear Models / Regression Analysis
    • Research Methodology (DSE)
    • Statistical computing practicals in R, SPSS and Excel
    How it is evaluated

    See your department's project guidelines.

    Also fits: BCU / Bangalore University SEP 2024, Kerala University FYUGP 2024.

    1 min read · 16 viva questions

  2. 2 min

    Synopsis

    Abstract

    This project studies the determinants of stunting among children aged 0–59 months using the unit-level National Family Health Survey (NFHS-5, 2019–21) children's recode. Stunting is defined as height-for-age z-score (HAZ) < −2 SD of the WHO Child Growth Standards. Using R's survey package to account for weights, clustering and stratification, we estimate weighted prevalence by background characteristics, test associations with design-adjusted (Rao–Scott) chi-square tests, and fit a survey-weighted logistic regression. Results are reported as adjusted odds ratios with 95 % confidence intervals, followed by diagnostics and a state-level map.

    Introduction

    Stunting reflects chronic undernutrition and repeated infection in early life. It is linked to poorer schooling and adult earnings, which is why it appears in national nutrition programmes such as POSHAN Abhiyaan. India's NFHS series gives nationally and state-wise representative anthropometric data, which makes it an ideal real dataset for an undergraduate statistics project that goes beyond textbook examples.

    Existing work and the gap

    Official NFHS reports publish one-way tables — stunting by wealth, by education, by state — but a one-way table cannot separate overlapping effects. Richer households also tend to have better-educated mothers and improved toilets, so a raw gap by education partly reflects wealth. Many student reports also analyse NFHS data without weights or design variables, which biases estimates and understates standard errors. Our project fills this gap at undergraduate level by (a) building a multivariable model and (b) doing every estimate under the correct complex-survey design.

    Proposed study

    1. Obtain NFHS-5 KR data from The DHS Program, restrict to living, de jure children with a valid HAZ.
    2. Recode covariates following the Guide to DHS Statistics definitions.
    3. Weighted descriptive analysis and Rao–Scott chi-square tests.
    4. Survey-weighted logistic regression (quasibinomial) with state fixed effects.
    5. Diagnostics: VIF, goodness of fit, ROC/AUC; sensitivity analysis.
    6. A choropleth map of weighted state prevalence and adjusted state effects.

    Feasibility

    • Technical: R, RStudio and all packages are free; SPSS is available in most DU college labs. The KR file is large but fits comfortably in 8 GB RAM once unnecessary variables are dropped.
    • Economic: zero cost — the data is free for registered research use.
    • Operational: a two-member team can finish the work in about 14 weeks alongside Semester 6 courses; data access takes a few working days, so it is requested in Week 1.
  3. 1 min

    Problem statement

    Although stunting has declined slowly over successive NFHS rounds, the level remains high and uneven across Indian states. Policy discussions often quote single-factor gaps — "children of uneducated mothers are more often stunted" — but these factors overlap strongly with household wealth, sanitation and rural residence. Without a multivariable model it is impossible to say which factors remain important after the others are held constant, or whether the differences between states are explained by their socioeconomic profile.

    A second problem is methodological. NFHS is a stratified two-stage cluster sample, and treating it as a simple random sample produces biased prevalence estimates and artificially narrow confidence intervals, which can make weak associations look significant.

    This project therefore sets out to estimate, from unit-level NFHS-5 data and under the correct survey design, the adjusted association between child stunting (HAZ < −2 SD) and mother's education, wealth quintile, sanitation, birth order, residence and state, and to report these associations with honest uncertainty and model diagnostics.

  4. 1 min

    Objectives & scope

    1. 01Obtain authorised access to the NFHS-5 children's recode through The DHS Program and document the data-access process.
    2. 02Construct the binary outcome 'stunted' (HAZ < −2 SD, WHO standards) and recode covariates following DHS definitions.
    3. 03Declare the complex-survey design (weights, clusters, strata) in R and produce weighted prevalence tables with 95 % confidence intervals.
    4. 04Test bivariate associations using design-adjusted (Rao–Scott) chi-square tests.
    5. 05Fit and interpret a survey-weighted logistic regression reporting adjusted odds ratios and confidence intervals.
    6. 06Check the model with VIF, a goodness-of-fit test, ROC/AUC and a sensitivity analysis.
    7. 07Present weighted state-level prevalence on a choropleth map and cross-check key tables in SPSS and Excel.

    Scope

    In scope

    • Children aged 0–59 months in the NFHS-5 children's recode (KR) with a valid height-for-age z-score; all states and union territories covered by the survey.
    • One binary outcome (stunted vs not stunted) and a fixed set of covariates: mother's education, wealth quintile, sanitation facility (improved / not improved / open defecation), birth order, child's age group and sex, residence and state.
    • Weighted descriptive statistics, Rao–Scott chi-square tests, survey-weighted logistic regression and diagnostics.
    • A state-level choropleth map and a short comparison of R output with SPSS/Excel output.

    Out of scope

    • Causal claims: the cross-sectional design supports associations, not cause and effect.
    • District-level modelling, multilevel (random-effects) models and trend analysis across NFHS rounds — listed under future scope.
    • Severe stunting (HAZ < −3 SD), wasting and underweight as separate outcomes.
  5. 1 min

    Methodology

    Research design: secondary analysis of a nationally representative cross-sectional household survey, following a quantitative, hypothesis-testing approach.

    Hypotheses (expected directions, to be tested — not results):

    • H1: the odds of stunting decrease as mother's years of schooling increase.
    • H2: the odds of stunting decrease across wealth quintiles from poorest to richest.
    • H3: children in households without improved sanitation have higher odds of stunting.
    • H4: higher birth order is associated with higher odds of stunting.
    • H5: state differences remain significant after adjustment for the above.

    Model specification

    logit P(stunted = 1) = β0 + β1·MotherEdu + β2·Wealth + β3·Sanitation + β4·BirthOrder + β5·AgeGroup + β6·Sex + β7·Residence + Σ γs·State_s

    Reference categories: no education, poorest quintile, improved sanitation, birth order 1, age 0–11 months, male, urban, and a large state chosen as the reference (justified in the report). Estimation uses svyglm() with family = quasibinomial(); adjusted odds ratios are exp(β) with Wald 95 % CIs.

    Timeline (14 weeks, team of two)

    WeeksActivityOwner
    1–2DHS registration, literature reading, variable listboth
    3–4Data import, cleaning, recoding, design objectMember A
    5–6Weighted tables, Rao–Scott tests, SPSS/Excel cross-checkMember B
    7–9Logistic model, model selection, interpretationMember A
    10–11Diagnostics, sensitivity analysis, state mapMember B
    12–13Report writing, figures, internal reviewboth
    14Final report, slides and viva preparationboth

    Quality control: every recode is checked with a frequency table before and after; the unweighted count of children in the analytic sample is logged at each filtering step so the sample flow chart in the report can be reproduced exactly.

  6. 3 min

    Architecture & tech stack

    • R
    • survey (R package)
    • tidyverse
    • ggplot2
    • haven
    • sf (state map)
    • SPSS
    • Excel
    • LaTeX or MS Word for the report

    The study is organised as a linear analysis pipeline. Each stage has a single R script (numbered 01–06) and writes its output to a folder, so any stage can be re-run without repeating the earlier ones.

    flowchart TD
      A["Register on The DHS Program website and request NFHS-5 (India) data"] --> B["Download children's recode (KR) file in Stata/SPSS format"]
      B --> C["01 Import with haven and keep required variables"]
      C --> D["02 Filter: living, de jure children with valid HAZ"]
      D --> E["03 Recode outcome and covariates (DHS definitions)"]
      E --> F["04 Declare survey design: weight v005, PSU v021, strata v022"]
      F --> G["Weighted prevalence tables and Rao-Scott chi-square"]
      F --> H["Survey-weighted logistic regression"]
      H --> I["Diagnostics: VIF, goodness of fit, ROC/AUC"]
      G --> J["State prevalence choropleth map"]
      G --> K["Cross-check key tables in SPSS and Excel"]
      I --> L["Report chapters and presentation"]
      J --> L
      K --> L

    Key variables (children's recode)

    ConceptDHS variableRecode used
    Height-for-age z-scorehw70 (z × 100; values ≥ 9996 are flags)stunted = 1 if hw70 < −200
    Sample weightv005weight = v005 / 1,000,000
    Cluster (PSU)v021as is
    Stratumv022as is (confirm with the recode manual)
    Mother's educationv106none / primary / secondary / higher
    Wealth quintilev190poorest … richest
    Toilet facilityv116improved / not improved / open defecation
    Birth orderbord1, 2, 3, 4+
    Residencev025urban / rural
    Statev024factor
    Child's age and sexhw1 or b19, b40–11, 12–23, 24–35, 36–47, 48–59 months; sex

    Core R code (abridged)

    library(haven); library(dplyr); library(survey)
    kr <- read_dta("IAKR7xFL.DTA") |>            # exact file name from the download page
      select(v005, v021, v022, v024, v025, v106, v190, v116, bord, b4, b5, hw1, hw70) |>
      filter(b5 == 1, hw70 < 9996) |>
      mutate(w = v005 / 1e6, stunted = as.integer(hw70 < -200))
    options(survey.lonely.psu = "adjust")
    des <- svydesign(ids = ~v021, strata = ~v022, weights = ~w, data = kr, nest = TRUE)
    svyby(~stunted, ~v190, des, svymean)            # weighted prevalence by wealth
    svychisq(~stunted + v106, des)                  # Rao-Scott test
    m <- svyglm(stunted ~ edu + wealth + san + bord4 + agegrp + sex + res + state,
                design = des, family = quasibinomial())
    exp(cbind(OR = coef(m), confint(m)))
    

    Results tables to fill (no numbers are pre-filled)

    Covariate levelWeighted % stunted (95 % CI)Adjusted OR95 % CIp
    Mother: no education (ref)…1.00——
    Mother: higher education…………
    Poorest quintile (ref)…1.00——
    Richest quintile…………
    Open defecation vs improved…………
  7. 5 modules

    Modules

    • Module 1 — Data access and cleaning (Member A)

      Member A completes the DHS Program registration, downloads the KR file, writes the import and filtering script, logs the sample flow (all children → living → de jure → valid HAZ) and documents every recode with before/after frequency tables.

    • Module 2 — Survey design and descriptive analysis (Member B)

      Member B builds the svydesign object with weights, clusters and strata, produces weighted prevalence tables with 95 % confidence intervals for every covariate, runs Rao–Scott chi-square tests and prepares the corresponding bar charts in ggplot2.

    • Module 3 — Weighted logistic regression (Member A)

      Member A fits the survey-weighted logistic model, compares nested models using design-based Wald tests, converts coefficients to adjusted odds ratios with confidence intervals and writes the interpretation of each covariate in plain language.

    • Module 4 — Diagnostics and sensitivity (Member B)

      Member B checks multicollinearity with VIF, runs a goodness-of-fit test and ROC/AUC, and repeats the model with an alternative sanitation coding and without the state term to show how stable the key odds ratios are.

    • Module 5 — State map and cross-software validation (both members)

      Both members join weighted state prevalence to a state boundary file with sf and draw a choropleth map; they recompute two key tables in SPSS (Complex Samples) and Excel pivot tables to confirm that the R output matches.

  8. Locked

    Presentation

    12 slides with speaker notes. The outline below is free; the bullets, notes and the generated .pptx unlock with the project.

    1. Determinants of Child Stunting across Indian States
    2. Why stunting?
    3. Research question and hypotheses
    4. Data: NFHS-5 children's recode
    5. Variables
    6. Why survey weights matter
    7. Descriptive analysis
    8. Logistic regression model
    9. Results: adjusted odds ratios
    10. Diagnostics and robustness
    11. State-level map
    12. Conclusions, limitations and future scope

    Bullets, speaker notes and the .pptx download unlock with the project.

    Presentation is locked: 12 slides, Speaker notes, .pptx download.

  9. 1 min

    Future scope

    • Multilevel logistic models (children nested in clusters, districts and states) to separate household and area-level variation.
    • District-level analysis using NFHS-5 district identifiers and small-area mapping.
    • Trend decomposition between NFHS-4 and NFHS-5 to see which factors explain the change in stunting.
    • Additional outcomes: severe stunting, wasting, underweight and anaemia, modelled jointly.
    • Interaction terms such as mother's education × wealth, or sanitation × residence.
    • Sharing a reproducible R Markdown / Quarto report so juniors can rerun the pipeline on the next NFHS round.
  10. 9 sources

    References

    1. International Institute for Population Sciences (IIPS) — National Family Health Survey (NFHS) portal
    2. The DHS Program — data access and India survey datasets
    3. World Health Organization — Child Growth Standards
    4. IIPS and ICF, National Family Health Survey (NFHS-5), 2019–21: India Report, Ministry of Health and Family Welfare, Government of India
    5. Croft, T. N. et al., Guide to DHS Statistics (DHS-7), ICF
    6. Thomas Lumley, Complex Surveys: A Guide to Analysis Using R, Wiley, 2010
    7. survey: Analysis of Complex Survey Samples (R package on CRAN)
    8. David W. Hosmer, Stanley Lemeshow & Rodney X. Sturdivant, Applied Logistic Regression, 3rd ed., Wiley, 2013
    9. William G. Cochran, Sampling Techniques, 3rd ed., Wiley

    Cite this bundle

    OnlyProjects. (2026). Determinants of Child Stunting across Indian States: A Weighted Logistic-Regression Study of NFHS-5 Data: B.Sc Statistics project bundle [Educational resource]. https://onlyprojects.online/projects/bsc-stats-child-stunting-determinants-nfhs5

Slides, diagrams & files

12 slides. Titles are free; bullets, speaker notes and the .pptx unlock with the project.

  1. SLIDE 1

    Determinants of Child Stunting across Indian States

  2. SLIDE 2

    Why stunting?

  3. SLIDE 3

    Research question and hypotheses

  4. SLIDE 4

    Data: NFHS-5 children's recode

  5. SLIDE 5

    Variables

  6. SLIDE 6

    Why survey weights matter

  7. SLIDE 7

    Descriptive analysis

  8. SLIDE 8

    Logistic regression model

  9. SLIDE 9

    Results: adjusted odds ratios

  10. SLIDE 10

    Diagnostics and robustness

  11. SLIDE 11

    State-level map

  12. SLIDE 12

    Conclusions, limitations and future scope

Architecture diagram

1
flowchart TD
  A["Register on The DHS Program website and request NFHS-5 (India) data"] --> B["Download children's recode (KR) file in Stata/SPSS format"]
  B --> C["01 Import with haven and keep required variables"]
  C --> D["02 Filter: living, de jure children with valid HAZ"]
  D --> E["03 Recode outcome and covariates (DHS definitions)"]
  E --> F["04 Declare survey design: weight v005, PSU v021, strata v022"]
  F --> G["Weighted prevalence tables and Rao-Scott chi-square"]
  F --> H["Survey-weighted logistic regression"]
  H --> I["Diagnostics: VIF, goodness of fit, ROC/AUC"]
  G --> J["State prevalence choropleth map"]
  G --> K["Cross-check key tables in SPSS and Excel"]
  I --> L["Report chapters and presentation"]
  J --> L
  K --> L

Files

Viva questions & answers

3 of 16 questions free. Explain each answer in your own words before you move on.

  1. Concept

    How is stunting defined in your study?

    A child is stunted when the height-for-age z-score is below minus two standard deviations from the median of the WHO Child Growth Standards reference population. In the NFHS-5 children's recode this is variable hw70, stored multiplied by 100, so stunted means hw70 below −200 after removing flagged values.

  2. Concept

    Why did you use logistic regression instead of linear regression?

    Our outcome is binary — stunted or not — so a linear model could predict probabilities outside 0 and 1 and would violate constant-variance assumptions. Logistic regression models the log-odds of stunting as a linear function of covariates, and exponentiated coefficients give odds ratios that are easy to interpret.

  3. Concept

    What is an odds ratio and how do you interpret it here?

    The odds ratio compares the odds of stunting in one category with the odds in the reference category, holding the other covariates constant. An adjusted odds ratio below one for the richest quintile, for example, would mean lower odds of stunting than the poorest quintile after accounting for education, sanitation and the other variables.

+13 more questions

They and the answers unlock with the project. Try answering the ones above yourself first. Your examiner will.

For educational purposes only. Use this bundle to understand how the project works, then build and write your own. Submitting it verbatim is between you, your conscience and your external examiner.