Skip to content

What Drives MGNREGA Work Demand? A District-Level Regression Study of Uttar Pradesh and Bihar (R)

  • 12 slides
  • 16 viva questions
  • 6 modules
  • No code needed

@mgnrega-work-demand-district-regression-up-biharUpdated Oct 2026

Secondary data from the MGNREGA MIS, Census 2011, NITI Aayog's MPI and IMD rainfall, 113 districts, OLS with robust errors

B.A., Economics · Sem 7 & 8 · Intermediate · 24 weeks · Solo

More info
Branch
Economics
Level
Intermediate · 24 weeks · Solo
Relevant for
Delhi
Common at
University of Delhi, University of Kerala, Bengaluru City University
Syllabus
Delhi University UGCF 2022 · Dissertation / Academic Project / Entrepreneurship · Semester 7 & 8
Tech stack
  • R 4.x + RStudio
  • tidyverse, readxl, janitor
  • lmtest, sandwich, car (diagnostics)
  • modelsummary / stargazer (regression tables)
  • MS Excel (data assembly)
  • MGNREGA MIS public reports
  • Census 2011 district tables
  • NITI Aayog National MPI 2023
  • IMD district rainfall statistics
For educational purposes only

Unlock this project

Full PPT + speaker notes, the step-by-step method, READMEFIRST, instructions and all 16 viva answers.

One-time. No subscription, no auto-renew, no drama.

Project packs

Credits never expire and work on any project. Use one here, save the rest for your friend who “will pay you back”.

  1. Pinned

    1 min

    Overview

    The Mahatma Gandhi National Rural Employment Guarantee Act (MGNREGA) promises every rural household up to 100 days of wage employment a year on demand. Yet the number of days households actually get varies enormously from one district to the next — even between neighbouring districts in the same state, with the same notified wage rate. This dissertation asks which district characteristics explain that variation.

    The study builds a cross-section of 113 districts — 75 in Uttar Pradesh and 38 in Bihar — for financial year 2022-23 entirely from public secondary sources: the MGNREGA MIS (person-days, households worked, households completing 100 days), Census 2011 district tables (agricultural labourers, SC/ST share, female literacy), NITI Aayog's National Multidimensional Poverty Index 2023 (district headcount ratios) and IMD district monsoon rainfall departures.

    Two outcomes are modelled with ordinary least squares in R: average person-days per working household and the share of working households that completed 100 days. The analysis reports heteroskedasticity-robust standard errors, checks multicollinearity (VIF), functional form (RESET) and influential districts (Cook's distance), and compares the two states with a state dummy and an interaction.

    The bundle gives you the full research design, a variable dictionary with sources, an R workflow you can reproduce from scratch, a chapter plan for the UGCF Semester 7–8 dissertation, a presentation outline and a viva bank. The data assembly is the real work here — and the part your supervisor will check most carefully.

    Syllabus alignment

    Delhi University · UGCF 2022

    Dissertation / Academic Project / Entrepreneurship · Semester 7 & 8

    Subjects this project applies
    • Research Methodology (DSE, Sem 6 or 7)
    • Introductory Econometrics
    • Statistical Methods for Economics
    • Indian Economy
    • Development Economics
    How it is evaluated

    See your department's project guidelines.

    Also fits: Kerala University FYUGP 2024, BCU SEP 2024 (NEP 2021 for 2021–24 batches).

    1 min read · 16 viva questions

  2. 2 min

    Synopsis

    Abstract

    MGNREGA is India's largest public-works programme, but its reach is uneven: some districts provide 60+ days of work to the average participating household while others struggle to cross 35. Using a cross-section of 113 districts of Uttar Pradesh and Bihar for FY 2022-23, this dissertation estimates the association between MGNREGA work outcomes and district-level distress, labour-market and social indicators. Outcomes are taken from the MGNREGA MIS; predictors from Census 2011, NITI Aayog's National MPI 2023 and IMD rainfall data. OLS models with robust standard errors are estimated in R. The findings are discussed in the light of the "self-targeting" argument for employment guarantees and the literature on administrative capacity and rationing.

    Introduction

    An employment guarantee is designed to be self-targeting: because the work is manual and the wage modest, only households that need it will ask for it. If the programme works as designed, districts with greater distress — more landless labourers, higher poverty, a failed monsoon — should generate more person-days. If instead outcomes track administrative capacity or social inequality in who gets heard, the pattern will be different. District data let an undergraduate test this idea with public, verifiable numbers.

    Literature and gap

    Earlier studies of MGNREGA have examined participation using household surveys (NSS, IHDS), documented rationing — households wanting work but not getting it — and analysed wage payment delays and the role of state capacity. Much of this work is either national in scope or based on survey rounds that are now a decade old. District-level, post-pandemic analyses that combine the MIS with the new NITI Aayog MPI are rare, particularly comparisons between two high-poverty states with very different administrative histories. This dissertation fills that gap on a modest scale.

    Research questions

    1. How much do MGNREGA person-days per household vary across districts of UP and Bihar in FY 2022-23?
    2. Are districts with a higher MPI headcount and a larger share of agricultural labourers associated with more person-days?
    3. Is a monsoon rainfall deficit associated with higher work demand in the same year?
    4. Does the association differ between Uttar Pradesh and Bihar?

    Feasibility

    • Data: every variable is published by a government body and downloadable free of cost.
    • Tools: R and RStudio are free; the econometrics needed (OLS, robust errors, diagnostics) is covered in Introductory Econometrics.
    • Time: two semesters — Sem 7 for review, data assembly and a pilot regression; Sem 8 for final models, robustness checks and writing — fit the UGCF 6 + 6 credit structure.
    • Risk: the main risk is inconsistent district names across sources; a concordance table is built in the first month to manage it.
  3. 1 min

    Problem statement

    MGNREGA is meant to provide work when and where it is needed, yet public data show large differences in the average number of days households receive across districts with similar poverty levels and identical wage rates. Policy discussions often attribute these gaps either to demand ("people there don't need the work") or to supply ("the administration doesn't open works"), usually without systematic evidence.

    For Uttar Pradesh and Bihar — two states that together account for a large share of India's multidimensionally poor — there is little recent district-level analysis that links MGNREGA outcomes to measurable indicators of need. This dissertation assembles a 113-district dataset from official sources and uses regression analysis to test whether work outcomes are associated with poverty, the agricultural labour share, social composition and monsoon shocks, and whether these associations differ between the two states. The results indicate whether the programme's outcomes are consistent with its self-targeting design and where closer field study is warranted.

  4. 1 min

    Objectives & scope

    1. 01To build a clean district-level dataset for FY 2022-23 combining MGNREGA MIS, Census 2011, NITI Aayog MPI 2023 and IMD rainfall data for 113 districts.
    2. 02To describe the distribution of person-days per working household and 100-day completion across districts of Uttar Pradesh and Bihar.
    3. 03To estimate the association between MGNREGA outcomes and poverty (MPI headcount), agricultural labour share, SC/ST share and female literacy.
    4. 04To test whether monsoon rainfall deficits are associated with higher same-year work demand.
    5. 05To test whether these associations differ between the two states using a state dummy and interaction terms.
    6. 06To carry out standard diagnostic and robustness checks and interpret the results in terms of the self-targeting argument.

    Scope

    In scope

    • All districts of Uttar Pradesh (75) and Bihar (38) as they existed in FY 2022-23.
    • A single-year cross-section (FY 2022-23), with FY 2021-22 used only as a robustness check.
    • Public secondary data; OLS regression with robust standard errors and standard diagnostics in R.
    • Interpretation in terms of need-based (self-targeting) versus capacity-based explanations.

    Out of scope

    • Household-level or village-level analysis, and any primary survey.
    • Causal identification strategies (instrumental variables, difference-in-differences) — noted as future work.
    • Wage-payment delays, social audits and corruption, which need different data.
    • Districts created after 2011 whose Census values cannot be separated cleanly (handled by merging with the parent district and flagging them).
  5. 2 min

    Methodology

    Research design

    A quantitative, cross-sectional, correlational design using secondary data at the district level. The unit of analysis is the district; the reference year is FY 2022-23.

    Variables and sources

    VariableDefinitionSource
    PD_HH (outcome 1)Total person-days ÷ households that workedMGNREGA MIS, R5 physical-progress reports
    HH100 (outcome 2)% of working households completing 100 daysMGNREGA MIS
    MPI_HCMultidimensional poverty headcount ratio (%)NITI Aayog, National MPI 2023 (NFHS-5 based)
    AGLABAgricultural labourers as % of total workersCensus 2011, B-series tables
    SCSTSC + ST population as % of totalCensus 2011, Primary Census Abstract
    FLITFemale literacy rate (%)Census 2011
    RAIN_DEVJune–September rainfall departure from normal (%)IMD district rainfall statistics
    BIHARState dummy (1 = Bihar)—

    Model

    PD_HH_i = b0 + b1 MPI_HC_i + b2 AGLAB_i + b3 SCST_i + b4 FLIT_i
              + b5 RAIN_DEV_i + b6 BIHAR_i + e_i
    

    Model 2 replaces the outcome with HH100; Model 3 adds BIHAR × MPI_HC to test whether the poverty association differs by state.

    Estimation and diagnostics

    • OLS with HC1 heteroskedasticity-robust standard errors (Breusch–Pagan test reported).
    • VIF for multicollinearity (MPI and female literacy are expected to correlate).
    • Ramsey RESET for functional form; log specification as an alternative.
    • Cook's distance to identify influential districts; results re-estimated without them.
    • Robustness: FY 2021-22 outcomes; dropping merged post-2011 districts.

    Timeline (two semesters, monthly progress reports)

    MonthWorkProgress report to supervisor
    1 (Sem 7)Topic, reading list, research questionsProposal note
    2Literature review draft; variable listReview chapter draft
    3Data download, district concordanceData log + concordance table
    4Cleaning, descriptive statistics, mapsDescriptive chapter
    5Pilot regressions; mid-term presentationSem 7 evaluation
    6–7 (Sem 8)Final models, diagnostics, robustnessResults tables
    8Discussion and conclusionFull draft
    9Revisions, formatting, similarity checkFinal submission and viva
  6. 1 min

    Architecture & tech stack

    • R 4.x + RStudio
    • tidyverse, readxl, janitor
    • lmtest, sandwich, car (diagnostics)
    • modelsummary / stargazer (regression tables)
    • MS Excel (data assembly)
    • MGNREGA MIS public reports
    • Census 2011 district tables
    • NITI Aayog National MPI 2023
    • IMD district rainfall statistics

    The study design runs from the research question to data assembly, estimation and interpretation. This flowchart is the "research design" figure for the Methodology chapter.

    flowchart TD
      Q["Research question: what explains MGNREGA work across districts?"] --> L[Literature review]
      L --> H["Hypotheses: self-targeting vs capacity"]
      H --> D1["MGNREGA MIS FY 2022-23"]
      H --> D2["Census 2011 district tables"]
      H --> D3["NITI Aayog MPI 2023"]
      H --> D4["IMD monsoon rainfall departure"]
      D1 --> C[District concordance and merge in R]
      D2 --> C
      D3 --> C
      D4 --> C
      C --> E["Descriptives, correlation matrix, maps"]
      E --> M["OLS models 1-3 with robust SE"]
      M --> G{Diagnostics pass?}
      G -- No --> R["Re-specify: logs, drop influential districts"]
      R --> M
      G -- Yes --> I[Interpretation and policy discussion]

    Data pipeline

    Raw files are never edited by hand. Each source is saved in data/raw/ with its download date, read into R, standardised (lower-case district names, spelling fixes through a concordance table) and joined on a district key. The concordance table is the most important artefact of the project: it documents every name mismatch (for example, spelling variants between the MIS and the Census) and every district created after 2011 that had to be merged with its parent district for Census variables.

    flowchart LR
      RAW[("data/raw: CSV and XLSX as downloaded")] --> R1["01_clean.R: read, rename, fix names"]
      CON[("concordance.csv")] --> R1
      R1 --> AN[("data/analysis.csv: 113 rows")]
      AN --> R2["02_describe.R: tables and plots"]
      AN --> R3["03_models.R: OLS, robust SE, diagnostics"]
      R3 --> OUT["tables/*.docx, figures/*.png"]

    Why OLS and not something fancier?

    The research question is descriptive-associational and the sample is 113 districts. OLS with robust errors is transparent, easy to diagnose and exactly what the Introductory Econometrics syllabus prepares a student to defend. Causal methods are discussed as limitations and future work.

  7. 6 modules

    Modules

    • Literature Review and Framework

      Review the design of MGNREGA, the self-targeting argument, evidence on rationing and administrative capacity, and earlier district- and household-level studies; end with the gap and testable hypotheses for each predictor.

    • Data Assembly and Concordance

      Download the MIS reports, Census tables, MPI annexure and IMD rainfall statistics; record source, table name and download date for each; build the district concordance and merge everything into one 113-row analysis file in R.

    • Descriptive Analysis

      Produce summary statistics by state, histograms of person-days per household, a correlation matrix, scatterplots of each predictor against the outcome, and a ranked table of the ten highest and lowest districts.

    • Regression and Diagnostics

      Estimate Models 1–3 by OLS with HC1 robust standard errors, report VIF, Breusch–Pagan and RESET tests, inspect Cook's distance, and present results in a single regression table with coefficients, robust SEs, R² and N.

    • Robustness Checks

      Re-estimate using FY 2021-22 outcomes, a log-transformed outcome, and a sample excluding merged post-2011 districts and influential observations; summarise which coefficients remain stable in sign and significance.

    • Discussion and Policy Implications

      Interpret findings against the self-targeting argument, discuss the UP–Bihar difference, state limitations such as the 2011 Census lag and ecological inference, and propose field-level questions for future research.

  8. Locked

    Presentation

    12 slides with speaker notes. The outline below is free; the bullets, notes and the generated .pptx unlock with the project.

    1. What Drives MGNREGA Work Demand?
    2. Motivation
    3. The Self-Targeting Idea
    4. Literature and Gap
    5. Research Questions
    6. Data
    7. Model and Method
    8. Descriptive Findings
    9. Regression Results
    10. Robustness
    11. Discussion and Limitations
    12. Conclusion and Future Work

    Bullets, speaker notes and the .pptx download unlock with the project.

    Presentation is locked: 12 slides, Speaker notes, .pptx download.

  9. 1 min

    Future scope

    • Panel data: extend to FY 2014-15 to FY 2023-24 and estimate district fixed-effects models to control for time-invariant administrative differences.
    • Causal design: use rainfall shocks as an instrument for agricultural distress, or compare districts before and after a policy change.
    • Spatial econometrics: test for spatial autocorrelation (Moran's I) and estimate a spatial lag model with the spdep package.
    • Field validation: follow up two high- and two low-performing districts with interviews of gram rozgar sevaks and job-card holders.
    • Gender lens: model women's share of person-days alongside the female labour-force indicators from PLFS.
  10. 9 sources

    References

    1. Ministry of Rural Development — Mahatma Gandhi NREGA public data portal and MIS reports
    2. The Mahatma Gandhi National Rural Employment Guarantee Act, 2005 (Act No. 42 of 2005)
    3. Office of the Registrar General & Census Commissioner, India — Census 2011 tables
    4. NITI Aayog (2023). National Multidimensional Poverty Index: A Progress Review 2023
    5. India Meteorological Department — rainfall statistics
    6. Wooldridge, J. M. Introductory Econometrics: A Modern Approach, 7th ed. Cengage.
    7. Gujarati, D. N., & Porter, D. C. Basic Econometrics, 5th ed. McGraw-Hill.
    8. Drèze, J., & Sen, A. (2013). An Uncertain Glory: India and its Contradictions. Allen Lane / Penguin.
    9. Wickham, H., Çetinkaya-Rundel, M., & Grolemund, G. (2023). R for Data Science, 2nd ed. O'Reilly.

    Cite this bundle

    OnlyProjects. (2026). What Drives MGNREGA Work Demand? A District-Level Regression Study of Uttar Pradesh and Bihar (R): B.A. Economics project bundle [Educational resource]. https://onlyprojects.online/projects/ba-economics-mgnrega-work-demand-district-regression-up-bihar

Slides, diagrams & files

12 slides. Titles are free; bullets, speaker notes and the .pptx unlock with the project.

  1. SLIDE 1

    What Drives MGNREGA Work Demand?

  2. SLIDE 2

    Motivation

  3. SLIDE 3

    The Self-Targeting Idea

  4. SLIDE 4

    Literature and Gap

  5. SLIDE 5

    Research Questions

  6. SLIDE 6

    Data

  7. SLIDE 7

    Model and Method

  8. SLIDE 8

    Descriptive Findings

  9. SLIDE 9

    Regression Results

  10. SLIDE 10

    Robustness

  11. SLIDE 11

    Discussion and Limitations

  12. SLIDE 12

    Conclusion and Future Work

Architecture diagrams · 2

1
flowchart TD
  Q["Research question: what explains MGNREGA work across districts?"] --> L[Literature review]
  L --> H["Hypotheses: self-targeting vs capacity"]
  H --> D1["MGNREGA MIS FY 2022-23"]
  H --> D2["Census 2011 district tables"]
  H --> D3["NITI Aayog MPI 2023"]
  H --> D4["IMD monsoon rainfall departure"]
  D1 --> C[District concordance and merge in R]
  D2 --> C
  D3 --> C
  D4 --> C
  C --> E["Descriptives, correlation matrix, maps"]
  E --> M["OLS models 1-3 with robust SE"]
  M --> G{Diagnostics pass?}
  G -- No --> R["Re-specify: logs, drop influential districts"]
  R --> M
  G -- Yes --> I[Interpretation and policy discussion]
2
flowchart LR
  RAW[("data/raw: CSV and XLSX as downloaded")] --> R1["01_clean.R: read, rename, fix names"]
  CON[("concordance.csv")] --> R1
  R1 --> AN[("data/analysis.csv: 113 rows")]
  AN --> R2["02_describe.R: tables and plots"]
  AN --> R3["03_models.R: OLS, robust SE, diagnostics"]
  R3 --> OUT["tables/*.docx, figures/*.png"]

Files

Viva questions & answers

3 of 16 questions free. Explain each answer in your own words before you move on.

  1. Concept

    What is meant by the self-targeting nature of MGNREGA?

    Self-targeting means the programme does not need to identify the poor administratively; because the work is manual and paid at a modest wage, mainly households with few better options will ask for it. If this holds, districts with more distress should show higher work demand, which is what my regressions test.

  2. Concept

    Why did you use robust standard errors?

    District-level cross-sections usually show heteroskedasticity — the spread of person-days is larger in some kinds of districts than others. OLS coefficients remain unbiased, but the usual standard errors become unreliable. HC1 robust errors correct the inference, and I also report the Breusch–Pagan test to show why they were needed.

  3. Concept

    How do you interpret the coefficient on MPI headcount?

    Holding the other variables constant, it is the expected change in average person-days per working household for a one percentage point higher multidimensional poverty headcount in the district. It is an association across districts, not the effect of making a particular district poorer.

+13 more questions

They and the answers unlock with the project. Try answering the ones above yourself first. Your examiner will.

For educational purposes only. Use this bundle to understand how the project works, then build and write your own. Submitting it verbatim is between you, your conscience and your external examiner.