Sources
Sources referenced by the active district dataset and its state/national context. Inclusion here is not proof that every source covers every district or period. Each source keeps its own licence.
Census of India 2011
Source ID: 1. Office of the Registrar General & Census Commissioner, Government of India
Govt. of India, open
Census of India 2011, Office of the Registrar General & Census Commissioner, Government of India. Open government data. The district and state tables are ingested from public compilations of the published census tables rather than a direct download from censusindia.gov.in.
District table: data/raw/india-districts-census-2011.csv (640 districts x 118 columns; via nishusharma1608 compilation). State table: data/raw/census_states_2011.csv (36 states/UTs; via covid19india deep-dive). Attribution given in the site footer.
NFHS-5 (2019-21) district factsheets
Source ID: 3. International Institute for Population Sciences (IIPS), for the Ministry of Health & Family Welfare
GODL-India
Contains information from the Ministry of Health & Family Welfare / IIPS, National Family Health Survey 5 (2019-21), obtained from data.gov.in and made available under the Government Open Data License - India (GODL).
MoHFW + IIPS compiled district XLS (resource uuid cf80173e-fece-439d-a0b1-6e9cb510593d), 111 fields. GODL attribution string required on the site. Factsheet PDFs (road 2): rchiips.org/nfhs/NFHS-5_FCTS/<STATE>/<District>.pdf via Wayback while rchiips is down. Licence verified in W2 spike 2026-07-28.
NFHS-5 (2019-21) state/UT factsheets
Source ID: 4. International Institute for Population Sciences (IIPS), for the Ministry of Health & Family Welfare
GODL-India
Contains information from the Ministry of Health & Family Welfare / IIPS, National Family Health Survey 5 (2019-21), obtained from data.gov.in and made available under the Government Open Data License - India (GODL).
MoHFW + IIPS compiled state/UT XLS (OGD resource uuid 7c568619-b9b4-40bb-b563-68c28c27a6c1, NFHS_5_Factsheets_Data.xls, 179712 bytes, 136 fields). Covers India + 36 states/UTs x Urban/Rural/Total. Licence verified 2026-07-29: the resource carries no per-resource licence override, so the site-wide OGD terms apply verbatim -- "The content published on data.gov.in is owned by the respective Ministry/State/Department/Organization and licensed under the Government Open Data License - India". GODL attribution string required on the site. NOT titled provisional (unlike the district resource); definitions of record are the NFHS-5 State Factsheet Compendium Phase-I/II PDFs.
NFHS-4 (2015-16) district factsheets (compiled)
Source ID: 34. International Institute for Population Sciences (IIPS), for the Ministry of Health & Family Welfare
Govt. of India (IIPS / MoHFW) publication; compilation MIT
National Family Health Survey 4 (2015-16) district fact sheets, International Institute for Population Sciences (IIPS) for the Ministry of Health & Family Welfare, Government of India. Ingested from a public compilation of the published fact sheets (February 2016 first release) rather than a direct IIPS download.
W14 dataset 1, 2026-08-03. LICENCE, IN FULL: the compilation is MIT (Copyright (c) 2017 Hindustan Times, verified against the LICENSE file). The official fact sheets carry no reuse statement and the IIPS host page footer says "Copyright (c) 2009 National Family Health Survey. All Rights Reserved" -- no open licence upstream. The file is HindustanTimesLabs/nfhs-data, nfhs_district-wise.csv, scraped by its authors from the IIPS district fact sheet PDFs. We rely on the underlying Government of India publication, not on the compiler: no permission is claimed beyond MIT and no rights of theirs are asserted over the numbers. VINTAGE: this is the FEBRUARY 2016 FIRST RELEASE; IIPS reissued the sheets in July 2018 with a minority of cells revised (verified: 36 of 2,232 cells over 8 districts). Same posture as the 2011 census ingest (migration 057). FLAGGED, NOT WAVED THROUGH, the owner prefers genuinely free/open sources and a third-party scrape is a retrieval route of last resort; replace it the moment an official IIPS or data.gov.in district file is located, and run NFHS-4 road 2 (the 640 July-2018 PDFs on the Wayback Machine, keyed by the census code in each filename) to supersede the vintage. SHAPE: long-format CSV, CP1252, bare-CR line endings, 637 of the 640 census-2011 districts, Chandigarh (55), Dadra & Nagar Haveli (496) and Lakshadweep (587) are absent from the compilation. Joins to the spine on the 2011 census district code, mechanically: no resolver, no aliases, no name matching. See warehouse/tools/standardize_nfhs4.py for the join gate.
Crop Area, Production & Yield (APY, DES-Agri)
Source ID: 35. Directorate of Economics & Statistics, Department of Agriculture & Farmers Welfare
ODC-BY (Open Data Commons Attribution License)
Area, Production and Yield (APY) crop statistics, Directorate of Economics & Statistics, Department of Agriculture & Farmers Welfare, Government of India; obtained via the India Data Portal under the Open Data Commons Attribution License.
W14 dataset 2, 2026-08-03. Licence verified against the CKAN API (package area-production-yield-apy): license_title "Open Data Commons Attribution License", package url https://data.desagri.gov.in/login. ODC-BY requires attribution and imposes NO non-commercial and NO share-alike term, so unlike SHRUG nothing derived from it constrains the site. RETRIEVAL: the CKAN resource UUID rotates, so re-fetch by resolving https://ckandev.indiadataportal.com/api/3/action/package_show?id=area-production-yield-apy and reading resources[].url, never by pinning the UUID. GRAIN: year x state x district x crop x season, district_code is the LGD code. TRAP: season = 'Total' is a SUBTOTAL of the seasonal rows (3,146 crop-district pairs in 2022-2023, reconciling in 3,145 of them) and is excluded by the standardizer; summing it in would double-count rice in 363 districts. NO suppression markers exist in this file: it is an administrative compilation, not a sample survey.
MGNREGA district monthly physical & financial progress (MoRD)
Source ID: 41. Ministry of Rural Development, Government of India (MGNREGA / NREGASoft)
Government Open Data Licence - India (GODL). MoRD/NREGA publishes the scheme's physical and financial progress as open data; attribution required, no NC and no ShareAlike.
Ministry of Rural Development, Government of India (MGNREGA / NREGASoft), via data.gov.in
data/raw/mgnrega_api/, 42 gzipped JSON pages + manifest.jsonl, 415,834 records pulled 2026-08-04T12:02Z by a background script that is NOT in this repository (wiki/raw-inventory.md flags the pulls as un-reproducible from a clean checkout). Each page is registered as its own raw.artifact. PRECEDENCE 60, deliberately below the 80 an official portal earns: data.gov.in showed a "sandbox environment" banner on the pull date and no sample has been re-verified against nrega.nic.in. Measures are FY cumulative-to-date, not monthly flows.
Udyam MSME registrations (Ministry of MSME)
Source ID: 43. Ministry of Micro, Small and Medium Enterprises, Government of India
Government Open Data Licence - India (GODL). The Ministry of MSME publishes the Udyam registration register as open data through data.gov.in; attribution required, no NC and no ShareAlike.
Ministry of Micro, Small and Medium Enterprises, Government of India (Udyam Registration Portal), via data.gov.in
data/raw/udyam_api/, 4,700 gzipped JSON pages + manifest.jsonl, 42,629,074 enterprise rows pulled 2026-08-04T11:59Z district by district BY LGD code, by a background script that is NOT in this repository (wiki/raw-inventory.md flags the API pulls as un-reproducible from a clean checkout). Every page is registered as its own raw.artifact through fetcher.py --from-file with this resource page as the uri. PRECEDENCE 60, deliberately below the 80 an official portal earns: data.gov.in showed a "sandbox environment" banner around the pull date and no sample has been re-verified against udyamregistration.gov.in. AS-OF SNAPSHOT of a live register, not a period return, a re-pull SUPERSEDES this one at a new period rather than contradicting it. UNIT-LEVEL ROWS ARE PERSONAL DATA AND NEVER LEAVE ing.
Census of India table A-2 (decadal variation in population, 1901-2011)
Source ID: 52. Office of the Registrar General & Census Commissioner, Government of India
Government of India, Office of the Registrar General & Census Commissioner. Published on the census website's table catalogue for public use. No explicit open-licence statement appears beside the tables, so this is recorded as "government publication, licence not stated" rather than claimed as GODL. Unlike SHRUG it carries no share-alike clause, so nothing downstream inherits an obligation.
Census of India 2011, table A-2 (Registrar General & Census Commissioner, India)
One Excel per state, 35 files, ten .xls and twenty-five .xlsx. Telangana is absent and that is correct rather than missing: it did not exist in 2011, so its districts appear inside Andhra Pradesh. The sheet is a REPORT rather than a table: a state or district name appears once and the eleven following rows carry only a year, so the name must be carried down. Codes are 2011 census state and district codes, which the LGD crosswalk already speaks (attrs->>'dtcode11'). The host's TLS chain does not verify for a stdlib client, so curl -k or equivalent is required. Parsed by warehouse/tools/load_census_a2.py.
SHRUG Census PCA panel 1991/2001/2011 (Development Data Lab)
Source ID: 53. Development Data Lab
Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0). "You are free to: Share — copy and redistribute the material in any medium or format; Adapt — remix, transform, and build upon the material. Under the following terms: Attribution — You must give appropriate credit, provide a link to the license, and indicate if changes were made. NonCommercial — You may not use the material for commercial purposes. ShareAlike — If you remix, transform, or build upon the material, you must distribute your contributions under the same license as the original." SHRUG additionally asks that users cite Asher, Lunt, Matsuura and Novosad (2021), "Development Research at High Geographic Resolution: An Analysis of Night-Lights, Firms, and Poverty in India using the SHRUG Open Data Platform", World Bank Economic Review 35(4). BOTH THE NC AND THE SA CLAUSE PROPAGATE to anything derived from these numbers, which is why migration 112 preferred the Registrar General's own A-2 tables for the population timeline. Recorded, never a reason to skip the dataset (owner ruling 2026-09-02).
Census of India 1991, 2001 and 2011 Primary Census Abstract, via SHRUG (Development Data Lab), CC BY-NC-SA 4.0.
The three census waves re-based by SHRUG onto CONSTANT 2011 district geometry, which is the whole reason to take them from here rather than from the Registrar General: a 1991 district and a 2011 district of the same name are frequently not the same place. Bridge: shrid2 -> (pc11_state, pc11_district) -> dtcode11 (sorted rank of the 640-row 2011 district file) -> dist_lgd via data/crosswalks/districts.json. Validated 640/640 on 2026-06-19, 2011 population within 0.009%. Loaded by warehouse/tools/standardize_shrug_census.py.
Census of India 2011 table A-1 (number of villages, towns, households, population and area)
Source ID: 54.
NO EXPLICIT LICENCE PUBLISHED. The only rights statement on the source page is the footer of the ORGI Digital Library: "ORGI Digital Library © ORGI Digital Library, All Rights Reserved." The table itself is a Government of India publication of the Office of the Registrar General & Census Commissioner, Ministry of Home Affairs, released for public download (618,853 downloads recorded on the catalogue page on 2026-09-03). Recorded as "government publication, licence not stated" rather than claimed as GODL — the same treatment source 52 (table A-2) already has. No share-alike and no non-commercial clause is asserted, so nothing downstream inherits an obligation from it; attribution is ours to give because it is right, not because a licence compels it.
Census of India 2011, table A-1 (Office of the Registrar General & Census Commissioner, India); area figures supplied by Survey of India at state and district level
One .xlsx, 20,024 rows, reference id PC11_A01. Geographic granularity is country / state / district / sub-district and every unit appears three times — Total, Rural and Urban — so the sheet fills `std.fact.area` as well as its value. Codes are 2011 CENSUS codes (state 01-35, district 001-640, sub-district 00001-05924), which the spine already speaks: state via attrs->>'stcode11', district via attrs->>'dtcode11', sub-district via attrs->'lgd_via_data_gov_in_api'->'lgd'->>'census2011_code'. AREA IS SURVEY OF INDIA'S AT STATE AND DISTRICT GRAIN AND ORGI'S OWN DIGITISATION AT SUB-DISTRICT GRAIN, by the workbook's note 2; the two pedigrees ride on every fact row's quality_note. Total is NOT the sum of Rural and Urban wherever a footnote says so (Arunachal Pradesh publishes no rural/urban split at all). The host's TLS chain does not verify for a stdlib client, so curl -k or equivalent is required — the same wall source 52 and the DCHB harvest already hit. Loaded by warehouse/tools/load_census_a01_area.py, promoted by warehouse/tools/standardize_census_a01_area.py.