- The paper introduces HICM, a framework that corrects nomenclature, coverage, and measurement biases in Indian Census migration data.
- It employs rigorous methods including temporal smoothing and sociologically-informed imputation to ensure statistical consistency.
- Empirical validation and network analysis demonstrate that harmonization preserves migration patterns and aids robust policy and demographic research.
HICM: A Data-Centric Framework for Harmonizing Indian Census Migration Data
Introduction
Indian census-based migration data are fraught with known inconsistencies, including missing values and structural biases across decadal rounds. These disparities affect analyses that rely on longitudinal consistency and comparability, particularly for interstate migration network studies central to population, urbanization, and policy research. The paper "HICM: An approach towards Harmonizing Indian Census Migration data and its applications" (2604.12324) introduces HICM, a systematic, reproducible framework that resolves these issues by harmonizing the core D-02 series of Indian Census migration tables (1991, 2001, 2011). The work provides a suite of mathematically rigorous tools and imputation strategies targeting representativeness and measurement bias. Furthermore, the authors validate the practical utility of harmonized datasets through empirical analyses and network-based temporal migration applications.
Diagnosis of Bias and Inconsistency Profiles
The paper identifies three primary sources of systematic bias:
- Measurement Bias: Manifested as unclassifiable migrants (with unknown last residence) and missing duration-of-stay attributes. The former accounts for less than 1% of migrants per decade, while the latter, "duration not stated," affects 8–16% of records—a substantial level that distorts duration-based migration analyses.
- Coverage Inconsistency: The 1991 dataset omits Jammu and Kashmir (J&K) from reported in-migrant flows due to historical enumeration disruptions, breaking the continuity of state-wise migration networks and introducing representativeness bias.
- Nomenclature and Structural Asymmetries: Disparities arise from differing indexing schemes, variable naming conventions, and inconsistent inclusion of aggregate totals across decades, increasing the risk of computational and analytical errors.
Harmonization Strategies
The HICM framework operationalizes three domain-informed correction layers:
- Nomenclature and Structure Correction: Standardizes state/UT names, sequential indexing, and inserts aggregate totals for the 1991 set, aligning the structural schema with later rounds and reducing the risk of propagation of indexing and mapping errors.
- Coverage Imputation via Temporal Smoothing: The authors propose a backward-temporal smoothing algorithm to fill unreported J&K in-migrant data for 1991, leveraging the observed ratios and their geometric mean across contiguous decades. This process exploits the empirical stability of migratory linkages and mitigates abrupt discontinuities in reconstructed profiles.

Figure 1: Kernel density estimation shows close concordance between ground-truth and imputed 1991 inflow distributions for major recipient states, validating the harmonization method.
Figure 2: Violin plots reveal the decade-wise distributional consistency after imputing previously missing 1991 inflows such as those for Jammu and Kashmir.
- Measurement Bias Correction for Missing Values:
Empirical Validation
The harmonized datasets are validated through statistical summaries and distributional overlaps before and after imputation. Key findings include:
- Proportional redistribution and temporal smoothing produce negligible shifts in the mean, quartiles, and variance of in-migrant and out-migrant statistics, except at the far extremes, confirming bias correction without artificial inflation of migration stocks.
- Imputation quality is visualized by close correspondence between the fitted and original distributions for high-migration states, as shown in Figure 1.
- Duration-based redistributions yield a more accurate demographic and policy-relevant stratification without masking structural patterns.
Applications to Migration Network Analysis
To demonstrate practical advances, the harmonized datasets are used to generate temporally coherent interstate migration networks for each decade. The resulting networks exhibit expected trends:
- Consistent Node Structure: After imputation, 1991 is realigned to a 35-node network, consistent with 2001 and 2011.
- Temporal Evolution: The average edge weight increases by ~30% per decade. The dominant roles of Maharashtra (in-migration) and Uttar Pradesh (out-migration) persist over thirty years.
- Community Structure: Louvain clustering of migration networks reveals both stable and evolving community profiles. Transition from three to four migration communities between 1991 and 2001/2011 signals substantive regime change, a pattern recoverable only with harmonized data.


Figure 4: Community assignments for Indian interstate migration networks, illustrating the temporal evolution and realignment of migration clusters over three decades.
Imputation enables robust analysis of dynamic properties such as centrality, flow asymmetry, and hierarchical structure, which are sensitive to missing edges and nodes. Analysis of short-duration migrants (residence < 1 year) highlights distinct community transitions otherwise obscured by raw data incompleteness.
Figure 5: Community assignments before and after imputation for the 1991 short-duration migrant network, demonstrating that harmonization enables consistent structural detection.
Implications and Future Directions
The HICM approach yields several theoretical and applied benefits:
- Longitudinal Network Consistency: By harmonizing temporal and structural properties, the datasets now enable high-fidelity longitudinal modeling and detection of regime shifts in migration flows.
- Policy Analysis: Harmonized migration tables underpin equity-aware planning and regional allocation models, critical for state and national policy, given the differential completeness in legacy data.
- Extensibility: The framework generalizes to non-interstate migration contexts (e.g., intrastate, rural/urban stratification) and can be adapted to additional demographic stratifiers.
- Grounded Imputation: The selective use of sociological priors and constraint preservation avoids the pitfalls of distributionally naive imputation, preserving salient network features for downstream network science and fairness-aware modeling.
As future census rounds grow in size and complexity, frameworks like HICM provide essential blueprints for responsible data correction, enabling reproducible, robust, and analytically sound migration research.
Conclusion
The paper presents a comprehensive, mathematically grounded protocol for the harmonization of Indian census migration data across three consecutive decades. By explicitly diagnosing measurement and structural inconsistencies and providing validated correction strategies, the HICM framework ensures longitudinal comparability and statistical fidelity. The resultant datasets empower rigorous migration network research and foster high-impact policy and demographic insights, setting a foundation for both future census harmonization and the development of methodologically transparent migration analytics.