Robust Machine Learning Imputation Estimator Using Multiple Auxiliary Variables for Handling Missing Data in Survey Sampling
Code:JOSSDA:202612.00010
Authors:H. A. Hamisu, O. O. Ishaq, A. A. Osi, U. A. Huguma, N. Yusuf, S. S. Suleiman, J. S. Ogheneovo, F. U. Muhammad, A. F. Kabara
Category:Survey Sampling
Publication date:2026-12-01
Keywords:gradient boosting machineimputationmissing at randommultiple auxiliary variablesnonresponsesurvey sampling
Missing data caused by survey nonresponse can bias finite-population estimates and reduce statistical efficiency, particularly when conventional imputation methods impose linearity or depend on a single auxiliary variable. This study develops a robust machine-learning imputation estimator of the finite-population mean using Gradient Boosting Machines (GBM) and multiple auxiliary variables. Under simple random sampling without replacement and a missing-at-random mechanism, a GBM model is trained on responding units and used to predict the study variable for nonrespondents. The completed-sample estimator is expressed in prediction-error form, from which its approximate bias and mean squared error (MSE) are derived. The bias is proportional to the nonresponse fraction and the average GBM prediction bias, while the MSE comprises the sampling variance, unexplained residual variance, prediction variance, and squared prediction-bias penalty. A Monte Carlo study with 10,000 replications compares the proposed estimator with mean, regression, Lee (1994), Singh and Horn (2000), Singh and Deo (2003), Kadilar and Cingi (2008), Singh (2009), Gira (2015), Prasad (2017), and Singh et al. (2022) estimators under linear, nonlinear, outlier-contaminated, and skewed data conditions. The proposed GBM estimator attained the smallest MSE under linear and nonlinear conditions, with MSEs of 1.890 and 12.494, respectively, and remained the second-best method under outlier and skewed conditions, where Prasad's estimator was marginally superior. Its MSE decreased consistently as the response rate increased from 50% to 90%. The results demonstrate that machine-learning-assisted imputation can substantially improve finite-population mean estimation when several informative auxiliary variables and complex relationships are present, although robust benchmark procedures remain important for heavily contaminated or skewed data.