Generalized Regression-Based Imputation Using Multiple Auxiliary Variables for Missing Data in Survey Sampling
Code:JOSSDA:202612.00009
Authors:U. A. Huguma, O. O. Ishaq, S. A. Sabo, H. A. Hamisu, I. A. Isah, U. A. Abubakar, N. Yusuf, S. S. Suleiman, J. S. Ogheneovo, F. U. Muhammad, A. F. Kabara
Category:Survey Sampling
Publication date:2026-12-01
Keywords:auxiliary informationfinite population mean
Non-response remains a major source of bias and efficiency loss in sample surveys, particularly when conventional imputation procedures use only one auxiliary variable or assume that missingness is completely random. This paper develops a generalised regression (GREG)-based imputation estimator for the finite-population mean when the study variable is subject to missingness and several auxiliary variables are completely observed. Under a missing-at-random mechanism, the respondent data are used to estimate a multiple linear regression model, predictions are generated for nonrespondents, and the completed-data estimator is expressed in the familiar model-assisted GREG form. The estimator is shown to be approximately unbiased to the first order, while its mean squared error is governed by the residual variation after projecting the study variable on the auxiliary vector. Four finite-population data sets are used for empirical evaluation against established imputation estimators. The proposed estimator recorded percentage relative efficiencies of 443.26% for Population I, 184.61% for Populations II and III, and approximately 184.81% across the complete response-rate scenarios for Population IV. These values correspond to substantial reductions in mean squared error relative to the respondent-mean baseline. The results further show that absolute mean squared error declines as the number of respondents increases, whereas the relative advantage of the proposed estimator remains stable. The findings support multiple-auxiliary GREG imputation as an efficient and implementable approach for finite-population estimation under non-response, provided that the working regression model and missing-at-random assumption are reasonable.