Paper by Ramanpreet Kaur and Dušan Gabrijelčič in Electric Power Systems Research by Elsevier. Date: April 2027.

Abstract:

Insights generated by advanced analytics in smart grids are critical for delivering safe and reliable power and providing value-added services in a cost-effective manner. However, electricity consumption data collected from advanced metering infrastructure (AMI) often contain missing values due to measurement errors, communication failures, device outages, or other unknown causes. Thus, missing data imputation is critical for ensuring the reliability of data-driven analysis. The existing literature has applied various imputation methods for electricity consumption data, but systematic comparison across different method families, missingness rates and contiguous gap sizes remain limited. In this work, we conducted a comparative analysis of eleven widely used imputation methods, including persistence-based, fixed statistical methods, time-series based, and machine learning based approaches, using high-resolution electricity consumption data from Slovenian households. The methods are evaluated using normalized root mean square error (NRMSE) as the key performance metric. Statistical significance testing is further used to determine whether the observed performance differences between the leading methods are consistent across households and experimental settings. The results show that the performance of imputation methods is strongly influenced by both missingness rate and gap size. While linear interpolation performs well for very short gaps, methods capturing historical and contextual consumption patterns perform better for medium and long gaps. Among the eleven evaluated methods, XGBoost achieves the best overall performance, with the lowest average NRMSE. Pairwise Wilcoxon signed-rank tests further confirm that XGBoost significantly outperforms the leading competing methods in most evaluated configurations. The consumer level analysis also shows that the most accurate method on average is not always the most consistent across households. Thus, researchers and practitioners should consider gap size, temporal structure, and consumer-level variability, in addition to the percentage of missing values, when selecting imputation methods for smart meter data.