Computational Intelligence for Missing Data Imputation, Estimation, and Management: Knowledge Optimization Techniques

This chapter develops and compares the merits of three different data imputation models by using accuracy measures. The three methods are auto-associative neural networks, a principal component analysis and support vector regression all combined with cultural genetic algorithms to impute missing variables. The use of a principal component analysis improves the overall performance of the auto-associative network while the use of support vector regression shows promising potential for future investigation. Imputation accuracies up to 97.4% for some of the variables are achieved.
The problem with data collection in surveys is that the data invariably suffers from some loss of information. This may for example be a consequence of problems such as incorrect data entry, or unfilled fields in surveys. This chapter explores three different methods for data imputation. These are the combination of cultural genetic algorithms with three learning methods i.e., neural networks, a principal component analysis and support vector regression.
The general approach pursued in this chapter is to use regression models to model the inter-relationships between data variables using neural networks (Chang and Tsai, 2008), a principal component analysis (Adams et al, 2002) and a support vector regression (Cheng, Yu, & Yang, 2007). Thereafter, a controlled and planned approximation of missing data is conducted using an optimization method, in this chapter a cultural genetic algorithm (Yuan and Yuan, 2006) is selected.
Data imputation using auto-associative neural networks as a regression model has been conducted as explained in earlier chapters by Abdella and Marwala (2006); Abdella (2005);