Computational Intelligence for Missing Data Imputation, Estimation, and Management: Knowledge Optimization Techniques

A number of techniques for handling missing data have been presented and implemented. Most of these proposed techniques are unnecessarily complex and, therefore, difficult to use. This chapter investigates a hot-deck data imputation method, based on rough set computations. In this chapter, characteristic relations are introduced that describe incompletely specified decision tables and then these are used for missing data estimation. It has been shown that the basic rough set idea of lower and upper approximations for incompletely specified decision tables may be defined in a variety of different ways. Empirical results obtained using real data are given and they provide a valuable insight into the problem of missing data. Missing data are predicted with an accuracy of up to 99%.
There are a number of general ways that have been used to approach the problem of missing data in databases (Little & Rubin, 1987: Rubin, 1976; Little & Rubin, 1989; Collins, Schafer Kam, 2001; Schafer & Graham, 2002). One of the simplest of these methods is the 'list-wise deletion', which simply deletes instances with missing values (Scheuren, 2005; King et al., 2001; Abdella, 2005). The major disadvantage of this method is the dramatic loss of information in data sets (King et al., 1988). Also, Enders and Peugh (2004) demonstrated that when there is a group of missing data, list-wise deletion gives biased parameters and standard errors. Another approach is 'pair-wise deletion' (Marsh, 1998).
Tsikritis (2005) observed that appropriately dealing with missing data has been underestimated by the...