The presence of missing or empty values can make problems for the modeler. There are several ways of replacing these values, but the best are those that not only are well understood by the modeler as to their capabilities, limits, and dangers, but are also in the modeler?s control. Even replacing the values at all has its dangers unless it is carefully done so as to cause the least damage to the data. It is every bit as important to avoid adding bias and distortion to the data as it is to make the information that is present available to the mining tool.
The data itself, considered as individual variables, is fairly well prepared for mining at this stage. This chapter discusses a way to fill the missing values, causing the least harm to the structure of the data set by placing the missing value in the context of the other values that are present. To find the necessary context for replacement, therefore, it is necessary to look at the data set as a whole.
8.1 Retaining Information about Missing Values
Missing and empty values were first mentioned in Chapter 2 , and the difference between missing and empty was discussed there. Whether missing or empty, many, if not most, modeling tools have difficulty digesting such values. Some tools deal with missing and empty values by ignoring them; others, by using some metric to determine ?suitable? replacements. As with normalization (discussed in...
Copyright Morgan Kauffmann Publishers, Inc. 1999 under license agreement with Books24x7