Sampling bias is a major bugaboo and very hard to detect, but it?s easy to describe. When a sampling method repeatedly takes samples of data from a population that differ from the true population measures in the same way and in the same direction, then that method is introducing sampling bias . It is a distortion of the true values in the sample from those in the population that is introduced by the selection method itself, independent of other factors biasing the data. It is difficult to avoid since it may be quite unconsciously introduced. Since miners often work with data collected for purposes uncertain, by methods unknown, and with measurements obscure, after the fact detection of sampling bias may be all but impossible. Yet if the data does not reflect the real world, neither will any model mined, regardless of how assiduously it is checked against test and evaluation sample data sets.
The best that can be had from internal evaluation of a data set are clues that perhaps the data is biased. The only real answer lies in comparing the data with the world! However, that said, what can be done? There are two main types of sampling bias: errors of omission and errors of commission.
Errors of omission, of course, involve leaving out data that should be put in, whereas errors of commission involve putting in what should be left out. For instance, many interest groups seem to be able to prove a point...
Copyright Morgan Kauffmann Publishers, Inc. 1999 under license agreement with Books24x7