Data Preparation for Data Mining

11.8 Novelty Detection

11.8 Novelty Detection
A novelty detector is not strictly part of the data survey, but is easily built from some of the various components that are discovered during the survey. The novelty detector is mainly used during the execution stage of using a model. Novelty detectors address the problem of ensuring that the execution data continues to resemble the training and test data sets.
Given a data set, it is moderately easy to survey it, or even to simply take basic statistics about the individual variables and the joint distribution, and from them to determine if the two data sets are drawn from the same population (to some chosen degree of confidence, of course). A more difficult problem arises when as assembled data set is not available, but the data to be modeled is presented on an instance-by-instance basis. Of course, each of the instances can be assembled into a data set, and that data set examined for similarity to the training data set, but that only tells you that the data set now assembled was or wasn?t drawn from the same population. To use such a method requires waiting until sufficient instances become available to form a representative sample. It doesn?t tell you if the instances arriving now are from the training population or not. And knowing if the current instance is drawn from the same population can be very important indeed, for reasons discussed in several places. In any case, if the distribution is not stationary (see

UNLIMITED FREE
ACCESS
TO THE WORLD'S BEST IDEAS

SUBMIT
Already a GlobalSpec user? Log in.

This is embarrasing...

An error occurred while processing the form. Please try again in a few minutes.

Customize Your GlobalSpec Experience

Category: Specialty Tapes
Finish!
Privacy Policy

This is embarrasing...

An error occurred while processing the form. Please try again in a few minutes.