Series variables have a number of characteristics that are sufficiently different from other types of variables that they need examining in more detail. Series variables are always at least two-dimensional, although one of the dimensions may be implicit. The most common type of series variable is a time series , in which a series of values of some feature or event are recorded over a period of time. The series may consist of only a list of measurements, giving the appearance of a single dimension, but the ordering is by time, which, for a time series, is the implicit variable.
The series values are always measured on one of the scales already discussed, nominal through ratio, and are presented as an ordered list. It is the ordering, the expression of the implied variable, that requires series data to be prepared for mining using techniques in addition to those discussed for nonseries data. Without these additional techniques the miner will not be able to best expose the available information. This is because series variables carry additional information within the ordering that is not exposed by the techniques discussed so far.
Up to this point in the book we have developed precise descriptions of features of nonseries data and various methods for manipulating the identified features to expose information content. This chapter does the same for series data and so has two main tasks:
1.
Find unambiguous ways to describe the component features of...
Copyright Morgan Kauffmann Publishers, Inc. 1999 under license agreement with Books24x7