Data Preparation for Data Mining

Chapter 10: Preparing the Data Set

Chapter 10: Preparing the Data Set
Overview
In this chapter, the focus of attention is on the data set itself. Using the term ?data set? places emphasis on the interactions between the variables, whereas the term ?data? has implied focusing on individual variables and their instance values. In all of the data preparation techniques discussed so far, care has been taken to expose the information content of the individual variables to a modeling tool. The issue now is how to make the information content of the data set itself most accessible. Here we cover using sparsely populated variables, handling problems associated with excessive dimensionality, determining an appropriate number of instances, and balancing the sample.
These are all issues that focus on the data set and require restructuring it as a whole, or at least, looking at groups of variables instead of looking at the variables individually.
10.1 Using Sparsely Populated Variables
Why use sparsely populated variables? When originally choosing the variables to be included in the data set, any in which the percentage of missing values is too high are usually discarded. They are discarded as simply not having enough information to be worth retaining. Some forms of analysis traditionally discard variables if 10 or 20% of the values are missing. Very often, when data mining, this discards far too many variables, and the threshold is set far lower. Frequently, the threshold is set to only exclude variables with more than 80 or 90% of missing values, if...

UNLIMITED FREE
ACCESS
TO THE WORLD'S BEST IDEAS

SUBMIT
Already a GlobalSpec user? Log in.

This is embarrasing...

An error occurred while processing the form. Please try again in a few minutes.

Customize Your GlobalSpec Experience

Category: Battery Monitors and Testers
Finish!
Privacy Policy

This is embarrasing...

An error occurred while processing the form. Please try again in a few minutes.