Data Preparation for Data Mining

Chapter 7: Normalizing and Redistributing Variables

Chapter 7: Normalizing and Redistributing Variables
Overview
From this point on in preparing the data, all of the variables in a data set have a numerical representation. Chapter 6 explained why and how to find a suitable and appropriate numerical representation for alpha values?that is, the one that either reveals the most information, or at least does the least damage to existing information. The only time that an alpha variable?s label values come again to the fore is in the Prepared Information Environment Output module, when the numerical representations of alpha values have to be remapped into the appropriate alpha representation. The discussion in most of the rest of the book assumes that the variables not only have numerical values, but are also normalized across the range of 0?1. Why and how to normalize the range of a variable is covered in the first part of this chapter.
In addition to looking at the range of a variable, its distribution may also make problems. The way a variable?s values are spread, or distributed, across its range is known as its distribution . Some patterns in a variable?s distribution can cause problems for modeling tools. These patterns may make it hard or impossible for the modeling tool to fully access and use the information a variable contains. The second topic in this chapter looks at normalizing the distribution, which is a way to manipulate a variable?s values to alleviate some of these problems.
The chapter, then, covers two key...

UNLIMITED FREE
ACCESS
TO THE WORLD'S BEST IDEAS

SUBMIT
Already a GlobalSpec user? Log in.

This is embarrasing...

An error occurred while processing the form. Please try again in a few minutes.

Customize Your GlobalSpec Experience

Category: Control Valves
Finish!
Privacy Policy

This is embarrasing...

An error occurred while processing the form. Please try again in a few minutes.