Data Preparation for Data MiningData Preparation for Data Mining addresses an issue unfortunately ignored by most authorities on data mining: data preparation. Thanks largely to its perceived difficulty, data preparation has traditionally taken a backseat to the more alluring question of how best to extract meaningful knowledge. But without adequate preparation of your data, the return on the resources invested in mining is certain to be disappointing. Dorian Pyle corrects this imbalance. A twenty-five-year veteran of what has become the data mining industry, Pyle shares his own successful data preparation methodology, offering both a conceptual overview for managers and complete technical details for IT professionals. Apply his techniques and watch your mining efforts pay off-in the form of improved performance, reduced distortion, and more valuable results. On the enclosed CD-ROM, you'll find a suite of programs as C source code and compiled into a command-line-driven toolkit. This code illustrates how the author's techniques can be applied to arrive at an automated preparation solution that works for you. Also included are demonstration versions of three commercial products that help with data preparation, along with sample data with which you can practice and experiment. |
Contents
Introduction | 1 |
Data Exploration as a Process | 9 |
Chapter 2 | 45 |
Supplemental Material | 87 |
Chapter 4 | 125 |
Normalizing and Redistributing Variables | 239 |
Chapter 8 | 275 |
Supplemental Material | 286 |
Trend | 323 |
Chapter 10 | 351 |
1 | 402 |
Supplemental Material | 446 |
Chapter 12 | 483 |
Appendix | 505 |
513 | |
About the Author | 537 |
Other editions - View all
Common terms and phrases
References to this book
Exploratory Data Mining and Data Cleaning Tamraparni Dasu,Theodore Johnson No preview available - 2003 |