TY - GEN
T1 - Yet another approach for completing missing values
AU - Ben Othman, Leila
AU - Ben Yahia, Sadok
PY - 2008
Y1 - 2008
N2 - When tackling real-life datasets, it is common to face the existence of scrambled missing values within data. Considered as "dirty data", it is usually removed during the pre-processing step of the KDD process. Starting from the fact that "making up this missing data is better than throwing it away", we present a new approach trying to complete the missing data. The main singularity of the introduced approach is that it sheds light on a fruitful synergy between generic basis of association rules and the topic of missing values handling. In fact, beyond interesting compactness rate, such generic association rules make it possible to get a considerable reduction of conflicts during the completion step. A new metric called "Robustness" is also introduced, and aims to select the robust association rule for the completion of a missing value whenever a conflict appears. Carried out experiments on benchmark datasets confirm the soundness of our approach. Thus, it reduces conflict during the completion step while offering a high percentage of correct completion accuracy.
AB - When tackling real-life datasets, it is common to face the existence of scrambled missing values within data. Considered as "dirty data", it is usually removed during the pre-processing step of the KDD process. Starting from the fact that "making up this missing data is better than throwing it away", we present a new approach trying to complete the missing data. The main singularity of the introduced approach is that it sheds light on a fruitful synergy between generic basis of association rules and the topic of missing values handling. In fact, beyond interesting compactness rate, such generic association rules make it possible to get a considerable reduction of conflicts during the completion step. A new metric called "Robustness" is also introduced, and aims to select the robust association rule for the completion of a missing value whenever a conflict appears. Carried out experiments on benchmark datasets confirm the soundness of our approach. Thus, it reduces conflict during the completion step while offering a high percentage of correct completion accuracy.
KW - Data mining
KW - Formal concept analysis
KW - Generic association rule bases
KW - Missing values completion
U2 - 10.1007/978-3-540-78921-5_10
DO - 10.1007/978-3-540-78921-5_10
M3 - Article in proceedings
AN - SCOPUS:41549121425
SN - 3540789200
SN - 9783540789208
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 155
EP - 169
BT - Concept Lattices and Their Applications - Fourth International Conference, CLA 2006, Selected Papers
PB - Springer
T2 - 4th International Conference on Concept Lattices and Their Applications, CLA 2006
Y2 - 30 October 2006 through 1 November 2006
ER -