SimFiller. Similarity-Based Missing Values Filling Algorithm

Fateh ur Rehman,Muhammad Abbas, Sajjad Murtaza,Wasi Haider Butt,Saad Rehman,Usman Qamar

2018 Thirteenth International Conference on Digital Information Management (ICDIM)(2018)

引用 2|浏览2
暂无评分
摘要
With the growth of heterogeneous data generation sources low-quality data volumes are expanding on a daily basis. This research proposed SimFiller: similarity-based missing (null) values filling algorithm, to enhance the quality of data for the data mining process. The proposed algorithm calculates the similarity of record pairs from the input data in such a way that at least one member of the pair has a non-null value for the attribute under consideration. After finding similar pairs, the algorithm fills the missing values by considering the pair having greatest similarity under the specified similarity threshold. The quality of resulted data is evaluated by analyzing the classification accuracy results for Audiology dataset. Five other missing values filling algorithms were selected and total six copies of filled Audiology dataset were created. All six copies of filled Audiology dataset were tested for their classification accuracy. Results show a huge boost in classification accuracy for the copy of the dataset filled with the proposed algorithm and indicate that the quality of the dataset is enhanced. The proposed algorithm can also be tested on other datasets for filling their missing (null) values and can also be extended to remove other inconsistencies from the datasets.
更多
查看译文
关键词
Data Preprocessing,Data Cleansing,Missing Values,Null Values,SimFiller,RM-RMVO (RapidMiner Replace Missing Values Operator)
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要