Missing Data Handling by Mean Imputation Method and Statistical Analysis of Classification Algorithm

Maheswari, K.; Packia Amutha Priya, P.; Ramkumar, S.; Arun, M.

doi:10.1007/978-3-030-19562-5_14

K. Maheswari⁶,
P. Packia Amutha Priya⁶,
S. Ramkumar⁶ &
…
M. Arun⁶

Part of the book series: EAI/Springer Innovations in Communication and Computing ((EAISICC))

848 Accesses
3 Citations

Abstract

The motive of data mining is to extract meaningful information from the large database. Because of the human errors, their high dimensionality, noisy data, and missing values, the process over dataset may degrade the performance. Therefore, the need for handling of those data in a proper way is important for improving the performance. There are many missing data handling methods available. Mean imputation is one of the methods for missing data in the dataset. This is the preprocessing operation performed before applying any machine learning algorithms. After applying mean imputation in a dataset, the decision is made either imputed mean value is good or bad. The rpart decision tree algorithm is applied on retailer dataset to handle more number of classes. From the experimental results, there is no significant difference among variables. The results of various GLM models with different were compared and analyzed to provide better performance.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 129.00; Price excludes VAT (USA)

Softcover Book: USD 169.99; Price excludes VAT (USA)

Hardcover Book: USD 169.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

References

I. Pratama, A. Erna Permanasari, I. Ardiyanto, R. Indrayani, A review of missing values handling methods on time-series data, in International Conference on Information Technology Systems and Innovation (ICITSI) (IEEE, Piscataway, NJ, 2016), INSPEC Accession Number: 16675571
Google Scholar
M. Moore, J.M. Carpenter, A decision tree approach to modeling the private label apparel consumer. Mark. Intell. Plan. 28(1), 59–69 (2010)
Article Google Scholar
K. Maheswari, P. Packia Amutha Priya, Predicting customer behavior in online shopping using SVM classifier, in IEEE International Conference on Intelligent Techniques in Control, Optimization, Signal Processing, INCOS’17, 1 Mar 2018
Google Scholar
K. Maheswari, P. Packia Amutha Priya, Analysis and implementation of text mining for different documents. Int. J. Scient. Res. Sci. Technol. 3(5), 109–113 (2017). ISSN: 2395-6011
Google Scholar
K. Maheswari, P. Packia Amutha Priya, Classification of twitter data set using SVM and KSVM. Int. J. Pure Appl. Math. 118(7), 675–680 (2018). ISSN: 1311-8080 (printed version); ISSN: 1314-3395 (on-line version), Scopus Indexed
Google Scholar
K. Maheswari, Improving accuracy of sentiment classification analysis in twitter data set using knn. Int. J. Res. Anal. Rev. 5(1), 422–425 (2018). E ISSN: 2348-1269, Print ISSN: 2349-5138, UGC Approved Journal
Google Scholar
R. Tinabo, Decision tree technique for customer retention in retail sector, in International Conference on Integrated Computing Technology INTECH 2011 (Springer, Berlin, 2011), pp. 123-131
Google Scholar
M. Rafiqul Islam, M. Ahsan Habib, A data mining approach to predict prospective business sectors for lending in retail banking using decision tree. Int. J. Data Min. Knowl. Manag. Process (IJDKP) 5(2), 13–22 (2015)
Article Google Scholar
B.N. Patel, S.G. Prajapati, K.I. Lakhtaria, Efficient classification of data using decision tree. Bonfring Int. J. Data Min. 2(1), 6–12 (2012)
Article Google Scholar
L. Li, X. Zhang, Study of data mining algorithm based on decision tree, in 2010 International Conference on Computer Design and Applications, 5 Aug 2010, INSPEC Accession Number: 11523965
Google Scholar
R. Senapati, K. Shaw, S. Mishra, D. Mishra, A novel approach for missing value imputation and classification of microarray dataset. Proc. Eng. 38, 1067–1071 (2012)
Article Google Scholar
R. Houari, A. Bounceur, T. Kechadi, T. Abdelkamel, R. Euler, A new method for estimation of missing data based on sampling methods for data mining, in Advances in Intelligent Systems and Computing (AISC), vol. 225, (Springer, Cham, 2012), pp. 89–100
Google Scholar
P.R. Houcka, S. Mazumdarb, E. Hartin, Classification of missing values handling method during data mining: review. Sigma Epsilon 21(2), 49–60 (2017). ISSN: 0853-9103
Google Scholar
Y.-Y. Song, Y. Lu, Decision tree methods: applications for classification and prediction. Shanghai Arch. Psychiatry 27(2), 130–135 (2015)
Google Scholar
S. Agarwal, G.N. Pandey, M.D. Tiwari, Data mining in education: data classification and decision tree approach. Int. J. e-Education, e-Business, e-Management and e-Learning 2(2), 140–144 (2012)
Google Scholar
C. Kishor Kumar Reddy, B. Vijaya Babu, A survey on issues of decision tree and non-decision tree algorithms. Int. J. Artif. Intell. Appl. Smart Dev. 4(1), 9–32 (2016)
Google Scholar
X. Wu, V. Kumar, J. Ross Quinlan, J. Ghosh, Q. Yang, H. Motoda, G.J. McLachlan, A. Ng, B. Liu, P.S. Yu, et al., Top 10 algorithms in data mining. Knowl. Inf. Syst. 14(1), 1–37
Article Google Scholar
H. Yang, S. Fong, Optimized very fast decision tree with balanced classification accuracy and compact tree size, in 2011 Third International Conference on Data Mining and Intelligent Information Technology Applications (ICMiA), 24–26 Oct 2011, pp. 57–64
Google Scholar

Download references

Author information

Authors and Affiliations

School of Computing, Kalasalingam Academy of Research and Education, Virudhunagar, India
K. Maheswari, P. Packia Amutha Priya, S. Ramkumar & M. Arun

Authors

K. Maheswari
View author publications
You can also search for this author in PubMed Google Scholar
P. Packia Amutha Priya
View author publications
You can also search for this author in PubMed Google Scholar
S. Ramkumar
View author publications
You can also search for this author in PubMed Google Scholar
M. Arun
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

Department of Computer Science & Engineering, Sri Eshwar College of Engineering, Coimbatore, Tamil Nadu, India
Anandakumar Haldorai
Department of Computer Science & Engineering, Presidency University, Bengaluru, India
Arulmurugan Ramu
Sri Eshwar College of Engineering, Coimbatore, Tamil Nadu, India
Sudha Mohanram
Department of Electrical Engineering, Faculty of Engineering, University of Malaysia, Kuala Lumpur, Kuala Lumpur, Malaysia
Chow Chee Onn

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Maheswari, K., Packia Amutha Priya, P., Ramkumar, S., Arun, M. (2020). Missing Data Handling by Mean Imputation Method and Statistical Analysis of Classification Algorithm. In: Haldorai, A., Ramu, A., Mohanram, S., Onn, C. (eds) EAI International Conference on Big Data Innovation for Sustainable Cognitive Computing. EAI/Springer Innovations in Communication and Computing. Springer, Cham. https://doi.org/10.1007/978-3-030-19562-5_14

Download citation

DOI: https://doi.org/10.1007/978-3-030-19562-5_14
Published: 19 October 2019
Publisher Name: Springer, Cham
Print ISBN: 978-3-030-19561-8
Online ISBN: 978-3-030-19562-5
eBook Packages: EngineeringEngineering (R0)

Publish with us

Policies and ethics