Applying Under-Sampling Techniques and Cost-Sensitive Learning Methods on Risk Assessment of Breast Cancer
Breast cancer is one of the most common cause of cancer mortality. Early detection through mammography screening could significantly reduce mortality from breast cancer. However, most of screening methods may consume large amount of resources. We propose a computational model, which is solely based on personal health information, for breast cancer risk assessment. Our model can be served as a pre-screening program in the low-cost setting. In our study, the data set, consisting of 3976 records, is collected from Taipei City Hospital starting from 2008.1.1 to 2008.12.31. Based on the dataset, we first apply the sampling techniques and dimension reduction method to preprocess the testing data. Then, we construct various kinds of classifiers (including basic classifiers, ensemble methods, and cost-sensitive methods) to predict the risk. The cost-sensitive method with random forest classifier is able to achieve recall (or sensitivity) as 100 %. At the recall of 100 %, the precision (positive predictive value, PPV), and specificity of cost-sensitive method with random forest classifier was 2.9 % and 14.87 %, respectively. In our study, we build a breast cancer risk assessment model by using the data mining techniques. Our model has the potential to be served as an assisting tool in the breast cancer screening.
KeywordsBreast cancer Cost-sensitive learning Sampling
Financial support for this study was provided in part by a grant from the National Science Council, Taiwan, under Contract No. NSC-102-2218-E-030-002. The funding agreement ensured the authors’ independence in designing the study, interpreting the data, writing, and publishing the report.
- 6.Lord, S.J., Lei, W., Craft, P., Cawson, J.N., Morris, I., Walleser, S., et al., A systematic review of the effectiveness of magnetic resonance imaging (MRI) as an addition to mammography and ultrasound in screening young women at high risk of breast cancer. Eur. J. Cancer 43 (13):1905–1917, 2007. Available from: http://www.sciencedirect.com/science/article/pii/S0959804907004844.CrossRefGoogle Scholar
- 7.Breast Cancer Screening (PDQ), Breast Cancer Screening Modalities Beyond Mammography (Health Professional Version) [homepage on the Internet]. National Cancer Institute; c2014 [updated 2014 Oct. 3; cited 2014 Oct. 6]. Available from: http://www.cancer.gov/cancertopics/pdq/screening/breast/healthprofessional/page9
- 10.Elkan, C.: The Foundations of cost-sensitive learning. In: Proceedings of the 17th International Joint Conference on Artificial Intelligence - Volume 2. IJCAI’01. Available from: http://dl.acm.org/citation.cfm?id=1642194.1642224, pp. 973–978. Morgan Kaufmann Publishers Inc., San Francisco, CA (2001)
- 11.Seiffert, C., Khoshgoftaar, T.M., van Hulse, J., Napolitano A.: A Comparative Study of Data Sampling and Cost Sensitive Learning. In: Proceedings of the 2008 IEEE International Conference on Data Mining Workshops, pp. 46–52 (2008)Google Scholar