Variable Importance in Nonlinear Kernels (VINK): Classification of Digitized Histopathology
Quantitative histomorphometry is the process of modeling appearance of disease morphology on digitized histopathology images via image–based features (e.g., texture, graphs). Due to the curse of dimensionality, building classifiers with large numbers of features requires feature selection (which may require a large training set) or dimensionality reduction (DR). DR methods map the original high–dimensional features in terms of eigenvectors and eigenvalues, which limits the potential for feature transparency or interpretability. Although methods exist for variable selection and ranking on embeddings obtained via linear DR schemes (e.g., principal components analysis (PCA)), similar methods do not yet exist for nonlinear DR (NLDR) methods. In this work we present a simple yet elegant method for approximating the mapping between the data in the original feature space and the transformed data in the kernel PCA (KPCA) embedding space; this mapping provides the basis for quantification of variable importance in nonlinear kernels (VINK). We show how VINK can be implemented in conjunction with the popular Isomap and Laplacian eigenmap algorithms. VINK is evaluated in the contexts of three different problems in digital pathology: (1) predicting five year PSA failure following radical prostatectomy, (2) predicting Oncotype DX recurrence risk scores for ER+ breast cancers, and (3) distinguishing good and poor outcome p16+ oropharyngeal tumors. We demonstrate that subsets of features identified by VINK provide similar or better classification or regression performance compared to the original high dimensional feature sets.
KeywordsPrincipal Component Analysis Mean Square Error Radical Prostatectomy Kernel Matrix Kernel Principal Component Analysis
- Ham, J., et al.: A Kernel View of the Dimensionality Reduction of Manifolds. Max Planck Institute for Biological Cybernetics, Technical Report No. TR-110 (2002)Google Scholar
- Ginsburg, S., Tiwari, P., Kurhanewicz, J., Madabhushi, A.: Variable Ranking with PCA: Finding Multiparametric MR Imaging Markers for Prostate Cancer Diagnosis and Grading. In: Madabhushi, A., Dowling, J., Huisman, H., Barratt, D. (eds.) Prostate Cancer Imaging 2011. LNCS, vol. 6963, pp. 146–157. Springer, Heidelberg (2011)CrossRefGoogle Scholar
- Esbensen, K.: Multivariate Data Analysis—In Practice: An Introduction to Multivariate Data Analysis and Experimental Design. CAMO, Norway (2004)Google Scholar
- Basavanhally, A., et al.: Multi–Field–of–View Framework for Distinguishing Tumor Grade in ER+ Breast Cancer from Entire Histopathology Slides. IEEE Trans. Biomed. Eng. (Epub ahead of print) (PMID: 23392336)Google Scholar
- Ali, S., et al.: Cell Cluster Graph for Prediction of Biochemical Recurrence in Prostate Cancer Patients from Tissue Microarrays. In: Proc. SPIE Medical Imaging: Digital Pathology (2013)Google Scholar