On Stability of Ensemble Gene Selection

  • Nicoletta Dessì
  • Barbara PesEmail author
  • Marta Angioni
Conference paper
Part of the Lecture Notes in Computer Science book series (LNCS, volume 9375)


When the feature selection process aims at discovering useful knowledge from data, not just producing an accurate classifier, the degree of stability of selected features is a very crucial issue. In the last years, the ensemble paradigm has been proposed as a primary avenue for enhancing the stability of feature selection, especially in high-dimensional/small sample size domains, such as biomedicine. However, the potential and the implications of the ensemble approach have been investigated only partially, and the indications provided by recent literature are not exhaustive yet. To give a contribution in this direction, we present an empirical analysis that evaluates the effects of an ensemble strategy in the context of gene selection from high-dimensional micro-array data. Our results show that the ensemble paradigm is not always and necessarily beneficial in itself, while it can be very useful when using selection algorithms that are intrinsically less stable.


Feature selection stability Ensemble paradigm Gene selection 



This research was supported by Sardinia Regional Government (project CRP‐17615, DENIS: Dataspaces Enhancing the Next Internet in Sardinia).


  1. 1.
    Saeys, Y., Inza, I., Larranaga, P.: A review of feature selection techniques in bioinformatics. Bioinformatics 23(19), 2507–2517 (2007)CrossRefGoogle Scholar
  2. 2.
    Kalousis, A., Prados, J., Hilario, M.: Stability of feature selection algorithms: a study on high-dimensional spaces. Knowl. Inf. Syst. 12(1), 95–116 (2007)CrossRefGoogle Scholar
  3. 3.
    Zengyou, H., Weichuan, Y.: Stable feature selection for biomarker discovery. Comput. Biol. Chem. 34, 215–225 (2010)CrossRefGoogle Scholar
  4. 4.
    Awada, W., Khoshgoftaar, T.M., Dittman, D., Wald, R., Napolitano, A.: A review of the stability of feature selection techniques for bioinformatics data. In: IEEE 13th International Conference on Information Reuse and Integration, pp. 356–363. IEEE (2012)Google Scholar
  5. 5.
    Abeel, T., Helleputte, T., Van de Peer, Y., Dupont, P., Saeys, Y.: Robust biomarker identification for cancer diagnosis with ensemble feature selection methods. Bioinformatics 26(3), 392–398 (2010)CrossRefGoogle Scholar
  6. 6.
    Saeys, Y., Abeel, T., Van de Peer, Y.: Robust feature selection using ensemble feature selection techniques. In: Daelemans, W., Goethals, B., Morik, K. (eds.) ECML PKDD 2008, Part II. LNCS (LNAI), vol. 5212, pp. 313–325. Springer, Heidelberg (2008)CrossRefGoogle Scholar
  7. 7.
    Kuncheva, L.I., Smith, C.J., Syed, Y., Phillips, C.O., Lewis, K.E.: Evaluation of feature ranking ensembles for high-dimensional biomedical data: a case study. In: IEEE 12th International Conference on Data Mining Workshops, pp. 49–56. IEEE (2012)Google Scholar
  8. 8.
    Haury, A.C., Gestraud, P., Vert, J.P.: The influence of feature selection methods on accuracy, stability and interpretability of molecular signatures. PLoS ONE 6(12), e28210 (2011)CrossRefGoogle Scholar
  9. 9.
    Dessì, N., Pes, B.: Stability in biomarker discovery: does ensemble feature selection really help? In: Ali, M., Kwon, Y.S., Lee, C.-H., Kim, J., Kim, Y. (eds.) IEA/AIE 2015. LNCS, vol. 9101, pp. 191–200. Springer, Heidelberg (2015)Google Scholar
  10. 10.
    Shipp, M.A., Ross, K.N., Tamayo, P., Weng, A.P., et al.: Diffuse large B-cell lymphoma outcome prediction by gene-expression profiling and supervised machine learning. Nat. Med. 8(1), 68–74 (2002)CrossRefGoogle Scholar
  11. 11.
    Dietterich, T.G.: Ensemble methods in machine learning. In: Kittler, J., Roli, F. (eds.) MCS 2000. LNCS, vol. 1857, pp. 1–15. Springer, Heidelberg (2000)CrossRefGoogle Scholar
  12. 12.
    Dessì, N., Pascariello, E., Pes, B.: A comparative analysis of biomarker selection techniques. BioMed Res. Int. 2013, Article ID 387673 (2013)Google Scholar
  13. 13.
    Wald, R., Khoshgoftaar, T.M., Dittman, D., Awada, W., Napolitano, A.: An extensive comparison of feature ranking aggregation techniques in bioinformatics. In: IEEE 13th International Conference on Information Reuse and Integration, pp. 377–384. IEEE (2012)Google Scholar
  14. 14.
    Kuncheva, L.I.: A stability index for feature selection. In: 25th IASTED International Multi-Conference: Artificial Intelligence and Applications, pp. 390–395. ACTA Press Anaheim, CA, USA (2007)Google Scholar
  15. 15.
    Dessì, N., Pes, B.: Similarity of feature selection methods: An empirical study across data intensive classification tasks. Expert Syst. Appl. 42(10), 4632–4642 (2015)CrossRefGoogle Scholar
  16. 16.

Copyright information

© Springer International Publishing Switzerland 2015

Authors and Affiliations

  1. 1.Dipartimento di Matematica e InformaticaUniversità degli Studi di CagliariCagliariItaly

Personalised recommendations