Advertisement

MIRACLE-GSI at ImageCLEFphoto 2008: Different Strategies for Automatic Topic Expansion

  • Julio Villena-Román
  • Sara Lana-Serrano
  • José Carlos González-Cristóbal
Part of the Lecture Notes in Computer Science book series (LNCS, volume 5706)

Abstract

This paper describes the participation of MIRACLE-GSI research consortium at the ImageCLEFphoto task of ImageCLEF 2008. For this campaign, the main purpose of our experiments was to evaluate different strategies for topic expansion in a pure textual retrieval context. Two approaches were used: methods based on linguistic information such as thesauri, and statistical methods that use term frequency. First a common baseline algorithm was used in all experiments to process the document collection. Then different expansion techniques are applied. For the semantic expansion, we used WordNet to expand topic terms with related terms. The statistical method consisted of expanding the topics using Agrawal’s apriori algorithm. Relevance-feedback techniques were also used. Last, the result list is reranked using an implementation of k-Medoids clustering algorithm with the target number of clusters set to 20. 14 fully-automatic runs were finally submitted. MAP values achieved are on the average, comparing to other groups. However, results show a significant improvement in cluster precision (6% at CR10, 12% at CR20, for runs in English) when clustering is applied, thus proving to be valuable.

Keywords

Image retrieval domain-specific vocabulary thesaurus linguistic engineering information retrieval indexing relevance feedback topic expansion ImageCLEF Photographical Retrieval task ImageCLEF CLEF 2008 

Preview

Unable to display preview. Download preview PDF.

Unable to display preview. Download preview PDF.

References

  1. 1.
    Arni, T., Clough, P., Sanderson, M., Grubinger, M.: Overview of the ImageCLEFphoto 2008 Photographic Retrieval Task. In: Peters, C., et al. (eds.) CLEF 2008. LNCS, vol. 5706, pp. 500–511. Springer, Heidelberg (2009)Google Scholar
  2. 2.
    Villena-Román, J., Lana-Serrano, S., González-Cristóbal, J.C.: MIRACLE-GSI at ImageCLEFphoto 2008: Experiments on Semantic and Statistical Topic Expansion. In: Working Notes of the 2008 CLEF Workshop, Aarhus, Denmark (2008)Google Scholar
  3. 3.
    García-Serrano, A., Benavent, X., Granados, R., Goñi, J.M.: Some results using different approaches to merge visual and text-based features in CLEF 2008 Photo Collection. In: Peters, C., et al. (eds.) CLEF 2008. LNCS, vol. 5706, pp. 568–571. Springer, Heidelberg (2009)Google Scholar
  4. 4.
    Apache Lucene project, http://lucene.apache.org (Visited 09/11/2008)
  5. 5.
    Eurowordnet: Building a Multilingual Database with Wordnets for several European Languages (March 1996), http://www.illc.uva.nl/EuroWordNet/ (Visited 09/11/2008)
  6. 6.
    Agrawal, R., Srikan, R.: Fast algorithms for mining association rules. In: Proceedings of the International Conference on Very Large Data Bases, pp. 407–419 (1994)Google Scholar
  7. 7.
    Park, H.-s., Lee, J.-s., Jun, C.-h.: A K-means-like Algorithm for K-medoids Clustering and Its Performance. In: Proceedings of the 36th CIE Conference on Computers & Industrial Engineering, Taipei, Taiwan, June 20-23, pp. 1222–1231 (2006)Google Scholar

Copyright information

© Springer-Verlag Berlin Heidelberg 2009

Authors and Affiliations

  • Julio Villena-Román
    • 1
    • 3
  • Sara Lana-Serrano
    • 2
    • 3
  • José Carlos González-Cristóbal
    • 2
    • 3
  1. 1.Universidad Carlos III de MadridSpain
  2. 2.Universidad Politécnica de MadridSpain
  3. 3.DAEDALUS - Data, Decisions and Language, S.A.Spain

Personalised recommendations