Ranking-Based Cited Text Identification with Highway Networks

Ou, Shiyan; Kim, Hyonil

doi:10.1007/978-3-030-43687-2_62

Part of the book series: Lecture Notes in Computer Science ((LNISA,volume 12051))

Included in the following conference series:

International Conference on Information

2761 Accesses

Abstract

In recent years, content-based citation analysis (CCA) has attracted great attention, which focuses on citation texts within full-text scientific articles to analyze the meaning of each citation. However, citation texts often lack the appropriate evidence and context from cited papers and are sometimes even inaccurate. Thus it is necessary to identify the corresponding cited text from a cited paper and examine which part of the content of the paper was cited in a citation. In this study, we proposed a novel ranking-based method to identify cited texts. This method contains two stages: similarity-based unsupervised ranking and deep learning-based supervised ranking. A novel listwise ranking model was developed with the use of 36 similarity features and 11 section position features. Firstly, top-5 sentences were selected for each citation text according to a modified Jaccard similarity metric. Then the selected sentences were ranked using the trained listwise ranking model, and top-2 sentences were selected as cited sentences. The experiments showed that the proposed method outperformed other classification-based and voting-based identification methods on the test set of the CL-SciSumm 2017.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 84.99; Price excludes VAT (USA)

Softcover Book: USD 109.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Notes

1.
https://www.nltk.org.
2.
In the fist-stage ranking, we wanted to identify more candidate sentences and thus used F_1.5 as the evaluation measure which gives more weight to recall than to precision. In the second-stage ranking, we wanted to identify cited sentences exactly and thus used F₁ as the evaluation measure which gives equal weight to recall and precision.

References

Ding, Y., Zhang, G., Chambers, T., Song, M., Wang, X., Zhai, C.: Content-based citation analysis: the next generation of citation analysis. J. Assoc. Inf. Sci. Technol. 65(9), 1820–1833 (2014)
Article Google Scholar
Cohan, A., Goharian, N.: Scientific document summarization via citation contextualization and scientific discourse. Int. J. Digit. Libr. 19(2-3), 287–303 (2017). https://doi.org/10.1007/s00799-017-0216-8
Article Google Scholar
Jaidka, K., Chandrasekaran, M.K., Rustagi, S., Kan, M.-Y.: Insights from CL-SciSumm 2016: the faceted scientific document summarization Shared Task. Int. J. Digit. Libr. 19(2), 163–171 (2017). https://doi.org/10.1007/s00799-017-0221-y
Article Google Scholar
Ma, S., Xu, J., Zhang, C.: Automatic identification of cited text spans: a multi-classifier approach over imbalanced dataset. Scientometrics 116(2), 1303–1330 (2018). https://doi.org/10.1007/s11192-018-2754-2
Article Google Scholar
Yeh, J.Y., Hsu, T.Y., Tsai, C.J., et al.: On identifying cited texts for citances and classifying their discourse facets by classification techniques. J. Inf. Sci. Eng. 35(1), 61–86 (2016)
Google Scholar
Pramanick, A., Mandi, S., Dey, M., Das, D.: Employing word vectors for identifying, classifying and summarizing scientific documents. In: Proceedings of the 3rd Computational Linguistics Scientific Summarization on Shared Task (CL-SciSumm 2017), pp. 94–98. CEUR-WS.org (2017)
Google Scholar
Jaidka, K., Chandrasekaran, K.M., Jain, D., Kan, M-Y.: The CL-SciSumm shared task 2017: results and key insights. In: Proceedings of the 3rd Computational Linguistics Scientific Summarization on Shared Task (CL-SciSumm 2017), pp. 1–15. CEUR-WS.org (2017)
Google Scholar
Ou, S.Y., Kim, H.I.: Unsupervised citation sentence identification based on similarity measurement. In: Chowdhury, G., McLeod, J., Gillet, V., Willett, P. (eds.) iConference 2018. LNCS, vol. 10766, pp. 384–394. Springer, Cham (2018). https://doi.org/10.1007/978-3-319-78105-1_42
Chapter Google Scholar
Cao, Z., Qin, T., Liu, T.Y., Tsai, M.F., Li, H.: Learning to rank: from pairwise approach to listwise approach. In: Proceedings of the 24th International Conference on Machine learning, pp. 129–136. ACM, New York (2007)
Google Scholar
Srivastava, R.K., Greff, K., Schmidhuber, J.: Training very deep networks. In: Proceedings of the 29th Annual Conference on Neural Information Processing Systems 2015 (NIPS 2015), pp. 2377–2385. Neural Information Processing Systems Foundation, Inc. (NIPS), California (2015)
Google Scholar
Felber, T., Kern, R.: Query generation strategies for CL-SciSumm 2017 shared task. In: Proceedings of the 3rd Computational Linguistics Scientific Summarization on Shared Task (CL-SciSumm 2017), pp. 67–72. CEUR-WS.org (2017)
Google Scholar
Li, L., Zhang, Y., Mao, L., Chi, J., Chen, M., Huang, Z.: CIST@CLSciSumm-17: multiple features based citation linkage, classification and summarization. In: Proceedings of the 3rd Computational Linguistics Scientific Summarization on Shared Task (CL-SciSumm 2017), pp. 43–54. CEUR-WS.org (2017)
Google Scholar
Lin, C.Y.: ROUGE: a package for automatic evaluation of summaries. In: Proceedings of the Workshop on Text Summarization Branches Out (Post-Conference Workshop of ACL 2004), pp. 74–81. Association for Computational Linguistics (2004)
Google Scholar

Download references

Acknowledgement

This paper is one of the research outputs of the project supported by the State Key Program of National Social Science Foundation of China (Grant No. 17ATQ001).

Author information

Authors and Affiliations

School of Information Management, Nanjing University, Nanjing, China
Shiyan Ou & Hyonil Kim

Authors

Shiyan Ou
View author publications
You can also search for this author in PubMed Google Scholar
Hyonil Kim
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Shiyan Ou .

Editor information

Editors and Affiliations

OsloMet – Oslo Metropolitan University, Oslo, Norway
Anneli Sundqvist
OsloMet – Oslo Metropolitan University, Oslo, Norway
Gerd Berget
University of Boras, Boras, Sweden
Jan Nolin
OsloMet – Oslo Metropolitan University, Oslo, Norway
Kjell Ivar Skjerdingstad

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Ou, S., Kim, H. (2020). Ranking-Based Cited Text Identification with Highway Networks. In: Sundqvist, A., Berget, G., Nolin, J., Skjerdingstad, K. (eds) Sustainable Digital Communities. iConference 2020. Lecture Notes in Computer Science(), vol 12051. Springer, Cham. https://doi.org/10.1007/978-3-030-43687-2_62

Download citation

DOI: https://doi.org/10.1007/978-3-030-43687-2_62
Published: 19 March 2020
Publisher Name: Springer, Cham
Print ISBN: 978-3-030-43686-5
Online ISBN: 978-3-030-43687-2
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics