Abstract
In recent years, content-based citation analysis (CCA) has attracted great attention, which focuses on citation texts within full-text scientific articles to analyze the meaning of each citation. However, citation texts often lack the appropriate evidence and context from cited papers and are sometimes even inaccurate. Thus it is necessary to identify the corresponding cited text from a cited paper and examine which part of the content of the paper was cited in a citation. In this study, we proposed a novel ranking-based method to identify cited texts. This method contains two stages: similarity-based unsupervised ranking and deep learning-based supervised ranking. A novel listwise ranking model was developed with the use of 36 similarity features and 11 section position features. Firstly, top-5 sentences were selected for each citation text according to a modified Jaccard similarity metric. Then the selected sentences were ranked using the trained listwise ranking model, and top-2 sentences were selected as cited sentences. The experiments showed that the proposed method outperformed other classification-based and voting-based identification methods on the test set of the CL-SciSumm 2017.
Access this chapter
Tax calculation will be finalised at checkout
Purchases are for personal use only
Notes
- 1.
- 2.
In the fist-stage ranking, we wanted to identify more candidate sentences and thus used F1.5 as the evaluation measure which gives more weight to recall than to precision. In the second-stage ranking, we wanted to identify cited sentences exactly and thus used F1 as the evaluation measure which gives equal weight to recall and precision.
References
Ding, Y., Zhang, G., Chambers, T., Song, M., Wang, X., Zhai, C.: Content-based citation analysis: the next generation of citation analysis. J. Assoc. Inf. Sci. Technol. 65(9), 1820–1833 (2014)
Cohan, A., Goharian, N.: Scientific document summarization via citation contextualization and scientific discourse. Int. J. Digit. Libr. 19(2-3), 287–303 (2017). https://doi.org/10.1007/s00799-017-0216-8
Jaidka, K., Chandrasekaran, M.K., Rustagi, S., Kan, M.-Y.: Insights from CL-SciSumm 2016: the faceted scientific document summarization Shared Task. Int. J. Digit. Libr. 19(2), 163–171 (2017). https://doi.org/10.1007/s00799-017-0221-y
Ma, S., Xu, J., Zhang, C.: Automatic identification of cited text spans: a multi-classifier approach over imbalanced dataset. Scientometrics 116(2), 1303–1330 (2018). https://doi.org/10.1007/s11192-018-2754-2
Yeh, J.Y., Hsu, T.Y., Tsai, C.J., et al.: On identifying cited texts for citances and classifying their discourse facets by classification techniques. J. Inf. Sci. Eng. 35(1), 61–86 (2016)
Pramanick, A., Mandi, S., Dey, M., Das, D.: Employing word vectors for identifying, classifying and summarizing scientific documents. In: Proceedings of the 3rd Computational Linguistics Scientific Summarization on Shared Task (CL-SciSumm 2017), pp. 94–98. CEUR-WS.org (2017)
Jaidka, K., Chandrasekaran, K.M., Jain, D., Kan, M-Y.: The CL-SciSumm shared task 2017: results and key insights. In: Proceedings of the 3rd Computational Linguistics Scientific Summarization on Shared Task (CL-SciSumm 2017), pp. 1–15. CEUR-WS.org (2017)
Ou, S.Y., Kim, H.I.: Unsupervised citation sentence identification based on similarity measurement. In: Chowdhury, G., McLeod, J., Gillet, V., Willett, P. (eds.) iConference 2018. LNCS, vol. 10766, pp. 384–394. Springer, Cham (2018). https://doi.org/10.1007/978-3-319-78105-1_42
Cao, Z., Qin, T., Liu, T.Y., Tsai, M.F., Li, H.: Learning to rank: from pairwise approach to listwise approach. In: Proceedings of the 24th International Conference on Machine learning, pp. 129–136. ACM, New York (2007)
Srivastava, R.K., Greff, K., Schmidhuber, J.: Training very deep networks. In: Proceedings of the 29th Annual Conference on Neural Information Processing Systems 2015 (NIPS 2015), pp. 2377–2385. Neural Information Processing Systems Foundation, Inc. (NIPS), California (2015)
Felber, T., Kern, R.: Query generation strategies for CL-SciSumm 2017 shared task. In: Proceedings of the 3rd Computational Linguistics Scientific Summarization on Shared Task (CL-SciSumm 2017), pp. 67–72. CEUR-WS.org (2017)
Li, L., Zhang, Y., Mao, L., Chi, J., Chen, M., Huang, Z.: CIST@CLSciSumm-17: multiple features based citation linkage, classification and summarization. In: Proceedings of the 3rd Computational Linguistics Scientific Summarization on Shared Task (CL-SciSumm 2017), pp. 43–54. CEUR-WS.org (2017)
Lin, C.Y.: ROUGE: a package for automatic evaluation of summaries. In: Proceedings of the Workshop on Text Summarization Branches Out (Post-Conference Workshop of ACL 2004), pp. 74–81. Association for Computational Linguistics (2004)
Acknowledgement
This paper is one of the research outputs of the project supported by the State Key Program of National Social Science Foundation of China (Grant No. 17ATQ001).
Author information
Authors and Affiliations
Corresponding author
Editor information
Editors and Affiliations
Rights and permissions
Copyright information
© 2020 Springer Nature Switzerland AG
About this paper
Cite this paper
Ou, S., Kim, H. (2020). Ranking-Based Cited Text Identification with Highway Networks. In: Sundqvist, A., Berget, G., Nolin, J., Skjerdingstad, K. (eds) Sustainable Digital Communities. iConference 2020. Lecture Notes in Computer Science(), vol 12051. Springer, Cham. https://doi.org/10.1007/978-3-030-43687-2_62
Download citation
DOI: https://doi.org/10.1007/978-3-030-43687-2_62
Published:
Publisher Name: Springer, Cham
Print ISBN: 978-3-030-43686-5
Online ISBN: 978-3-030-43687-2
eBook Packages: Computer ScienceComputer Science (R0)