Single-Sentence Readability Prediction in Russian

Karpov, Nikolay; Baranova, Julia; Vitugin, Fedor

doi:10.1007/978-3-319-12580-0_9

Nikolay Karpov⁶,
Julia Baranova⁶ &
Fedor Vitugin⁶

Part of the book series: Communications in Computer and Information Science ((CCIS,volume 436))

Included in the following conference series:

International Conference on Analysis of Images, Social Networks and Texts

1263 Accesses
10 Citations

Abstract

In an effort to make reading more accessible, an automated readability formula can help students to retrieve appropriate material for their language level. This study attempts to discover and analyze a set of possible features that can be used for single-sentence readability prediction in Russian. We test the influence of syntactic features on predictability of structural complexity. The readability of sentences from SynTagRus corpus was marked up manually and used for evaluation.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 39.99; Price excludes VAT (USA)

Softcover Book: USD 54.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Notes

1.
http://texts.cie.ru

References

Flesch, R.: A new readability yardstick. J. Appl. Psychol. 32, 221 (1948)
Article Google Scholar
Kincaid, J.P., Fishburne, Jr., R.P., Rogers, R.L., Chissom, B.S.: Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel. DTIC Document (1975)
Google Scholar
Chall, J.S.: Readability Revisited: The New Dale-Chall Readability Formula. Brookline Books Cambridge, MA (1995)
Google Scholar
Collins-Thompson, K., Callan, J.: Predicting reading difficulty with statistical language models. J. Am. Soc. Inf. Sci. Technol. 56, 1448–1462 (2005)
Article Google Scholar
Schwarm, S.E., Ostendorf, M.: Reading level assessment using support vector machines and statistical language models. In: Proceedings of the 43rd Annual Meeting on Association for Computational Linguistics, pp. 523–530. Association for Computational Linguistics (2005)
Google Scholar
Oborneva, I.: Automatic assessment of the complexity of educational texts on the basis of statistical parameters (2006)
Google Scholar
Krioni, N., Nikin, A., Filippova, A.: Automated system for analysis of the complexity of educational texts. Manag. Soc. Econ. Syst. 11, 101–107 (2008)
Google Scholar
Verhelst, N., Van Avermaet, P., Takala, S., Figueras, N., North, B.: Common European Framework of Reference for Languages: Learning, Teaching. Assessment. Cambridge University Press, Cambridge (2009)
Google Scholar
Bocharov, V., Stepanova, M., Ostapuk, N., Bichineva, S., Granovsky, D.: Quality assurance tools in the OpenCorpora project. In: Computational Linguistics and Intelligent Technology: Proceeding of the International Conference « Dialog–2011 » , pp. 10–17 (2011)
Google Scholar
Francois, T.L.: Combining a statistical language model with logistic regression to predict the lexical and syntactic difficulty of texts for FFL. In: Proceedings of the 12th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop, pp. 19–27. Association for Computational Linguistics (2009)
Google Scholar
Nevdah, M.: Development of a method of automated evaluation of the complexity of educational texts for higher school (2008)
Google Scholar
Kent, J.T.: Information gain and a general measure of correlation. Biometrika 70, 163–173 (1983)
Article MATH MathSciNet Google Scholar
Nivre, J., Boguslavsky, I.M., Iomdin, L.L.: Parsing the SynTagRus treebank of Russian. In: Proceedings of the 22nd International Conference on Computational Linguistics, vol. 1, pp. 641–648. Association for Computational Linguistics (2008)
Google Scholar

Download references

Acknowledgment

This study comprises research findings from the «Adaptation of texts from the Russian National Corpus» for the electronic textbook «Russian language as a foreign one» carried out within The National Research University Higher School of Economics’ Academic Fund Program in 2013, grant No 13-05-0031.

Author information

Authors and Affiliations

National Research University Higher School of Economics, Nizhny Novgorod, Russia
Nikolay Karpov, Julia Baranova & Fedor Vitugin

Authors

Nikolay Karpov
View author publications
You can also search for this author in PubMed Google Scholar
Julia Baranova
View author publications
You can also search for this author in PubMed Google Scholar
Fedor Vitugin
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Nikolay Karpov .

Editor information

Editors and Affiliations

National Research University Higher School of Economics, Moscow, Russia
Dmitry I. Ignatov
Krasovsky Inst. of Math. and Mechanics, Yekaterinburg, Russia
Mikhail Yu. Khachay
Université catholique de Louvain, Moscow, Russia
Alexander Panchenko
University of Wolverhampton, Wolverhampton, United Kingdom
Natalia Konstantinova
National Research University, Moscow, Russia
Rostislav E. Yavorsky

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Karpov, N., Baranova, J., Vitugin, F. (2014). Single-Sentence Readability Prediction in Russian. In: Ignatov, D., Khachay, M., Panchenko, A., Konstantinova, N., Yavorsky, R. (eds) Analysis of Images, Social Networks and Texts. AIST 2014. Communications in Computer and Information Science, vol 436. Springer, Cham. https://doi.org/10.1007/978-3-319-12580-0_9

Download citation

DOI: https://doi.org/10.1007/978-3-319-12580-0_9
Published: 07 November 2014
Publisher Name: Springer, Cham
Print ISBN: 978-3-319-12579-4
Online ISBN: 978-3-319-12580-0
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics