Abstract.
Extraction of some meta-information from printed documents without carrying out optical character recognition (OCR) is considered. It can be statistically verified that important terms in technical articles are mainly printed in italic, bold, and all-capital style. A quick approach to detecting them is proposed here. This approach is based on the global shape heuristics of these styles of any font. Important words in a document are sometimes printed in larger size as well. A smart approach for the determination of font size is also presented. Detection of type styles helps in improving OCR performance, especially for reading italicized text. Another advantage to identifying word type styles and font size has been discussed in the context of extracting: (i) different logical labels; and (ii) important terms from the document. Experimental results on the performance of the approach on a large number of good quality, as well as degraded, document images are presented.
Similar content being viewed by others
Author information
Authors and Affiliations
Additional information
Received July 12, 2000 / Revised October 1, 2000
Rights and permissions
About this article
Cite this article
Chaudhuri, B., Garain, U. Extraction of type style-based meta-information from imaged documents. IJDAR 3, 138–149 (2001). https://doi.org/10.1007/PL00013557
Issue Date:
DOI: https://doi.org/10.1007/PL00013557