A Robust OCR for Degraded Documents
In the last two decades, many advances have been made in the field of document image analysis and recognition. In the recent past, several methods for recognizing Latin, Chinese, Japanese, and Arabic scripts have been proposed [7–9]. Until now, most of the OCR work has concentrated on high quality images and great success has been achieved by character recognition systems. Apart from these successes, there still exist two challenging problems in the field of recognition. The first one is optical character recognition (OCR) for low-quality images. Images having luminance variations, noise, and random degradation of text are difficult to read by OCR systems. The second open problem is that of recognizing off-line cursive handwritten character recognition . Our work concentrates on the former one particularly for Devanagari script, which is the script for Hindi, Nepali, Marathi, and several other Indic languages. Together, these languages have a user base exceeding 500 million people.
KeywordsMachine Intelligence Character Recognition Document Image Optical Character Recognition High Dimensional Feature Space
Unable to display preview. Download preview PDF.
- 1.Bansal V, Sinha RMK (2001) A Devanagari OCR and a brief review of OCR research for Indian scripts. Proceedings of STRANS01Google Scholar
- 2.Chaudhari BB, Pal U (1997) An OCR system to read two Indian Languages scripts. Proc. of 4th Int. Conf. on Document Analysis and Recognition, 1011–1015Google Scholar
- 3.Atul Negi, Chakravarthy Bhagvati, Krishna B (2001) An OCR System for Telugu, ICDAR, 1110Google Scholar
- 4.Jawahar CV, Pavan Kumar MNSSK, Ravi Kiran SS (2003) A Bilingual OCR for Hindi-Telugu Documents and its Applications, ICDAR. 408–412Google Scholar
- 5.Xuewen Wang, Xiaoqing Ding, Changsong Liu (2002) Optimized Gabor Filters Based Feature Extraction for Character Recognition, Proc.16th International Conference on Pattern Recognition, 223–226Google Scholar
- 6.Qiang Huo, Yong Ge and Zhi Dan Feng, (2001) High Performance Chinese OCR Based on Gabor Features, Discriminative Feature Extraction and Model Training. Proc. IEEE International Conference on Accoustic, Speech and Signal Processing, 1517–1520Google Scholar
- 16.Kanungo T. et al. (2000) A Statistical, Nonparametric Methodology for Document Degradation Model Validation. IEEE Trans. on Pattern Analysis and Machine Intelligence 20:1209–1223Google Scholar