Nom Optical Character Recognition using Pseudo-Skeleton Feature LE HONG TRANG Faculty of Information Technology University of Engineering and Technology Vietnam National University, Hanoi Supervised by Assoc. NGUYEN NGOC BINH A thesis submitted in fulfillment of the requirements for the degree of Master of Computer Science December, 2009 ĐẠI HỌC QUỐC GIA HA NOI Nom Optical Character Recognition using Pseudo-Skeleton Feature LE HONG TRANG Faculty of Information Technology University of Engineering and Technology Vietnam National University, Hanoi Supervised by Assoc. NGUYEN NGOC BINH A thesis submitted in fulfillment of the requirements for the degree of Master of Computer Science December, 2009 ĐẠI HỌC QUỐC GIA HA NOI Table of Contents 1 Introduction 1.1 A Brief Introduction to N6m Characters .22 OCR and Nôm Characters Recognition ee mmm A“ 6 8 8. 13 The Striuctiive ot THER 6 4 6h 5 KR CRRWERRE AS HER HES 2 Related Works 2.1 Other Approaches to Nôm OCR.1 Neural Network Approach.
212 ‘TesseraciOCR Bhgite 2: g@icawnea eee eee wns mores 22 Printed Chie OCr.50 s mormon Bee wee ee HM 2. 0202 eee ee ee ee 2.2 Stricture AGaAlyai6 «<< <8 64% SSE RR ESM RAR RS REE S 22:0 P¥ojection POH 2a ia wine wes wines Hee wee wees 2.3 Problems to N6m OCRs .0 0000 2 eee eee 3 Pseudo-Skeleton Feature Extraction a: Skeleton Featars cc 4 eS CER WED E MER S HERE HS HE es $:1.Asae Trasisformationi: « as: 2 asc ee ee ee ee 3.2 ° Medial Axis Computation .2 TH TẾ HỆ «cane eee we ee Re KY BO we 3.2 Pseudo-Skeleton Feature for Nôm OCR. 6 648 685 eK HET HER HES: 3.3 Encoding the P-Skeleton Features.4 Pre-filtering using statistical properties. iv TABLE OF CONTENTS 4 Maximum Entropy Model for Nom OCR 20 Al Overview.
ee 20 42 Maximum Entropy Model << 225 ssie ewes weeawee wes 21 4.2 Elements of Maximum entropy Model. 2 eee eR REE RESK OER RES 22 4.2:3: COnstreiite «ccc cesses eee EER RRR S 23 4.3 Maximum Entropy Model for NômOCR. 24 5 Implementation and Experiments 25 SL: Syetein OVvVervieW 2c awe eeev ava eva wesw raw evra wea 25 5.2 Building the Sets of Data for N6m Characters. 25 eo POEYWBHINICHM & s ác de vớ Vy HE v m KH 3E 8A HẢU HN Tả KẾ MÔ R Y là BE ar 5.1 Test of Printed Nôm Characters.2 Test of Kiéu Story 2.002082 eee 30 6 Conclusion 32 A Source code of P-Skeleton Logical operators 33 B Source code of Pseudo-Skeleton Encoding Module 37 List of Figures tJ An example image for literary worksin Nôm.1 Components of Chinese characters.2 Examples of projection profiles.3 Survey tovelated OCRS: cs ae ea 88 HEN BEES Ge aOR ER eS 10 3.1 Some Ïlllustrations for Medial Axis.2 The computation of Medial Axis.3 Example of thinning proces.4 Example of a reduction operation that does not preserve the topology 15 3.5 Illustration of P-Skeleton.6 An example of Pseudo-Skeleton of Nom character.
17 A process of full feature of Nom character. 19 =~] w SVSteL OVERVIEW «§ bc pe oa MRTG BERETA Ae He 26 Cn — Getting the set of Nom character from UniHan database. 26 Cty tò Converting N6m character from Unicode points to images. 27 Œt Bw Testing of original NOm images.
28 ss hm Noised Immages by erasing black dots. 29 an Co Noised images by adding blackdots. 29 G72 ONS Experimental results of noised images by erasing black dots. 30 or Results of Kiéu story recognition .00 eee eee 31 on List of Tables 5.1 Statistics of recognition in common font styles 5.2 Comparision with neural network approach .1 A Brief Introduction to N6m Characters Noin characters are an obsolete writing system of the Vietnamese language.
Nom characters are among the ideograph word systems, and are based on Chinese char- acters. The first Ném characters appeared in the 13th century and were used widely to record historical and cultural documents of Vietnam. Today, N6ém characters are replaced entirely by the writing system based on Latin alphabet. Nom characters are among the kind of words in square blocks - i.
the entire character is composed in a square. They are built from some materials of Chinese characters and but are read in Vietnamese sound. Nom character became popular in the 13-15 century. Since then many documents in various fields such as literature, history, law have been written in N6m characters.
A literary work that used Noém characters is in Figure 1.1 From the second half of the 19th century when Vietnam was invaded by French colonialism, French scholars have prevented to use of classical Chinese. Gradually to the 1915 and 1918 to 1919, Chinese characters was eliminated, drag the exclusion of the N6ém character. In the early 20th century, the national language is increasingly complete, popular and replaced completely the Ném character. An overview of NOm characters shown as figure 1.2 In this figure, two Chinese characters (Han) are in the left.
The first one is “nam” in Vietnamese (“year” in English) and the second is “nam” in Vietnamese. The Nom character in the middle is composed from the two Chinese characters and one represents the N6m sound and the other for the meaning. The right word “ndm” Chapter 1. Introduction wk SA $8 Fic Vịnh Người Chửa Hoang ñ† tí % A PAT Cả nề cho nên hóa dở dang ñ% 48 ‡/ 1A HEHL Nổi lòng chàng có biết chăng chàng te AiG VM OA Duyên thiên chưa thấy nhỏ đầu dọc 2} ƒ †R 16 iÑ i1 hì Phận liễu sao đà đây nét ngang {let AFR wae TL, 2 SF FF 1 12 SM Se A HE Chữ tình một khỏi thiếp xin mang 1 E1 im TH WS fi HEE Quản bao miệng thể nhời chênh lệch 2 (Ad 7 Pied (A) Ee Không có nhưng mà có mới ngoan BAe Hồ Xuân Hương Figure 1.1: An example image for literary works in Nom Chinese (Han) Ném National Language 13th century S| - Figure 1.
OCR and Nom Characters Recognition 3 is the current Vietnamese writing.2 OCR and N6ém Characters Recognition Optical character recognition (OCR) is one of most important problems of paper- based or image-based documents digitalization. Today there are many good OCR tools for Latin characters, after a long history of research and development. Recently, many approaches and tools for recognizing logographic (ideograph) languages like Japanese and Chinese have been developed based on, for example, k-nearest neigh- bors (k-NN) (Dasarathy, 1991) and artificial neural network (ANN) (Arbib, 1995; Barber, 2003). In Vietnam, several groups have been researching and developing OCR software for the Vietnamese national language, whose characters are based on Latin char- acters.
One of the fñrst produects is VnDOCR (công nghệ thông tin, 1997 1998) , which can convert images in different formats to Microsoft Word files while preserv- ing characters’ format. Recently, an open source project named VietOCR ! that uses Tesseract OCR engine (Smith, 2007) as its back-end can also recognize very well the Vietnamese national language. However, a lot of historical documents of Vietnamese were written in Nom, a lan- guage rooted in Chinese language. Hence, Nom optical character recognition (OCR) problem is meaningful for preserving the national cultural heritage.
Up to now, re- search and development for the Nom OCR are still limited while many OCR tools for Chinese (Dan Klein, 2003) and Japanese languages have been developed with good quality. Their successes encourage us to research and develop a Nom OCR software and we hope that it will be useful for digitalization of a lot of Nom doc- wmnents in libraries, for young generations to learn and understand our traditional culture. We also aim at building the software that can run on handheld devices so that Vietnamese and foreigners can install on their mobile phone and use it when visiting historical sites. Toward these goals, we propose an approach to No6m OCR based on maximum entropy (Dan Klein, 2003) with pseudo-skeleton feature and then present some ini- tial experimental results.
We use 8488 common N6m characters as the training data and test recognition on the several paragraphs in “Kiéu Story” (the 1866 edition). Introduction The initial result is promising, and gives some us conclusions about the advantages and disadvantages of this approach so that we can improve the quality of the recog- nition for practical application. The Structure of Thesis The rest of this thesis is organized as follows. Chapter 2 introduces related works which consist of approaches to Nom OCR with two developed method are using Neural Network and Tesseract engine.
There are some works in offline Chinese OCRs which are among kind of the same character with N6ém character. Chapter 3 represents pseudo-skeleton feature extraction. In this chapter, first we deal with term of skeleton and its classical extraction method. In addition, it represent of pseudo-skeleton feature and operations to extract it from character images.
Chapter 4 represents the maximum entropy model which is applied to building Ném OCR system. The contents of this chapter focus in brief of concepts of maximum entropy model such as modeling, training data, statistics, features, constrains and principle of maximum entropy model. The important part of this content is building the maximum entropy model for Nom OCR that is represented in the last chapter. Chapter 5 shows the implementation and experimental results.
Chappter 6 is the conclusion; This chapter indicates the archieved results of thesis and it also raises some problems and works. Appendices A and B show the source code of main functions in Nom OCR such as P-Skeleton operators and feature coding function. Chapter 2 Related Works 2.1 Other Approaches to N6m OCR Several groups have been researching and reserving N6m heritage, and noteably 4232 characters have been accepted in Unicode standard!. Based on this standard source, we can save a lot of effort in building training data with different styles and fonts.
In our previous work (Pham, 2008) we have applied several common methods for Nôm character recognition such as Tesseract OCR engine and neural network, and tested recognition with Kiéu story. The results of those methods are reasonable with the accuracy at approximately 90%.1 Neural Network Approach In neural network approach, a Multi-Layer Perceptron (MLP) network is built for recognition. The network structure consists of three layers. The input layer has 32x32 signals represented by a 32x32 bit array and it is built from analysis of a Ném char- acter.
The output layer has 16 signals represented by an array of 16 bits.The hidden layers is further divided into several hidden (sub)layers with different numbers of neural nodes per each hidden layer. The hyperbolic tangent function is used as the activation function of the network as follows: f(t) = 3 ~Ä (2.1) Training process is based on supervised learning principle with a sample set ‘http://www.com/en/dict/viet.Ys) where X, is a bit array of 32x32 elements representing a Ném character image and Y, is an array of 16 bits representing the Unicode index of the character. The purpose of training process is to find an array W of weights such that out, = f(X,,W) = Y, for all learning samples. After training process is completed, the weight set is stored for later recognition.
Recognition is a simple process which conwerts input a sample X to an output sample Y based on the weight set W created by the training process. The output sample Y is recognised if it matches a stiandard sample output that has been used to train the network. Otherwise, the network cannot recognise the sample.2 Tesseract OCR Engine Tesseract OCR (Smith, 2007) is an open source optical character recognition engine developed by HP between 1984 and 1994, now it is sponsored by Google *. A feature extraction of Tesseract is described as follows.
The first, it decomposes an input image into set of subparts which are called components. The components after are analyses and stored in outline form.