VIETNAM NATIONAL UNIVERSITY, HO CHI MINH CITY HO CHI MINH UNIVERSITY OF TECHNOLOGY DEPARTMENT OF COMPUTER SCIENCE AND ENGINEERING GRADUATION THESIS DEVELOPING A PIPELINE FOR TABLE EXTRACTION IN DOCUMENT IMAGES MAJOR: COMPUTER SCIENCE THESIS COMMITTEE: COMPUTER SCIENCE 1 SUPERVISORS: Dr. Tran Tuan Anh Mr. Nguyen Nam Quan Dr. Nguyen Tien Thinh REVIEWER: Dr.
Le Thanh Sach STUDENT: Lu Anh Khoa (1852112) HO CHI MINH CITY, 02/2023 EAI HQC QIJOC GIA TP.ICM CQNG HOA XA HOI CHU NCHIA VIET NAM DQc lap - Tu do - Hanh Phic TRU,.,NG DAI HQC BACH KHOA KHOA: KH & KT M6y tinh- NHIEM VU LUAN AN TOT NGHIEP s0 N,roN: KHMT_ Chi i: Sinh vian phqi dltn ki nd), vdo tang nhd cia ban rhuyir trinh HQ VA TEN: Lu Anh Khoa MSSV:1852 ll2 NGANH: Khoa hoc mril tinh - l6P l. DAu dd luAn rin: XAy dung trQihOng nhdn d4ng bdng trong dnh tdi liQu Developing a pipeline for table extraction in the document image 2. Nhi$m vg (y6u ciu v6 nQi dung vir s5 liQu ban dAu;: Table extraction is one ofthe most critical components ofdocument imagesi we can see it almost everywhere in every report and documenl. l-his is also the topic ofmost concem today in digital transformation.
Table recognition and comprehension require many techniques and research. it is still a massive challenge tbr scientists and industrial applications. This thesis focuses on studying the problem oftable information extraction in a whole document image processing system There are three main tasks in this research: - Build a comprehensive pipeline lor tablc cxtraction in thc document image. - Develop a model lbr detecting tablc regions in the document inlagc.
- Develop a classification model to classity the detected table into a borderless table and a bordered table for extracting the table's cells. - Benchmark the models with public/private datasets. Ngiry giao nhiQm vg lu$n 6n: 2011212021 4. Nglry hoirn thirnh nhiQm vq:20/1212022 5.
Hg tOn giing vi6n hufng dAn: Phin hu6ng din: I )Trin TuinAnh IIuong dan dinh hutrng phdt tri€n. kiim tra drinh gi[ 2) Nguy6n Nam Quin - I lucrng dAn ph6t tri6n model c6ng ngh€, thu thap dt liQu 3)Nguy6n Tiiln Thinh - Huong dAn nQi dung trinh biy. phuong ph6p nghidn cuu N6i dung vd y6u cAu I-Vl N de dugc thtlng qua 86 m6n - Ngity. CHU NHIEM BQ MON GIANG VIEN HU,ONG DAN CHiNH (Ki, vd ghi rd ho rAn) (K! vtt hi rb ho l0rt , )^' 44t4 /LaD ,r4 PH4N D)NH (,LIO KIIOA.
BO MON: Nguoi duyQt (chim so bO): Don v i: - Ngd)' bdo v4: Di6m t6ng ktit:- Noi luu tri lu6n dn: _ __ - TRU.ONG DAI HQC BACH KHOA cQNG HoA xA HQl cHU NGHie vpr Nelvt KHOA KH & KT MAY TINH DQc l{p - Tr,r do - Hanh Phuc Nsdy 26 ttuing 12 ndm2022 PHIEU CHAM BAO VE LVTN lDinh cho nguii hmmg din phun biQnt l H9 vd t6n SV: Lu Anh Khoa MSSV: 1852112 Ngdnh (chuy€n ngdnh): Khoa Hqc M6y. Od tui, Developing a pipeline for table extraction in the document image (X6y dUng hQ th6ng nhan dang bang trong anh tai liQu). Hq t6n nguoi huong din/phan biQn: Trin tuAn enl 4. T6ng qurit v6 ban thuyOt minh: 55 trang: 56 chuong: tigu s6 Uang s6 Sdhinh.vE: kheo: 56 tai liQu tham PhAn m6m tinh to6n: HiQn v4t (san ph6m) 5.
Tiing qtuit vd c6c bdn v€: - Sti-ban vC: BanAl: Ban42: Kh6 kh6a: - Sii ban vC vC tay 56 ba'n vC tren m6y tinh: 6. Nhtng uu di6m chinh cua LVTN: - This Gsis presents a pipeline for table extraction in the document image. Student also develops a model for deiecting tabie regions in the document image and a classification model to classifr the detected table into a borderless table and a bordered table for extracting the table's cells. - The presented pipeline is also used in many applications in an industrial environment.
- This thesis hai researched many different methods and approaches and made appropriate assessments and analYzes. - The thesis has also completed some practical data and put it into the evaluation, the construction of diverse data is very necessary. - Students have had good experiments, evaluations, and demos 7. Nhirng thiliu s6t chinh cua LVTN: - This thisis is inclined towards application, so the direction of research development has not been made clear.
- The overview of the works has not been fully detailed, nor is the assessment comprehensive. DE ngh!: Dugc bio vQ E 86 sung th6m d6 bio v€ tr Kh6ng dugc b6o vQ tr 9. 3 cAu h6i SV phni tra loi trudc HQi dti,ng: - Make clear the future works including research/pipeline development. Drinh gi6 chung GAng chft gi6i, klui TB): Gi6i Di6m : 9.
Ho vd t6n SV: Lu Anh Khoa MSSV: 1852112 Ngdnh (chuydn ngdnh): Computer Science Z. Od al: Developing a pipeline for table extraction in documenl images 3. Ld Thinh Sfch 4. T6ng qudt v6 ban thuyCt minh: 56 trang: 31 56 chuong: 56 bang s6 liQu 56 hinh v6: 56 tdi liQu tham kh6o: PhAn mAm tinh to6Ln: HiQn vft (san phAm) 5.
T6ng qu6t vO c6c bdn v6: - 56 ban vC: Ban Al: Ban A.2: Kh6 kh6c: - 56 ban ve ve tay SO ban ve tr6n miiy tinh: 6. Nhtng uu tli6m chinh cia LVTN: o The author has a strong background in deep leaming and its applications for computer vision. o The author has proposed a pipeline for table extraction in document images and focused (implemented) on two main tasks inside: table detection and table classification. o Table detection: consists of two steps: detecting tables using YoloVT model and post-processing detected tables with morphology and connected-component analysis.
o Table classification: classifying extracted tables into two classes: bordered vs borderless using MobilenetV3 (Large) model o The proposed method can produce accurate results and be able to compared with other methods in table extraction. Nhring thii5u s6t chinh cira LVTN: o The thesis needs to be re-written: o to add a survey for related works in the research fields (table detection and classification) o to add detail explaination for the proposed method and its results (quantitative and qualitative) o to add an evaluation for the processing time 8. Dd nghi: Du-o. c bdo vQ:x 16 sung th6m di: bao vC tr Kh6ng cluoc bAo v€ tr 9.
3 cdu hdi SV phAi trd ldi truoc H6i d6ng: Gi6i 10. Drinh giri chung (b5ng chir: gi6i, khri, TB): Di6m : 9 ll0 Kf t6n (ghi 16 ho ttn) LE Thdnh SAch Ho Chi Minh University of Technology Faculty of Computer Science and Engineering DECLARATION OF AUTHORSHIP I hereby declare that this thesis was carried out by myself under the guidance and supervision of Dr.Tran Tuan Anh, Dr. Nguyen Tien Thinh, and Mr. Nguyen Nam Quan; and that the work contained and the results in it are true by its author and have not violated research ethics.
The data and figures presented in this thesis are for analysis, comments, and evaluations from various resources by my work and have been fully acknowledged in the reference part. In addition, other comments, reviews, and data used by other authors, and organizations have been acknowledged, and explicitly cited. I will take full responsibility for any fraud detected in my thesis. HO CHI MINH CITY, Dec 2022 Author Thesis, Semester 1, Academic year 2022 - 2023 Page 1 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering ACKNOWLEDGMENT I would like to acknowledge and give my warmest thanks to my advisors, Dr.
Tran Tuan Anh, Dr. Nguyen Tien Thinh, and Mr. Nguyen Nam Quan who made this work possible. They are the ones who build the first bricks of my scientific career.
Besides, I would like to also acknowledge all of the instructors of Ho Chi Minh University of Technology, who have given me motivation, encouragement, and precious knowledge during the long road of my university life. Last but not least, I would like to thank my family who always be there, support me throughout my life. Thesis, Semester 1, Academic year 2022 - 2023 Page 2 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering ABSTRACTION The ”Developing pipeline for table extraction in document images” research topic aims to develop a system to extract tabular regions from scanned/captured document images (invoices, reports, research papers .) with high accuracy and reasonable response time. This thesis will propose a pipeline consists of several steps to detect, classify and to extract data from tabular regions.
Thesis, Semester 1, Academic year 2022 - 2023 Page 3 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering Contents 1 Introduction 7 1.1 The need for table extraction .1 Convolution and cross-correlation in image processing .2 Convolution neural network (CNN) .5 Object detection - Metric .8 YOLO approach for object detection .1 CSP-ize a block .15 YOLOv7 - training techniques .1 Label assignment: Simple Optimal Transport Assignment (SimOTA) .16 YOLOv7 - Loss function .18 Image classification - Metric .1 Depthwise Separable Convolutions .2 Inverted residual block .3 Squeeze and excite (SE) .2 Convolutional Neural Networks (CNN) .3 Postprocess for table detection .5 Table structured recognition. 41 Thesis, Semester 1, Academic year 2022 - 2023 Page 4 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering 4.6 Structure of a table .7 Bordered table - just use some image processing .8 Putting everything together .1 Table detection - mAP .2 Table detection - Weighted Average F1 .3 Table classification - Model comparisons. 49 List of Figures 1 A typical CNN. 9 2 Maxpool 2x2: an example of pooling layer.
9 3 Some activation function. 10 4 Example of Dilated Convolution. 12 6 Example of data augmentation. 12 7 An example of image segmentation output.
13 8 Different IOU with Red bounding boxes are ground truths while the Green ones are pre- dictions. 16 10 Faster RCNN achitecture. 16 11 RPN in Faster-RCNN. 17 13 Detection model in Faster-RCNN.
23 18 Darknet53 vs CSPDarknet53. 24 20 Left: Sample image, Center: DropOut, Right: DropBlock. 24 22 SPP in YOLOv4. 25 23 Hard label vs Smooth label.
25 24 Cosine Annealing Learning rate .07 for all 3 cases but their IOU are different by a large margin. IOU is also used to evaluate an object detection model. So using IOU as a loss is a logical improvement. 27 28 CSP-ized a block.
30 32 CSP-OSA block. 32 35 YOLOv7 head, with implicit knowledge. 32 Thesis, Semester 1, Academic year 2022 - 2023 Page 5 Ho Chi Minh University of Technology Faculty of Computer Science and Engineering 36 An example of mosaic augmentation, 4 images are ”merged” together to create a new sample. Red bounding boxes annotate labels.
34 37 Mixup augmentation example. 34 38 Depthwise Convolution, visualized. 36 39 Left: Normal convolution. Right: Depthwise separable convolution.
37 40 Squeeze and excitation block. 37 41 CascadeTabNet architecture from [23]. Notice that on the left the table does not have enough outer bordered lines. 40 44 Examples of bordered tables.
41 45 Examples of borderless tables. 41 46 Original bordered table. 43 47 After cell indexing, each cell has the form start row, start col. end row, end col.
44 48 Excel result of figure 46. Note that tesseract cannot read the text in some cells. 44 49 Correct detection on publaynet dataset. 47 50 Correct detection on fintabnet dataset.
48 51 Wrong cases: Partial detection (top), missed tables(bottoms). Green boxes are ground truth while red boxes are predictions. 49 List of Tables 1 Data statistic for table detection. 45 2 Data statistic for table classification.
45 3 Result on test set of each dataset, table detection .