MINISTRY OF EDUCATION AND TRAINING HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY AND EDUCATION GRADUATION THESIS MAJOR: ROBOTICS AND ARTIFICIAL INTELLIGENCE DEVELOPMENT OF AN AI SYSTEM FOR DATA EXTRACTION FROM VIETNAMESE PRINTED DOCUMENTS INSTRUCTOR: BUI HA DUC STUDENT: HUYNH VINH PHUC NGUYEN XUAN PHI CHU NHAT MINH QUAN Ho Chi Minh city, July 2024 MINISTRY OF EDUCATION AND TRAINING HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY AND EDUCATION ---------------------------------- FACULTY OF MECHANICAL ENGINEERING GRADUATION THESIS DEVELOPMENT OF AN AI SYSTEM FOR DATA EXTRACTION FROM VIETNAMESE PRINTED DOCUMENTS Supervisor: BUI HA DUC, PhD Student: HUYNH VINH PHUC Student ID: 20134005 Student: NGUYEN XUAN PHI Student ID: 20134004 Student: CHU NHAT MINH QUAN Student ID: 20134021 Year of Admission: 2020-2024 Ho Chi Minh city, July 2024 HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY AND EDUCATION FACULTY OF MECHANICAL ENGINEERING ---------------------------------- DEPARTMENT OF MECHATRONICS GRADUATION THESIS DEVELOPMENT OF AN AI SYSTEM FOR DATA EXTRACTION FROM VIETNAMESE PRINTED DOCUMENTS Supervisor: BUI HA DUC, PhD Student: HUYNH VINH PHUC Student ID: 20134005 Student: NGUYEN XUAN PHI Student ID: 20134004 Student: CHU NHAT MINH QUAN Student ID: 20134021 Class: 20134 Year of Admission: 2020 - 2024 Ho Chi Minh City, July 2024 COMMITMENT • Project: DEVELOPMENT OF AN AI SYSTEM FOR DATA EXTRACTION FROM VIETNAMESE PRINTED DOCUMENTS • Lecturer: Bui Ha Duc, PhD • Student: Huynh Vinh Phuc • Student ID: 20134005 - Class: 20134 • Adress: Hiep Phu Ward, Thu Duc City, Ho Chi Minh City • Phone number: 0357642052 • Email: phuchuynhvinh.com • Student: Nguyen Xuan Phi • Student ID: 20134004 - Class: 20134 • Adress: Linh Chieu Ward, Thu Duc City, Ho Chi Minh City • Phone number: 094540064 • Email: xphi.com • Student: Chu Nhat Minh quan • Student ID: 20134021 - Class: 20134 • Adress: Hiep Phu Ward, Thu Duc City, Ho Chi Minh City • Phone number: 0938822147 • Email: quancnm.com • Graduation thesis submission date: 04/07/2024 • Commitment: “I affirm that the graduation thesis presented here is the result of my research and efforts. I have not replicated any content from published articles without appropriate citations. Should any breach be identified, I acknowledge full accountability for the consequences.” Ho Chi Minh City, July 4, 2024 i ACKNOWLEDGEMENT We would like to begin by expressing our profound gratitude, on behalf of our team, to our supervisor, Bui Ha Duc, PhD. Your unwavering commitment, expertise, and guidance have been instrumental in shaping our research and leading us to success.
Your mentorship, patience, and continuous encouragement have been vital in navigating our academic journey. Your extensive knowledge and valuable insights have helped us overcome challenges and expand our intellectual horizons. We deeply appreciate your dedication and support. Additionally, we extend our sincere thanks to Ho Chi Minh City University of Technology and Education for providing an exceptional learning environment and resources.
The university has consistently fostered growth, innovation, and academic excellence. The dedicated faculty and staff have greatly influenced our academic and personal development, and we are grateful for their commitment to nurturing future scholars and leaders. Our heartfelt thanks go to our families for their enduring love, support, and understanding. Your unwavering belief in us has been the foundation of our journey, and we are forever grateful for your sacrifices and the countless ways you have encouraged us.
Your steadfast support has empowered us to overcome obstacles and strive for excellence. Finally, we wish to express our gratitude to all those who have contributed to our growth and development, both directly and indirectly. Your encouragement, advice, and confidence in our abilities have been invaluable throughout this challenging yet rewarding journey. Sincerely, Huynh Vinh Phuc Nguyen Xuan Phi Chu Nhat Minh Quan ii ABSTRACT Document AI refers to the use of machine learning technique to automatically analyze the structure of documents and extract relevant information.
Its purpose is to streamline the processing of large volumes of documents by identifying key elements, such as tittle, images, tables, question and answer, and converting them into structured data. This enhances efficiency and accuracy in data extraction, enabling businesses to automate workflows and improve decision-making processes. Currently, in Vietnam, paper-based documents remain the predominant medium for storing and conveying information, surpassing electronic alternatives. However, these paper documents come with inherent limitations, including inaccessibility, high maintenance costs, difficulty in managing data, and the need for extensive physical storage space.
Consequently, the digitization of documents is crucial for all organizations seeking to enhance information management efficiency. With an average processing volume of up to one thousand unstructured documents per day, typically stored in formats like PDFs or scanned images, the imperative for digitization becomes even more apparent. However, common practices in Vietnam often involve manual data entry methods, which are not only time-consuming but also reliant on human labor. The objectives of this project are to research, design and develop a software compatible with various types of scan devices to capture the scanned image of a document, and apply artificial intelligence to digitalization and extract key information from those documents.
The extracted information is then stored in the database for display and management. By applying many techniques including image processing, document layout analysis, text detection and text recognition, we have successfully created a system that is capable of extracting information from various types of forms with high performance for typed texts and acceptance performance for handwritten texts. iii TABLE OF CONTENTS COMMITMENT. iii TABLE OF CONTENTS.
iv LIST OF TABLES .vii LIST OF FIGURES. viii LIST OF ACRONYMS. Scientific and Practical Significances. Scope of The Thesis.
Scientific Research Methods. Structure of The Report. Introduction to Document Imaging Method. Introduction to Document AI.
Introduction to Document Layout Analysis. Document Layout Types. Document Layout Analysis. Introduction to Optical Character Recognition.
Scale Invariant Feature Transform. Metrics Evaluation in Document AI. Document Layout Analysis Metrics. Optical Character Recognition Metrics.
DOCUMENT IMAGE ACQUISITION. Application Programming Interface. Create Connection to Desired Scanner. DEVELOPMENT OF DATA GENERATION.
Glyph Template Creation. Conversion of Template Images to Character Images. Conversion of Character Images to SVG Format Images. Conversion of SVG Format Character Images to Handwritten Fonts.
Synthetic Data Generation. Advantages and Disadvantages of Synthetic Data. Real Data Collection. DEVELOPMENT OF TEMPLATE CREATION.
Template Creation Workflow. Optical Character Recognition. Document Layout Analysis. Data Storage Development.
Template Creation Sequence Diagram. Template Creation GUI. DEVELOPMENT OF INFORMATION EXTRACTION. Information Extraction Workflow.
Information Extraction Sequence Diagram. Information Extraction GUI. EXPERIMENTAL RESULTS AND DISCUSSION. Document Layout Analysis Model.
Text Detection Model. Text Recognition Model. 83 CONCLUSION AND FUTURE WORKS. I vi LIST OF TABLES Table 3.1: Baseline hardware requirements .2: Comparison of document layout analysis models .3: Dataset description for training text recognition model .4: Training parameter for text recognition model .5: Comparison of text recognition models .6: Illustration of correctly recognized images .7: Illustration of misidentified images.
83 vii LIST OF FIGURES Figure 2.1: Sample for image capturing (Source: Internet).2: Some type of scanners (Source: Internet) .3: OCR is the process of transform image to text (Source: [1]) .4: Document image classification examples (Source: [1]) .5: Document layout analysis examples (Source: [1]) .6: Table detection examples (Source: [1]) .8: General document layout analysis framework (Source: [2]) .9: Document layout analysis taxonomy (Source: [2]) .10: A document is passed through a generic layout analysis model, resulting in a layout segmentation mask with the following classes: title (blue), text (red), table (green), and figure (grey) .11: An example of energy map text line segmentation (Source: [2]) .13: The architecture and pre-training objectives of LayoutLMv3 (Source: [24])31 Figure 2.14: Architecture of DiT (Source: [25]) .15: General OCR process (Source:[26]) .16: Vietnames text recognition results of VietOCR (Source: [34]) .17: TransformeOCR architecture in VietOCR (Source: [34]) .18: AttentionOCR architecture in VietOCR (Source: [34]) .1: Design of scanner interface .2 Block diagram of Application Programming Interface architecture .3 Single-User to Multi-Scanner .4: Multi-User to Multi-Scanner (Server-Based) .5: Block diagram of scanning process .6: Bit and byte order of image data .1: Glyph template with characters background page 1 (Source: [36]) .2: Area containing characters .3: QR code image is cut from the image.4: Character image in PNG format (65.5: Character image in BMP format (65.6: Character image in SVG format (65.7: A glyph after automatically designing both side bearings values (left_value, right_value) = (60, -50) .8: A sample image of synthetic data .1: The template creation system .2: Image preprocessing workflow .4: A sample of document parsing .5: A sample of document layout analysis .6: Template formatting workflow .7: A new table with the same name as the form's title is created .8: Sequence diagram of template creation system .9: Template Creation system GUI. Each box can be clicked and adjust the size, position. The text each box carries can be rewritten in the top left input field of the interface (the field is “Enter text for selected box”) .1: The Information extraction system .2: Template matching with single points (blue) and matched lines (green) .3: The information after being extracted is stored in the database with its corresponding table.4: Sequence diagram of information extraction system .5: Information extraction system GUI .1: A sample result of YOLOv8 (left) and LayoutLMv3 (right) .2: Text detection results in typed and handwritten forms. 80 ix LIST OF ACRONYMS ADF Automatic Document Feeder AI Artificial Intelligence ANN Artificial Neural Network API Application Programming Interface BMP Bitmap CER Charater Error Rate CNN Convolutional Neural Network DiT Document image Transformer DLA Document Layout Analysis FCNN Fully Convolutional Neural Network GUI Graphical User Interface IoU Intersection over Union LSTM Long Short-Term Memory mAP mean Average Precision OCR Optical Character Regconition PNG Portable Network Graphics QR Quick Response RLSA Run Length Smearing Algorithm SIFT Scale Invariant Feature Transform SVG Scalable Vector Graphics WER Word Error Rate YOLO You Only Look Once x CHAPTER 1.
Motivation Digital transformation in Vietnam is occurring at a rapid pace. The effective operation of the national data sharing and integration platform enables over 1.6 million transactions daily. The national insurance database has compared and verified information for 91 million citizens from the national population database. Additionally, the civil servant and public employee database has collected data on almost 2.1 million individuals, achieving a 95% connectivity rate with all ministries, branches, and localities.
A key component of digital transformation is digitization, which involves converting information from analog to digital formats. Examples include converting handwritten text to digital format or analog audio recordings to digital format. Digitization is more than just scanning documents; it encompasses extracting information from documents, storing these digital files in repositories or the cloud, and performing ongoing maintenance and management. Digitization allows for effective information storage in databases, automating the registration and retrieval process.
This enhances document security by mitigating the risks of data loss, theft, or damage associated with paper formats, thereby preventing serious data breaches. Another benefit of digitization is cost savings. Paper documents are expensive to produce, store, and access. Digitizing these files reduces costs and minimizes paper waste, making it environmentally friendly.
Furthermore, digitization frees employees from manual data entry processes, reducing time spent on repetitive tasks and enhancing organizational productivity. Minimizing human errors improves data accuracy. Automating data entry, retrieval, classification, and sorting reduces the need for labor-intensive tasks, replacing time- consuming work with swift and precise automated algorithms. Managing vast amounts of paper documents poses significant challenges in the contemporary era.
Effective solutions necessitate support from advanced computer 1 technologies. However, computers lack the inherent ability to read, understand, and analyze paper documents autonomously to extract and digitize data. This presents a substantial difficulty, as it often requires considerable human effort and time. Consequently, researchers and companies are striving to automate this process or, at the very least, minimize human involvement.
This drive has led to the development of systems designed to digitize and distill information from documents efficiently.