Khóa luận tốt nghiệp khoa học máy tính thích ứng miền tăng tiến cho bài toán nhận diện văn bản ngoại cảnh

Khóa luận tốt nghiệp khoa học máy tính nghiên cứu phương pháp thích ứng miền tăng tiến cho bài toán nhận diện văn bản ngoại cảnh, ứng dụng hiệu quả.

Chuyên ngành

Computer Science

Tác giả

Ho Chung Duc Khanh

Người đăng

Ẩn danh

Thể loại

Bachelor thesis

2023

73
3
0

Phí lưu trữ

30 Point

Mục lục chi tiết

ACKNOWLEDGEMENTS

Abstract

1. CHƯƠNG 1: INTRODUCTION

1.1. Scene Text Recognition

1.2. Domain Adaptation

1.3. Gradual Domain Adaptation

1.4. Scope

1.5. Structure of This Thesis

2. CHƯƠNG 2: RELATED WORKS

3. CHƯƠNG 3: GRADUAL DOMAIN ADAPTATION IN SCENE TEXT RECOGNITION VIA PSEUDO-LABELING

3.1. Domain Adaptation in Scene Text Recognition via Pseudo-labeling

3.2. The Scene Text Recognition Framework

3.3. Gradual Domain Adaptation and Domain Routing

3.3.1. Unsupervised Domain Adaptation on synthetic-trained model

3.3.2. Gradual Domain Adaptation on synthetic-trained model

3.4. Comparing domain routing approach with state-of-the-art models

4. CHƯƠNG 4: EXPERIMENTS

List of Figures

List of Tables

Tóm tắt

I. Khóa Luận Tốt Nghiệp

Khóa luận tốt nghiệp này tập trung vào việc áp dụng khoa học máy tính để giải quyết vấn đề nhận diện văn bản ngoại cảnh (Scene Text Recognition - STR). Nghiên cứu đề xuất phương pháp thích ứng miền tăng tiến (Gradual Domain Adaptation - GDA) nhằm giảm thiểu khoảng cách giữa dữ liệu tổng hợp và dữ liệu thực tế. Khóa luận này được thực hiện bởi sinh viên Ho Chung Duc KhanhNguyen Thi Minh Phuong dưới sự hướng dẫn của PhD. Thanh Duc Ngo.

1.1. Mục tiêu nghiên cứu

Mục tiêu chính của khóa luận là đánh giá hiệu quả của thích ứng miền tăng tiến trong việc cải thiện độ chính xác của mô hình nhận diện văn bản ngoại cảnh. Nghiên cứu tập trung vào việc thêm các miền trung gian, thay đổi thứ tự miền, và tìm kiếm định tuyến miền phù hợp để tối ưu hóa hiệu suất mô hình.

1.2. Phương pháp tiếp cận

Phương pháp thích ứng miền tăng tiến được đề xuất nhằm giảm thiểu khoảng cách miền bằng cách huấn luyện mô hình trên nhiều miền trung gian trước khi chuyển sang miền đích. Phương pháp này được đánh giá thông qua các thí nghiệm so sánh với các mô hình hiện đại khác.

II. Nhận Diện Văn Bản Ngoại Cảnh

Nhận diện văn bản ngoại cảnh là một nhiệm vụ quan trọng trong lĩnh vực xử lý ngôn ngữ tự nhiênthị giác máy tính. Nghiên cứu này tập trung vào việc nhận diện văn bản trong các bối cảnh đa dạng như biển báo giao thông, biển số xe, và các hình ảnh kỹ thuật số.

2.1. Thách thức trong nhận diện văn bản

Một trong những thách thức lớn nhất trong nhận diện văn bản ngoại cảnh là sự đa dạng của điều kiện đầu vào, bao gồm font chữ, kích thước, màu sắc, và hướng văn bản. Ngoài ra, nhiễu nền và các yếu tố gây nhiễu khác cũng làm tăng độ phức tạp của nhiệm vụ.

2.2. Ứng dụng thực tế

Nhận diện văn bản ngoại cảnh có nhiều ứng dụng thực tế như nhận diện biển số xe, hệ thống kiểm soát truy cập, và tự động hóa quy trình xử lý tài liệu. Công nghệ này cũng hỗ trợ người khiếm thị trong việc tiếp cận thông tin.

III. Khoa Học Máy Tính và Thích Ứng Miền

Khoa học máy tính đóng vai trò quan trọng trong việc phát triển các mô hình thích ứng miền để giảm thiểu khoảng cách giữa dữ liệu tổng hợp và dữ liệu thực tế. Nghiên cứu này tập trung vào thích ứng miền tăng tiến, một phương pháp mới trong lĩnh vực này.

3.1. Thích ứng miền tăng tiến

Thích ứng miền tăng tiến là phương pháp huấn luyện mô hình trên nhiều miền trung gian để giảm thiểu khoảng cách miền trước khi chuyển sang miền đích. Phương pháp này đã được chứng minh hiệu quả trong việc cải thiện độ chính xác của mô hình nhận diện văn bản ngoại cảnh.

3.2. Định tuyến miền

Định tuyến miền là quá trình lựa chọn thứ tự các miền trung gian để tối ưu hóa hiệu suất mô hình. Nghiên cứu này đánh giá ảnh hưởng của việc thay đổi thứ tự miền và tìm kiếm định tuyến miền phù hợp.

IV. Phân Tích Văn Bản và Xử Lý Ngôn Ngữ Tự Nhiên

Phân tích văn bảnxử lý ngôn ngữ tự nhiên là các công nghệ nền tảng hỗ trợ nhận diện văn bản ngoại cảnh. Nghiên cứu này sử dụng các kỹ thuật tiên tiến để cải thiện độ chính xác và hiệu suất của mô hình.

4.1. Mô hình hóa dữ liệu

Mô hình hóa dữ liệu là quá trình biến đổi dữ liệu thô thành các đặc trưng có thể sử dụng trong huấn luyện mô hình. Nghiên cứu này sử dụng các mô hình học sâu để trích xuất đặc trưng từ hình ảnh văn bản.

4.2. Học máy trong nhận diện văn bản

Học máy đóng vai trò quan trọng trong việc phát triển các mô hình nhận diện văn bản ngoại cảnh. Nghiên cứu này sử dụng các thuật toán học sâu để cải thiện độ chính xác và khả năng thích ứng của mô hình.

V. Giá Trị và Ứng Dụng Thực Tiễn

Nghiên cứu này mang lại giá trị lớn trong việc cải thiện độ chính xác của các mô hình nhận diện văn bản ngoại cảnh. Phương pháp thích ứng miền tăng tiến có tiềm năng ứng dụng rộng rãi trong các lĩnh vực như an ninh, giao thông, và tự động hóa.

5.1. Ứng dụng trong thực tế

Phương pháp thích ứng miền tăng tiến có thể được áp dụng trong các hệ thống nhận diện biển số xe, kiểm soát truy cập, và tự động hóa quy trình xử lý tài liệu. Công nghệ này cũng hỗ trợ người khiếm thị trong việc tiếp cận thông tin.

5.2. Đóng góp khoa học

Nghiên cứu này đóng góp vào lĩnh vực khoa học máy tính bằng cách đề xuất phương pháp thích ứng miền tăng tiến mới, giúp cải thiện hiệu suất của các mô hình nhận diện văn bản ngoại cảnh. Kết quả nghiên cứu cũng cung cấp cái nhìn sâu sắc về việc lựa chọn định tuyến miền phù hợp.

Tóm tắt và mô tả trên trang này được tạo với sự hỗ trợ của AI. Nếu bạn thấy nội dung không chính xác hoặc có vấn đề, vui lòng Báo lỗi nội dung.

21/02/2025

Trích đoạn nội dung tài liệu

VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY FACULTY OF COMPUTER SCIENCE BACHELOR THESIS Gradual Domain Adaptation in Scene Text Recognition Bachelor of Computer Science (Honors degree) HO CHUNG DUC KHANH- 19520624 NGUYEN THI MINH PHUONG- 19522065 Supervised by PhD. THANH DUC NGO HO CHI MINH CITY, 2023 VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY FACULTY OF COMPUTER SCIENCE BACHELOR THESIS Gradual Domain Adaptation in Scene Text Recognition Bachelor of Computer Science (Honors degree) HO CHUNG DUC KHANH- 19520624 NGUYEN THI MINH PHUONG- 19522065 Supervised by PhD. THANH DUC NGO HO CHI MINH CITY, 2023 COMMITTEE The Thesis Defense Committee has been carefully established in accordance with the Decision 155/QD-DHCNTT, issued on 01/03/2023 by the President of the University of Information Technology. This committee is comprised of eminent individuals who possess a great deal of expertise and knowledge in the specific area of study that is relevant to the thesis defense.

To ensure that all aspects of the thesis defense are properly addressed, the following personnel have been carefully chosen to comprise the committee: ¢ Chairman: PhD. Le Dinh Duy. Nguyen Thanh Son. Le Minh Hung.

ACKNOWLEDGEMENTS We would like to express our deepest gratitude to all those who have been in- volved in the successful completion of this thesis. First and foremost, we would like to thank our supervisor, PhD. Thanh Duc Ngo, for their invaluable guidance and support throughout the course of this project. Thanh has been an incredible mentor, providing insightful feedback and suggestions, and challenging us to think outside of the box.

we are deeply grate- ful for their commitment and dedication to our project. We are also thankful to all the members of our research group for their help and encouragement. We are especially grateful to Tan and Hung, whose technical guidance and knowledge of domain adaptation techniques have been instrumen- tal in the successful completion of this thesis. We are thankful to the reviewers for their valuable feedback and suggestions, which have been extremely helpful in improving the quality of this thesis.

We would also like to thank our family and friends for their support and encour- agement throughout the entire process. Finally, we would like to express our gratitude to the computing resources pro- vided at MMLAB UIT and the Faculty of Computer Science, which enabled us to develop the algorithms and experiments presented in this thesis. All in all, this project would not have been possible without the assistance and help of the people mentioned above. Our sincere appreciation goes out to each of them.

Contents IAbstractl 1_ Introductioni 1.2 Scene Text Recogniion|.4 Gradual Domain Adaptation|. eee 11 [6 Structure of This Thesisl.1 Language-free methods|.2 Language-based methodsl. Gradual Domain Adaptation]. xa 24 BAL 5yntheicDataseBl.2 Unlabeled Real-world Datasetl.

00000 eee ee 26 3_ Gradual Domain Adaptation in Scene Text Recognition via Pseudo-labeling] 28 3.1 Domain Adaptation in Scene Text Recognition via Pseudo-labeling| 28 3.2 The Scene Text Recognition Framework|.2_ Gradual Domain Adaptation and Domain Routing].1 Unsupervised Domain Adaptation on synthetic-trained model].2 Gradual Domain Adaptation on synthetic-trained model]. 46 [£4 Comparing domain routing approach with state-of-the-art models] 52 54 See 54 ¬ 55 56 List of Figures 1.1 Scene Text Recognition task: output the text content in the image. The sample image is taken from ICDAR 2015 dataset|.2 Common challenges in Scene Text Recognition}.4 Gradual Domain Adaptation for Scene Text Recognition] .1 Hlustration of Scene Text Recognition framework].2 Transformation stage from 2] THỦ: - - - - --- ---------- 15 [2.3 Some samples of three unlabeled real-world đatasets|.1 Overview of Pseudo-labeling approach|.2 Two model combinations according to the STR Framework. (Image [rom Do eee 32 3.

Structure of Spatial Transform Networks.4 Structure of BiLSTM.5 Ilustration of Attention mechanisml.6 The Portraits dataset of students from 1905 to 2005 [17]|.1 Illustration of Experiment Resultsl|.2 Illustration of Experiment ResultslH|. 43 Samples of three unlabeled real-world datasets|.4 Illustration of Experiment ResultsM]. List of Tables 1 Experimental resultsl|.3 Compare three unlabeled real-world datasets| .4 Experiment results on different domain routings}. 51 5 Comparison between our domain routing approach and state-of- the-art methodsl.

53 Abstract Textual information is essential to virtually all aspects of our daily lives, and au- tomating the process of bringing this information onto the digital world is a key goal of research in the field of computer vision. Scene Text Recognition (STR) is a particularly important task in this domain, as it has numerous applications in areas such as automated number plate recognition for vehicles, access control systems, and much more. However, the challenge of Scene Text Recognition is that it requires large amounts of annotated data, which is very expensive and time-consuming to collect. As such, a common approach is to leverage generated synthetic data during train- ing, and test the result on real-world data.

Unfortunately, this approach can be ineffective, as there is often a large discrepancy between the synthetic data used for training and the data in the real world, referred to as a “domain gap”. Recent approaches to Scene Text Recognition have attempted to address the do- main gap by adopting Domain Adaptation techniques, which try to minimize the discrepancy between the two domains in a semi-supervised manner. However, when the gap between the two domains is too large, Domain Adaptation may not be effective. This thesis proposes and evaluates a Gradual Domain Adaptation approach, which trains the model on multiple intermediate domains in order to minimize the gap before training on the final domain.

This technique is evaluated in terms of its ef- fectiveness in reducing the domain gap, as well as the impact of adding interme- diate domains, changing domains, and finding the appropriate domain routing. As a result, the experiments are able to confirm the following: 1. Gradual Domain Adaptation improves Scene Text Recognition baseline per- formance by up to 3.12%, compared to Domain Adaptation approach. 1 List of Tables 2.

Switching domains order improves performance by 1. Performance is consistently improved when adapting with increasingly sta- ble domain routings. We observe a performance boost of up to 5.81% in our experiments. Moreover, our method was able to outperform state-of-the-art approaches by leveraging the domain routing approach.

This demonstrates the potential for this approach in scene text recognition. Keywords: Gradual Domain Adaptation, Domain Routing, Domain Adaptation, Scene Text Recognition, Unsupervised Domain Adaptation, Pseudo-labeling.1 Overview Textual information is essential to virtually all aspects of our daily lives, and au- tomating the process of bringing this information onto the digital world is a key goal of research in the field of computer vision. Scene Text Recognition (STR) is a particularly important task in this domain, as it has numerous applications in areas such as automated number plate recognition for vehicles, access control systems, and much more. However, the challenge of Scene Text Recognition is that it requires large amounts of annotated data, which is very expensive and time-consuming to collect.

As such, a common approach is to leverage generated synthetic data during train- ing, and test the result on real-world data. Unfortunately, this approach can be ineffective, as there is often a large discrepancy between the synthetic data used for training and the testing data in the real world, referred to as a “domain shift”. Recent approaches to Scene Text Recognition have attempted to address the do- main shift by adopting Domain Adaptation techniques, which try to minimize the discrepancy between the two domains in a semi-supervised manner. How- ever, when the gap between the two domains is too large, Domain Adaptation 3 Chapter 1.

Introduction may not be effective. This thesis proposes and evaluates a Gradual Domain Adaptation approach, which trains the model on multiple intermediate domains in order to minimize the gap before training on the final domain. This technique is evaluated in terms of its ef- fectiveness in reducing the domain gap, as well as the impact of adding interme- diate domains, changing domains, and finding the appropriate domain routing. The main research questions of this thesis are: 1.

What is the performance of Gradual Domain Adaptation in Scene Text Recog- nition? 2. Does intermediate domain routings affect overall performance of Scene Text Recognition model using Gradual Domain Adaptation? 3. How to choose a good domain routing when applying Gradual Domain Adaptation for Scene Text Recognition? 1.2 Scene Text Recognition Scene Text Recognition (STR) is an important task in the fields of computer vision and natural language processing, and has been heavily studied due to its many useful applications. Scene Text Recognition is capable of recognizing text com- ponents in a wide range of settings, from street signs and license plates, to news- paper headlines, advertisements, and digital images and videos.

This powerful technology enables tasks that would otherwise be impossible, such as quickly and accurately searching, translating, and verifying documents. In addition to aiding law enforcement, Scene Text Recognition can also be used to help compa- nies read customer reviews and detect text in security documents for automated Chapter 1. It is also a powerful tool for a variety of industries, from automotive to healthcare, and even retail, as it can be used to build more efficient and secure systems. Scene Text Recognition is a versatile technology that has numerous ap- plications in multiple fields and can be used to greatly improve the speed and accuracy of various tasks, making it an invaluable tool.

The input of this task is an image containing a text instance and the output of the task is the corresponding text sequence. As an illustration, figure [I-1]shows an example of the input and output of the Scene Text Recognition task. Specif- ically, the image contains the text "MOVING" which is the input of the Scene Text Recognition task. The corresponding output of the task is the text sequence "MOVING" which can be seen in the output box.

By recognizing the text compo- nents present in a scene, the Scene Text Recognition task can be used to perform a range of tasks from automating document processing to providing assistance for the visually impaired. INPUT OUTPUT FIGURE 1.1: Scene Text Recognition task: output the text content in the image. The sample image is taken from ICDAR 2015 dataset. Despite the progress made on Scene Text Recognition, the task still faces several challenges.

A key challenge is the diversity of the conditions of the input images. The text components can be in various fonts, sizes, colors, orientations, and even shapes. This means that a single recognition algorithm may not work optimally Chapter 1. Introduction across different types of images.

Additionally, the background can contain noise, complex patterns and other distracting elements, making the task more difficult. For instance, figure [1.2] demonstrates some of the common challenges of Scene Text Recognition, such as irregular fonts (Fig. noisy background (Fig. irregular text orientation (Fig.3cp, and uneven lighting/obstructed texts (Fig.

These conditions make it difficult to process the images correctly, as the system must take into account the context of the image in order to accurately recognize the text components. As a result, a robust and accurate Scene Text Recognition algorithm must be able to handle possible variations in the input images. An additional challenge lies in the limited amount of annotated data available for training. To address this issue, researchers often resort to the use of synthetic data to train their models, since it is possible to generate large amounts of data in this manner [29] [20].

While this approach can be successful in some cases, the models trained on synthetic data tend to have poor performance when applied to real-world data due to the domain gap To reduce this domain gap and increase the accuracy of the models, recent works have adopted domain adaptation techniques to bridge the difference between synthetic and real-world data 1. Domain Adaptation Domain Adaptation (DA) is a powerful machine learning technique for bridging the gap between two different datasets, particularly in cases where the training and testing data have different distributions (Fig.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ

Khóa luận tốt nghiệp với tiêu đề "Nhận Diện Văn Bản Ngoại Cảnh Với Khoa Học Máy Tính Thích Ứng Miền Tăng Tiến" tập trung vào việc áp dụng các phương pháp khoa học máy tính để nhận diện và phân tích văn bản trong các bối cảnh khác nhau. Tài liệu này không chỉ cung cấp cái nhìn sâu sắc về các kỹ thuật hiện đại trong lĩnh vực nhận diện văn bản mà còn nhấn mạnh tầm quan trọng của việc thích ứng với các miền dữ liệu đa dạng. Độc giả sẽ được trang bị kiến thức về các thuật toán và công nghệ tiên tiến, từ đó có thể áp dụng vào các dự án thực tiễn trong lĩnh vực công nghệ thông tin.

Để mở rộng thêm kiến thức, bạn có thể tham khảo tài liệu Luận văn thạc sĩ khoa học máy tính sử dụng active learning trong việc lựa chọn dữ liệu gán nhãn cho bài toán speech recognition, nơi bạn sẽ tìm hiểu về cách lựa chọn dữ liệu hiệu quả trong các bài toán nhận diện giọng nói. Ngoài ra, tài liệu Luận văn thạc sĩ khoa học máy tính nghiên cứu các phương pháp trích xuất thông tin trong ảnh tài liệu và ứng dụng sẽ giúp bạn khám phá các phương pháp trích xuất thông tin từ hình ảnh, một khía cạnh quan trọng trong nhận diện văn bản. Cuối cùng, bạn có thể tìm hiểu thêm về Luận văn thạc sĩ kỹ thuật viễn thông phân loại chủ đề bản tin online sử dụng máy học, tài liệu này sẽ cung cấp cái nhìn về việc áp dụng máy học trong phân loại văn bản trực tuyến. Những tài liệu này sẽ giúp bạn mở rộng hiểu biết và ứng dụng trong lĩnh vực khoa học máy tính.