Học Tập Giám Sát Yếu trong Trích Xuất Thông Tin: Giải Pháp Mới cho Dữ Liệu

Tài liệu nghiên cứu Weak supervision learning for information extraction, tổng hợp lý thuyết và thực hành, cung cấp kiến thức chuyên sâu về .

Người đăng

Ẩn danh

Thể loại

master's thesis

2022

70
4
0

Phí lưu trữ

30 Point

Mục lục chi tiết

Declaration of Authorship and Topic Sentences

Declaration of Authorship

Acknowledgments

Abstract

Contents

1. Introduction

1.2. Goals of the thesis

1.3. Thesis contributions

1.4. Main content and Structure of the thesis

2. Problems and Solutions

2.1. Problems with the Scrapy Crawler

2.2. Requirements for the AI Crawler

2.3. Solutions analysis

2.3.1. Page Classification

2.3.2. Extract information from web page

2.4. Overall solution

2.1. System Architecture

2.2. Website Explorer

2.3. Parser Crawler

3. Page Classification

3.1. Main Content Detection Model

3.1.1. Requirements

3.1.2. Problem analysis and Solution direction

3.2. Related Work

4. Information Extraction Model: Background

4.2. Problems analysis and Solutions direction

4.3. Framework and library

5. Information Extraction Model: Implementation and Results

5.6. Train Final Model

5.7. Result and Discussion

6. System Implementation Result

7. Conclusion

Bibliography

Glossary

Some source code were used in thesis

List of Figures

List of Tables

Tóm tắt

I. Tổng quan về Học Tập Giám Sát Yếu cho Trích Xuất Thông Tin

Học Tập Giám Sát Yếu (Weak Supervision Learning) là một phương pháp học máy đang được nghiên cứu và áp dụng rộng rãi trong lĩnh vực trích xuất thông tin. Phương pháp này cho phép xây dựng các mô hình học máy mà không cần phải có một tập dữ liệu được gán nhãn hoàn chỉnh. Thay vào đó, nó sử dụng các nguồn thông tin không chính xác hoặc không đầy đủ để tạo ra các nhãn cho dữ liệu. Điều này đặc biệt hữu ích trong các tình huống mà việc gán nhãn dữ liệu là tốn kém hoặc khó khăn.

1.1. Khái niệm về Học Tập Giám Sát Yếu

Học Tập Giám Sát Yếu là một phương pháp học máy cho phép sử dụng các nhãn không chính xác để huấn luyện mô hình. Điều này giúp giảm thiểu chi phí và thời gian trong việc chuẩn bị dữ liệu.

1.2. Tầm quan trọng của Trích Xuất Thông Tin

Trích xuất thông tin là quá trình tự động thu thập và tổ chức dữ liệu từ các nguồn khác nhau. Nó đóng vai trò quan trọng trong việc phân tích và khai thác dữ liệu lớn.

II. Vấn đề và Thách thức trong Học Tập Giám Sát Yếu

Mặc dù Học Tập Giám Sát Yếu mang lại nhiều lợi ích, nhưng vẫn tồn tại một số thách thức lớn. Một trong những vấn đề chính là độ chính xác của các nhãn được tạo ra từ các nguồn không chính xác. Điều này có thể dẫn đến việc mô hình học không đạt được hiệu suất mong muốn.

2.1. Độ chính xác của nhãn

Nhãn không chính xác có thể gây ra sự nhầm lẫn trong quá trình huấn luyện mô hình, dẫn đến kết quả không chính xác trong trích xuất thông tin.

2.2. Chi phí gán nhãn dữ liệu

Việc gán nhãn dữ liệu thủ công là một quá trình tốn kém và tốn thời gian, đặc biệt là khi khối lượng dữ liệu lớn.

III. Phương pháp Học Tập Giám Sát Yếu cho Trích Xuất Thông Tin

Để giải quyết các vấn đề liên quan đến việc gán nhãn dữ liệu, nhiều phương pháp đã được phát triển. Một trong những phương pháp hiệu quả nhất là sử dụng Học Tập Giám Sát Yếu để xây dựng tập dữ liệu huấn luyện cho mô hình trích xuất thông tin.

3.1. Sử dụng mô hình Học Tập Giám Sát Yếu

Mô hình Học Tập Giám Sát Yếu cho phép kết hợp nhiều phương pháp gán nhãn khác nhau để tạo ra một tập dữ liệu huấn luyện chất lượng cao hơn.

3.2. Tích hợp các mô hình học máy

Việc tích hợp các mô hình học máy khác nhau giúp cải thiện độ chính xác của quá trình trích xuất thông tin từ các nguồn dữ liệu khác nhau.

IV. Ứng dụng thực tiễn của Học Tập Giám Sát Yếu

Học Tập Giám Sát Yếu đã được áp dụng trong nhiều lĩnh vực khác nhau, từ thương mại điện tử đến phân tích dữ liệu lớn. Các ứng dụng này cho thấy khả năng của phương pháp trong việc cải thiện hiệu suất trích xuất thông tin.

4.1. Trích xuất thông tin từ trang web

Học Tập Giám Sát Yếu cho phép trích xuất thông tin từ các trang web mà không cần phải biết trước cấu trúc của chúng.

4.2. Phân tích dữ liệu lớn

Phương pháp này giúp các doanh nghiệp khai thác dữ liệu lớn một cách hiệu quả hơn, từ đó đưa ra các quyết định kinh doanh chính xác.

V. Kết luận và Tương lai của Học Tập Giám Sát Yếu

Học Tập Giám Sát Yếu là một lĩnh vực đang phát triển mạnh mẽ và có tiềm năng lớn trong tương lai. Với sự phát triển của công nghệ và các phương pháp học máy mới, khả năng ứng dụng của nó trong trích xuất thông tin sẽ ngày càng mở rộng.

5.1. Xu hướng phát triển

Các nghiên cứu hiện tại đang tập trung vào việc cải thiện độ chính xác và hiệu suất của các mô hình Học Tập Giám Sát Yếu.

5.2. Tác động đến ngành công nghiệp

Học Tập Giám Sát Yếu có thể thay đổi cách thức mà các doanh nghiệp thu thập và phân tích dữ liệu, từ đó tạo ra giá trị lớn hơn.

16/07/2025

Trích đoạn nội dung tài liệu

HANOI UNIVERSITY OF SCIENCE AND TECHNOLOGY Master’s Thesis in Data Science and Artificial Intelligence Weak Supervision Learning for Information Extraction NGUYEN HOANG LONG Long.vn Supervisor: Dr. Tran Viet Trung Department: Information System Ha Noi, 07/2022 Declaration of Authorship and Topic Sentences 1. Personal information Full name: Nguyen Hoang Long Phone number: 096 320 7903 Email: Long.vn Major: Data Science and Artificial Intelligence 2. Topic Weak Supervision Learning for Information Extraction 3.

Contributions • We introduce a new data collection system that uses machine learning to automatically extract information from websites without knowing the structure. • We apply the Weak Supervision learning method in building training data set for Information Extraction problem to reduce costs labeling. Declaration of Authorship I hereby declare that my thesis, titled “Weak Supervision for Information Ex- traction”, is the work of myself and my supervisor Dr. Tran Viet Trung.

All papers, sources, tables,. used in this thesis have been thoroughly cited. Ha Noi, July 2022 Supervisor Dr. Tran Viet Trung Acknowledgments I would like to thank my supervisor, Dr.

Tran Viet Trung for guiding and helping me during this research. I would also like to thank the professors at the School of Information and Communication Technology, Hanoi University of Science and Technology, especially the instructors who guided me throughout this master’s course. I would like to thank Cengroup for creating favorable conditions and work- ing environment for me to carry out this research. I am grateful for my family, friends, and colleagues who have always sup- ported me to complete my Master’s program.

ii Abstract Data collection system is a system that plays an important role in helping businesses and data laboratories to actively exploit data from websites on the Internet. With the rapid development of the Internet today, data collection systems based on information extraction rules for each website have revealed weaknesses in expanding information exploitation on more websites. To solve this problem, we introduce a new design for our data collection system called AI Crawler, which allows us to automatically detect and extract informa- tion from multiple websites without defining the structure in advance. In this study, we also apply weak supervision in data labeling for the Information Extraction problem, thereby significantly reducing the cost of data labeling.

Keywords: Information Extraction, Weak Supervise Learning. Author Nguyen Hoang Long iii Contents List of Figures List of Tables 1 Introduction 1 1.2 Goals of the thesis .4 Main content and Structure of the thesis. 2 2 Problems and Solutions 3 2.1 Problems with the Scrapy Crawler .2 Requirements for the AI Crawler .2 Extract information from web page .1 Main Content Detection Model .2 Problem analysis and Solution direction .5 Results and Discussion .2 Content Classification Model .2 Brief about Fasttext .4 Results and Discussion .3 URL Classification Model .3 Result and Discussion. 25 4 Information Extraction Model: Background 26 4.2 Problems analysis and Solutions direction .4 Framework and library.

32 5 Information Extraction Model: Implementation and Results 36 5.6 Train Final Model .2 Result and Discussion. 42 6 System Implementation Result 43 7 Conclusion 45 Bibliography 46 Glossary 48 A Some source code were used in thesis 51 List of Figures 2.1 Scrapy Crawler Architecture .2 A website has many pages, hence many categories .3 AI Crawler Architecture .4 AI Crawler Flow .5 Website Explorer survey new Website .6 Website Explorer train URL Classification Model .7 Parser Crawler flow .3 VIPS Semantic Structure [4] .4 Web2text pipeline [13] .5 Collapsed DOM procedure example [13] .1 Weak Supervision Flow [16] .2 Overview of the Snorkel system [10] .3 Overview of Fonduer [15] .4 Component and Pipeline Process in Fonduer .5 Parsing Component in Fonduer [15] .6 Fonduer Data Model [15]. 44 List of Tables 3.2 Data use for train Main Content Detection Model .3 Result of the Dragnet Model .4 Result of the Web2text Model .5 Data for Content Classification Model .6 Result of Content Classification Model .7 Result of URL Classification Model .1 Analyst Label Function Address .2 Result of Information Extraction Model.1 Problem overview As a leading real estate service business company, Cengroup follows the real estate market closely, which helps us decide on proper business strategies for any situation that arises. Working in the data technology department, our responsibility is to make sure the real estate market information is up-to-date.

To do so, we developed a Scrapy Crawler, crawling data from several large Websites. However, when it comes to scaling it up to the increasing number of Websites, we can no longer program to extract data from every Website as before. Therefore, we upgraded the Scrapy Crawler system to a smarter version called the AI Crawler. In this update, we integrated several machine learning models, allowing the crawler to automatically classify web pages and extract information.

In the process of building machine learning models for the AI Crawler System, especially the Information Extraction Model, we ran into a training data hurdle: The data is unlabeled, while manual labeling is very expensive. We have reviewed and considered several methods for reducing labeling costs, such as Active Learning, Transform Learning, and Weak Supervision Learn- ing. Ultimately, we chose Weak Supervision because of its suitability for this particular problem.2 Goals of the thesis The goals of the thesis are: • Building an AI Crawler that automatically detects and extracts infor- mation from many different Websites. • Applying Weak Supervision in data labeling to effectively reduce labeling costs.3 Thesis contributions The work in this thesis proposes improvements to the existing system in information extraction.

Generally, the research approach can be applied to similar systems and implemented in other Domains.4 Main content and Structure of the thesis In this thesis, we introduce the problems of the Scrapy Crawler system that reduce system scalability with an increasing number of Websites and our proposal for improvement in Chapter 2. In Chapter 3, Chapter 4 and Chapter 5, we present the machine learning models, which we have built to improve the system, in detail. Especially in Chapter 4 and Chapter 5, we provide an in-depth description of the Weak Supervision method in building the training data set for the machine learning model. Finally, Chapter 7 gives the conclusions about the achieved results and future development directions.

2 Chapter 2 Problems and Solutions In this chapter, we present the problems associated with the Scrapy Crawler and solutions to them.1 Problems with the Scrapy Crawler The Scrapy Crawler is designed based on Scrapy [1] platform (Fig 2. Within this architecture, each Website uses a separate data extraction code called Spider, in which we develop programs to retrieve data from a page. The source code at Listing 2.1 is an example. In the parser function of WebSpider, we defined the rules to check if a page has the required data and then extract data by using css-selector [2].

There are three problems with this design: • When we needed to crawl a new Website, we had to write a new Spider to execute it. • When a Website changes its appearance, we have to reprogram the Spi- der for it. • When more information is needed, we have to update all the Spiders. These problems make the system difficult to scale to more Websites.

Fur- thermore, scaling in this approach makes maintenance more expensive.1: Scrapy Crawler Architecture 1 import scrapy 2 class WebSpider ( scrapy. Spider ) : 3 name = ’ web_spider ’ 4 start_urls = [ 5 ’ https :// example. com / ’ , 6 ] 7 8 def parse ( self , response ) : 9 // Check this page has needed data 10 if response. get () is not None : 11 yield { 12 ’ address ’: response.

real - estate ␣ >␣ div. get () , 13 ’ price ’: response. real - estate ␣ >␣ div. get () , 14 ’ acreage ’: response.

real - estate ␣ >␣ div. get () 15 } 16 17 next_pages = response. all () 18 for next_page in next_pages : 19 yield response. follow ( next_page , self .1: WebSpider example code 4 Figure 2.2: A website has many pages, hence many categories 2.2 Requirements for the AI Crawler In the AI Crawler, we set up the crawler to be able to classify Pages and extract information automatically, with the following requirements: • The AI Crawler needs to know which Pages needed to get data from since there are many different types (Categories) of Pages on a Website.

For instance, in this system the AI Crawler only needs to get data from real estate classified advertising sites (REAL_ESTATE_PAGE Category), hence, we need to program so that the crawler can distinguish between classified advertising sites and others (Fig 2. For better optimization, the crawler needs to distinguish the Page’s URL before downloading that Page. Therefore, the Page Classification has to be based on the URL of the Page. • The AI Crawler should be able to automatically extract information from the Raw HTML without the predefined rules as in Scrapy Crawler.3 Solutions analysis To fulfill the above requirements, our approach is to use machine learn- ing models for the AI Crawler.

In this section, we present the solutions we 5 have experimented with and analyze the advantages and drawbacks of these solutions.1 Page Classification For Page Classification requirements, we have experimented with building a text classifier model based on the Websites’ URLs. The results show that in the trained Websites, the model was fairly good; but when it was tested on other websites, the result was poor. Although preprocessing and normalizing were done on URLs before training, the results were not improved. To explain this, we suppose that the URLs of different websites are 1) various and 2) too short for the model to have enough features to distinguish between them.

One upside to this is that the model was doing well on the trained Web- sites, so if we use one model for each Website separately, the result could be better. But what about new Websites? We continued experimenting with Page Classification by using Title and Description. This time, the results were good even with new Websites. However, two problems remained: • Data needs to be retrieved from the Page before we can classify it.

• Some Websites do not have Title and Description. Regarding the first problem, we have an idea is to train a Page Classi- fication model for all known Websites using Title and Description. For new Websites, we will use this model to build a training dataset with URLs, then train a URL Classification Model by using this dataset for the new Website. In terms of the second problem, we found a way to get the Main Text of the Page to replace the Title and Description.

We will present this solution in more detail in Chapter 3 To summarize, the solution to the Page Classification problem is as follows: • We build a Content Classification Model based on Title, Description, or Main Text on each Page of all known Websites. • With this model, we will build a training dataset with URL for each new Website then train URL Classification Model using this dataset for them.2 Extract information from web page Regarding the second requirement, AI Crawler must extract information from the Raw HTML automatically without the predefined rules. To do this, the crawler needs to use an Information Extraction Model for HTML data. When building this model, we ran into a training data hurdle where the data was unlabeled.

To solve this data labeling problem, we used Weak Supervision, a train- ing method that allows different labeling methods to be combined to form a better-labeled dataset. We present the details for building this model in Chapter 4. When experimenting with this model for data of real estate detail Pages, we saw that in the same Page, the model extracted many Values of the same Label.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ

Tài liệu "Học Tập Giám Sát Yếu cho Trích Xuất Thông Tin" cung cấp cái nhìn sâu sắc về phương pháp học máy trong việc trích xuất thông tin từ dữ liệu không có cấu trúc. Bài viết nhấn mạnh tầm quan trọng của việc áp dụng các kỹ thuật học tập giám sát yếu, giúp cải thiện độ chính xác và hiệu quả trong việc nhận diện và phân loại thông tin. Độc giả sẽ tìm thấy những lợi ích rõ ràng từ việc áp dụng các phương pháp này, bao gồm khả năng xử lý dữ liệu lớn và tối ưu hóa quy trình trích xuất thông tin.

Để mở rộng kiến thức của bạn về chủ đề này, bạn có thể tham khảo tài liệu Luận văn thạc sĩ xây dựng hệ thống trích chọn tên riêng cho văn bản tiếng việt bằng phương pháp học thống kê. Tài liệu này sẽ cung cấp thêm thông tin về cách áp dụng các phương pháp học thống kê trong việc trích xuất tên riêng, từ đó giúp bạn có cái nhìn toàn diện hơn về lĩnh vực này.