VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY ADVANCED PROGRAM IN INFORMATION SYSTEMS LE DUC TIN — 19522348 NGUYEN VĂN QUOC VIET -— 19522518 NATURAL LANGUAGE QUESTION-ANSWERING SYSTEM ABOUT TOURISM LOCATION BACHELOR OF ENGINEERING IN INFORMATION SYSTEMS THESIS ADVISOR Prof. DO PHUC HO CHI MINH CITY, 2023 ASSESSMENT COMMITTEE The Assessment Committee is established under the Decision. by Rector of the University of Information Technology. KLTN-8 DAI HOC QUOC GIA TP.HCM CONG HOA XA HỘI CHỦ NGHĨA VIỆT NAM TRƯỜNG ĐẠI ,HỌC Độc lập— Tự do- Hạnh phúc CÔNG NGHỆ THÔNG TIN BIEN BẢN HOI DONG KHÓA LUẬN TOT NGHIỆP Đợt: 1 - Năm học: 2023-2024 Khoa Hệ thống thông tin I.
Thời gian - Địa điểm: 1. Thời gian: bắt đầu lúc. ngày 22 tháng 01 năm 2024 2. Dia điểm: phòng F3.
Thành phan hội đồng: 1. Nguyễn Đình Thuân. Ngô Đức Thành. Nguyễn Thanh Bình.:: Thông tin (các) đề tài: CÁN BỘ HƯỚNG DẪN PHẢN BIỆN 19522348 | Lé Dire Tin IV.
Hội đồng làm việc a. Thư ký đọc quyết định thành lập hội đồng. Chủ tịch hội đồng điều khiển buổi bảo vệ. Sinh viên trình bay dé tài.
CBPB (hoặc thư ký) đọc nhận xét. Sinh viên trả lời các câu hỏi của CBPB và hội đồng. CBHD (hoặc thư ky) đọc nhận xét. Hội dong trao đôi và cham điểm.
Y kiến trao đôi của hội đồng we Da Stahl Adu. rer mem are ide khz/sh. j i ie _ den ith. NÀ N Lk ue en re es er oe no eo ee f2] h.
g Ls6 p a Mil ag.35)/ u i aco a e Le Ade.ẻ h3 ie ¬ ao e t ‘lay. Tra „4 ude Aplin dul, Shin, a oth. hi th ấu. a e a Th d n ad we t dlê n a aida ttDaihen.
Kết luận Hội đồng thố ng nhấ t két quả khó a luận tốt nghiệp của các sin h viê n nh ư sau : TT Họtên sin h viê n Đi ểm | Di em | Di ém | Đi ểm | Đi ểm | Đi êm CBHD | CBPB | CTHĐ |TKHD | UVHD tông kết RẺ tee HỆ Hội đồng kết thúc lúc. Thư ký Chủ tịch hội đồng (ký và ghi rõ họ tên) (ký và ghi rõ họ tên) TS. Ngô Dire Thanh Đình Thuân PGS. Xe _ Lưu ý: Số thành viên Hội đồng là 03.
Công thức tính điểm tông kết: e - Trường hợp CBPB khong có trong Hội đồng Điểm KLTN = (CTHĐ+TKHĐ+UVHĐ+CBHD*2+CBPB*2)/7 « _ Trường hợp CBPB là ủy viên Hội đồng Điểm KLTN = (CTHĐ+TKHĐ+CBHD*2+CBPB*2)/6 ACKNOWLEDGMENTS The thesis was completed at the University of Information Technology, Vietnam National University, Ho Chi Minh City. During the process of writing my graduation thesis, I received a lot of help from teachers at the school and department to complete the thesis. Firstly, we would like to send our sincere and special thanks to Prof. Đỗ Phúc for enthusiastically guiding and imparting knowledge and experience to us throughout the process of writing this graduation thesis.
I would like to thank the teachers of the Department of Information Systems, University of Information Technology, who have imparted valuable knowledge to me during the past 4 years of studying at the school. Finally, we would like to thank our family, friends, and students of CTTT2019.2 class for always encouraging and helping us during the thesis writing process. At the same time, I would like to thank the respondents who enthusiastically participated in answering the survey questions to help me complete this graduation thesis. However, the condition of personal capacity is still limited, the research topic certainly cannot avoid shortcomings.
We hope to receive comments from teachers in the Department of Information Systems. Once again sincerely thank you! Ho Chi Minh City, September 1, 2023 Group of authors Nguyen Van Quoc Viet — Le Duc Tin TABLE OF CONTENTS Lie ABSTRACT 1 CHAPTER 1 INTRODUCTIONN. The problems and its significance. S1 TH TH TH HH 6 1.
Structure of the theS1S.- -- s11 Tnhh ng nh 6 CHAPTER2 BACKGROUND AND RELATED WORKS. 4g cố CV ca) mm n. cece cece ceeesceseceeecseecsecsesseeeceeessessesseseeeseeeseeeeseeaaeegs 12 2. 5 s1 9v nh ng ng nghiện 18 2.-- s9 gg nHHngnhệt 23 "hy Y co naỀ.
Position wise Feed-Forward NGfWOTK. Evaluation Methods for Word Embedding Techniques. 36 CHAPTER 3 SYSTEM DESIGN. Training Model ATrChIt€C(UTC.
System ATChI(€CfUTG.- Ăn HH HH ng Hư,40 3. Data DTOC€SSINE. SG TH TH TH TH HH 41 3. TH TH HH ng ng Hết 41 3.- Ăn HH HH như42 ch).
Parallel Setence Dataset. SG SH HHHHHTHnHnHngnHnệt 44 3. Pre-traIned mO€ÌS.- óc 11 991193119111 vn ng nến 44 “hs. SH TH TH HH HH Hi, 46 CHAPTER4 SYSTEM IMPLEMENTA TION.
Training model DaTaIT€fTS.- s5 1v vn nhe, 54 CHAPTER 5 CONCLUSIONS AND FUTURE WORKS. 2G G1 TH TH HH Hệ 57 5. HT HT gi HH nh it 58 REFERENCES.- G0 cọ HH G000 0090059 APPENDIX ASOME EXAMPLE TYPES OF QUESTIONS TO TEST THE CHATBOT00 15155. 61 APPENDIXB GUIDE TO SET UP OUR MODEL BASED ON PRE- TRAINED MODEL PHO B,ERT,.-- << << <9 56505 038956956 056656 05666 63 APPENDIX C SOME IMAGE ABOUT OUR WEBSITE AND MOBILE .66 APPENDIX D HIGHLIGHTS IN THIS THESIS .--<=<<=<<<< 71 APPENDIX E SEVERAL NOVELS APPROACHES .ccscsscsssssssscssossenee 73 LIST OF FIGURES Lz Figure 2-1: The transformation of vectors using SOffIax.
--- 5< +<<c+<x+x+ 11 Figure 2-2: Comparison between cross-entropy function and squared distance.12 Figure 2-3: Work embedding.-- ---- - -- + + + 13311119111 91 1 91119 1v ng ng 14 Figure 2-4: Data preprocessing 1n BIEIRÏT.-- 5 6+ 3k1 1991 9 1v re, 16 Figure 2-5: Dynamic masKITE.- 6 6 + 3 931911911 1E 1 vn ngưng ng gưệp 18 Figure 2-6: VnCoreNLP tasK.- - Ăn HH TH TH HH HH HH, 20 Figure 2-7: VnCoreNLP token1Zaf(IOH.-- cece 1E ngư 20 Figure 2-8: Transformer €TCO(T. --- 5 + E21 191899118391 E911 9 1 ve 21 Figure 2-9: The mechanism of operation of the Transformer Encode. 22 Figure 2-10: Deep Averaging ñ€fWOTK.- --- -- + kh HH ng gi tưệt 23 Figure 2-11: SBERT ModelL.-- --- 6 <Ex+ SE k 39 111g HH HH ngư 24 Figure 2-12: Classification Objective FUNCTION.- «+ ss + ssssvesseese 26 Figure 2-13: Regression Objective FUncfIOH.-- ---- «+ s + vs sssseeeseeee 26 Figure 2-14: Triplet Objective FUNCTION.-- --- 3+ *+*‡+*EE+eeeeeeeeeerereers 27 Figure 2-15: The architecture of the Transformer model.- --«-s«+<s+ 29 Figure 2-16: Self Atf€nfIOH.- ch HH HH 30 Figure 2-17: Illustration Wk, Wq, WV. LH HH TH ng 3l Figure 2-18: Attention weight computation ((Ï).- - «+ « «+ x+ssexsseeeeseeseeers 31 Figure 2-19: Attention weight computation (2).-- - -- + +s + +++++eeseeeeeeeeseeers 32 Figure 2-20: The process of computing the self-attention mechanIsm.
33 Figure 2-21: Multi-head Attention. cccccccscessecesceeeseeeeeeseneceeeeeaeeseaecssneeeseeeaes 35 Figure 2-22: Multi-head Atf€nfIOH.- 5 5c kg TH ng HH tiệt 36 Figure 2-23: STS Benchmark (Ìafa.- (5 + E1 E991 E391 E311 911 9 re 37 Figure 3-1: Training student model mimics teacher model.- - --‹---««+ 39 Figure 3-2: System ATChIf€CfUTC.- - -- c3 11331118311 18911 1191111 11 9 1 191 1g ng 40 Figure 3-3: Question-Answer training data. sành rikt 41 Figure 3-4: Training data for similar meaning senf€nCeS .- - -- «+ +-s«++<<ss+ 41 Figure 3-5: Average Pooling ExampÏe.- --- << + 3k k* nh ng trên46 Figure 4-1: Training time CaFT. - -- c5 + 3 31 1911 91 1v HH HH ng ng rưệt 49 Figure 4-2: Evaluation Models CaTt.- -- <5 1311191113 11 9 1 9 re 50 Figure 4-3: Pearson correlation coefficient CHaTI.-- 5555 s + + ++eexsseseeers 51 Figure 4-4: Application arChIf€CfUT€.
56 Figure C-1: Home PBe. - ce -- <1 HH TH HH 66 Figure C-2: Login P 4G.- -- c1 TH TH HH nh 67 Figure C-3: Sign Up Page.ccccescsssccssseceseesseceseecesceceeeeeseesceceeeceeeseaeceeaeeeeeesaes 67 Figure C-4: Home PB€. n9 91h TH TH HH HH HH, 68 Figure C-5: Detail Dag. LG HH ng TH TH 69 Figure C-6: Wikipedia DaB€.
- nh HH TH ng HH tiệt 70 Figure D-1: Question and answer with chatTP T. --- «<< + ssseeeskrrse 71 Figure D-2: Question and answer With OUT app.-- --- 5 + S<*+ksseeeeseesreers 71 LIST OF TABLES Lie Table 3-1: Example removing non-ASCII charaCt€TS.-- 55 eee s<<<+£<scese 42 Table 3-2: Example removing SfODWOTS. cv ng ng 42 Table 3-3: Example capitalize WOTS.- -- c1 1 1v HH ng ng ng rg 43 Table 3-4: Parallel Sentence Dafaset. cv HH, 43 Table 3-5: T'Ok€ñ1ZAfIOI.
- G1 01019 HH 44 Table 005cii) 0 11. 51 LIST OF ABBREVIATIONS Abbreviation Full form 1 NLP Natural Language Processing 2 AI Artificial Intelligence 3 QA Question and Answering 4 RNN Recurrent Neural Network 5 BERT Bidirectional Encoder Representations from Transformers 6 NLI Natural Language Inference 7 STS Semantic Textual Similarity 8 SBERT Sentences BERT 9 NSP Next Sentence Prediction 10 BPE Byte Pair Encoding ABSTRACT In the comtemporary era, the ubiquity of the Internet has revolutionized the way we access and consume information, with a vast majority of it being readily available for free. However, in this diversity of resources, the aspect of ensuring the accuracy of information has become an increasingly challenging endeavor. In the pursuit of knowledge, it is imperative that we navigate through the vast sea of data by far- sighted eyes, and a clear mind which can distinguish between reliable and incredulous.
This thesis embarks on a significant exploration of the question-answer problem within the context of the Vietnamese language and toursim location domain, employing deep learning techniques. The specific focus of our research centers on the rich landscapes and tourist destinations that characterize the vibrant nation of Vietnam. By delving into this specific domain, we aim to find fact resources and extract meaningful information and provide accurate responses to user queries. Vietnam, renowned for its picturesque landscapes and culturally significant landmarks, serves as an attractive context and topic for us to exploit and investigate.
The country’s diverse topography, ranging from the serene beauty of Ha Long Bay to the historical significance of Hue’s acient citadel, presents a huge of challenges and opportunities in the development of a robust question-answer system. To bridge the gap between theoretical exploration and practical application, our research endeavors extend beyond the confines of traditional academic discourse. We are committed to bringing our findings to user by developing website and mobile which specifically designed to address inquiries related to landscapes and famous landmarks in Vietnam. Through these platforms, users will have the easy access and take the good responses from our big dataset resources what we didn’t know before.
Context In the current era of digitalization and information technology development, Natural Language Processing (NLP) and Artificial Intelligence (AI) are becoming one of the important and rapidly growing fields. The power of NLP and AT is being widely applied to many aspects of life, from mobile applications to computational thinking, from data prediction to human-computer interaction. In the field of NLP, one of the most important applications is Question Answering (QA) systems. QA systems help computers understand and answer questions posed in natural language, creating intelligent interaction between humans and computers.
QA systems have had widespread applications in many fields, including online search, customer support, healthcare, education, and many others. However, within the framework of this thesis, we will focus on a specific application of the QA system: answering questions about landscapes and tourist attractions in Vietnam in Vietnamese. Scenic landscapes and tourist attractions are an important part of culture and daily life in Vietnam. The unique features and beauty of tourist destinations in Vietnam have attracted the attention of many people both at home and abroad.
Therefore, developing a QA system specializing in landscape and tourism in Vietnam capable of providing information and answering questions about famous landmarks and scenic spots will bring great value to the community.