i HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY PHẠM ĐIỀN KHOA APPLICATION OF VISUAL QUESTION ANSWERING USING BERT INTEGRATED WITH KNOWLEDGE BASE TO ANSWER EXTENSIVE QUESTION Major:. MASTER’S THESIS HO CHI MINH CITY, July 2023 ii HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY PHẠM ĐIỀN KHOA APPLICATION OF VISUAL QUESTION ANSWERING USING BERT INTEGRATED WITH KNOWLEDGE BASE TO ANSWER EXTENSIVE QUESTION Major:. MASTER’S THESIS HO CHI MINH CITY, July 2023 i THIS THESIS IS COMPLETED AT HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY – VNU-HCM Supervisor(s): Assoc. Quản Thành Thơ Dr.
Bùi Hoài Thắng Examiner 1: Dr Trần Tuấn Anh Examiner 2: Dr Bùi Thanh Hùng This master’s thesis is defended at HCM City University of Technology, VNU- HCM City on 13th July, 2023. Master’s Thesis Committee: 1. Võ Thị Ngọc Châu 2. Secretary: Dr Phan Trọng Nhân 3.
Examiner 1: Dr Trần Tuấn Anh 4. Examiner 2: Dr Bùi Thanh Hùng 5. Commissioner: Dr Bùi Công Giao Approval of the Chairman of Master’s Thesis Committee and Dean of Faculty of … Computer Science and Engineering …after the thesis being corrected (If any). CHAIRMAN OF THESIS COMMITTEE HEAD OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING i VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness THE TASK SHEET OF MASTER’S THESIS Full name: Phạm Điền Khoa Student ID: 2170462 Date of birth: 18/05/1998 Place of birth: An Giang Major: Computer Science Major ID: 8480101 I.
THESIS TITLE (In Vietnamese): Ứng Dụng Của Visual Question Answering Sử Dụng BERT Tích Hợp Với Knowledge Base Để Trả Lời Câu Hỏi Mở Rộng. THESIS TITLE (In English): : APPLICATION OF VISUAL QUESTION ANSWERING USING BERT INTEGRATED WITH KNOWLEDGE BASE TO ANSWER EXTENSIVE QUESTION III. TASKS AND CONTENTS: Researching Visual Question Answering in natural language processing Proposing suitable approaches for Visual Question Answering Experimenting and evaluating proposed approaches IV.THESIS START DAY: 06/02/2023 V. THESIS COMPLETION DAY: 09/06/2023 VI.
Quản Thành Thơ, Dr. Bùi Hoài Thắng Ho Chi Minh City, date ……… SUPERVISOR 1 SUPERVISOR 2 CHAIR OF PROGRAM COMMITEE ( (Full name and signature) (Full name and signature) (Full name and signature) DEAN OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING (Full name and signature) ii ACKNOWLEDGEMENTS I would like to express my deepest gratitude to my parents for their unwavering support and encouragement throughout my academic journey. Their love, patience, and belief in me have been a constant source of motivation and strength. I am also immensely grateful to my thesis instructor, Assoc.
Quản Thành Thơ, for his invaluable guidance, expertise, and continuous support throughout the research process. His insightful feedback, constructive criticism, and dedication to excellence have greatly contributed to the success of this thesis. I would like to extend my appreciation to all the faculty members and staff of the Department of Computer Science for providing me with a conducive learning environment and valuable resources. Their expertise and assistance have been instrumental in shaping my research work.
Finally, I would like to thank my friends and colleagues for their friendship, encouragement, and insightful discussions, which have enriched my understanding and perspectives on the subject matter. To everyone who has played a part, big or small, in the completion of this master thesis, I am sincerely thankful for your support and assistance. iii ABSTRACT Fusing different modalities, such as image and text, to obtain important information has long been an issue in artificial intelligence. Visual Question Answering (VQA) is an emerging field that aims to develop intelligent systems capable of understanding and answering questions based on visual content.
This dissertation presents a comprehensive study on VQA with the intergration of knowledge base, focusing on the fusion of language understanding, image processing, and knowledge retrieval techniques. The primary objective of this research is to investigate the effectiveness of different baselines in the Knowledge-based Visual Question Answering task and explore their strengths and weaknesses. Two baselines were considered: KBVQA with BERT base and CNN, KBVQA with BERT large and CLIP. These models were evaluated using the ViQuAE dataset, which contains pre-classified questions with ground truth answers.
To improve the KBVQA system, several future directions are proposed. First, training the BERT models with larger and more diverse datasets could enhance their language understanding capabilities. Additionally, efforts should be made to address the challenges associated with object detection and face recognition, which were not fully implemented within the limited timeframe of this research. In conclusion, this dissertation contributes to the field of Knowledge-based Visual Question Answering by evaluating and comparing different baseline models.
It highlights the effectiveness of BERT large and CLIP in achieving accurate and relevant answers. The findings and proposed future directions provide valuable insights for researchers and practitioners interested in advancing KBVQA systems. iv TÓM TẮT LUẬN VĂN Việc hợp nhất các phương thức khác nhau, chẳng hạn như văn bản và hình ảnh, để có được các thông tin cần thiết từ lâu đã là một vấn đề trong lĩnh vực trí tuệ nhân tạo. Visual Question Answering (VQA) là một lĩnh vực mới nổi gần đây, nhằm phát triển các hệ thống AI có khả năng hiểu và trả lời các câu hỏi dựa trên hình ảnh.
Luận văn này trình bày một nghiên cứu về VQA tích hợp với Knowledge Base, tập trung vào sự kết hợp giữa việc đọc hiểu ngôn ngữ, xử lý hình ảnh và các kỹ thuật truy xuất thông tin. • Mục tiêu chính của nghiên cứu này là điều tra hiệu quả của các baseline khác nhau trong nhiệm vụ trả lời câu hỏi theo hình ảnh dựa trên knowledge base và khám phá điểm mạnh và điểm yếu của chúng. Hai baseline đã được xem xét: KBVQA với BERT base và CNN, KBVQA với BERT large và CLIP. Các mô hình này được đánh giá bằng cách sử dụng dataset ViQuAE, bao gồm các câu hỏi được phân loại trước với các câu trả lời ground truth.
• Để cải thiện hệ thống KBVQA, một số hướng tiếp theo được đề xuất. Đầu tiên, đào tạo mô hình BERT với bộ dữ liệu lớn hơn và đa dạng hơn có thể nâng cao khả năng hiểu ngôn ngữ. Ngoài ra, cần nỗ lực giải quyết các thách thức liên quan đến phát hiện đối tượng và nhận dạng khuôn mặt, vốn không được triển khai đầy đủ trong thời gian giới hạn của nghiên cứu này. Tóm lại, luận văn này đóng góp vào lĩnh vực Knowledge-based Visual Question Answering bằng cách đánh giá và so sánh các baseline khác nhau.
Thực nghiệm chỉ ra hiệu quả của baseline sử dụng BERT large và CLIP trong việc đạt được câu trả lời chính xác và phù hợp. Các nghiên cứu tiếp tục trong tương lai sẽ cung cấp những hiểu biết có giá trị cho các nhà nghiên cứu và những người thực hành quan tâm đến việc cải tiến các hệ thống KBVQA. v THE COMMITMENT OF THE THESIS’S AUTHOR I declare that this thesis, written under the supervision of Assoc. Quan Thanh Tho, was created to suit the needs of society and my abilities to obtain information.
External support should be documented, referenced, and cited. AUTHOR (Full name and signature) vi TABLE OF CONTENTS CHAPTER I: INTRODUCTION .2 OVERVIEW OF VQA AND KBVQA PROBLEM .1 Visual Question Answering (VQA) .2 Knowledge-based Visual Question Answering (KBVQA): .3 TARGET AND SCOPE OF THE THESIS.4 LIMITS OF THE THESIS .5 CONTRIBUTION OF THE THESIS. 7 CHAPTER 2: BACKGROUND KNOWLEDGE .1 CONVOLUTIONAL NEURAL NETWORKS (CNNS): .1 Architecture of CNN: .2 LONG SHORT-TERM MEMORY (LSTM): .1 Overview about Recurrent Neural Networks (RNNs): .2 Long-short term memory: .4 CONTRASTIVE LANGUAGE-IMAGE PRE-TRAINING (CLIP) .5 WIKIPEDIA KNOWLEDGE BASE. 26 CHAPTER 3: RELATED WORKS .1 APPROACHES OF KNOWLEDGE-BASED VISUAL QUESTION ANSWERING.
33 CHAPTER 4: PROPOSED SOLUTION .1 Metric Evaluation: Precision, Recall and F1 score: .2 Dataset for evaluation .1 Late fusion technique .4 SETTING-UP EXPERIMENT – TRAINING BERT:.5 FIRST BASELINE: BERT BASE WITH CNN.1 Motivation and idea:.3 Result and Discussion .6 SECOND BASELINE: BERT LARGE AND CLIP .1 Motivation and idea:.3 Result and Discussion .7 EXPERIMENT WITH GPT-2 FOR 2ND BASELINE .1 Motivation and idea:.3 Result and Discussion. 69 viii LIST OF TABLES Table 1.1: Computer vision sub-tasks required to be solved by VQA.1: Google Colabs’s Specs.2: GPUs available in Colab, Colab Pro, and Colab Pro+.3: KBVQA system (version BERT base + CNN)’s parameters.4: Result of KBVQA system (BERT base + CNN).5: KBVQA system (version BERT large + CLIP)’s parameters.6: Result of KBVQA system (BERT large + CLIP).61 ix TABLE OF FIGURES Figure 1.1: Illustration of VQA's tasks.2: The question and relevant item in the Knowledge Base.1: Convolutional Neural Network.2 Comparison of 20-layer vs 56-layer architecture.3 Skip (Shortcut) connection of ResNet.5 Recurrent Neural Networks.6: Long-short-term-memory’s structure.9: Contrastive Pre-training CLIP.10: Create dataset classifier from lable text and use it for zero-prediction .1: Knowledge-based Visual Question Answering’s Taxonomy.1: Types of question that you can ask between traditional VQA and Knowledge-based VQA.2: The overview of "Show, Ask, Attend, and Answer: A Strong Baseline For Visual Question Answering" VQA system.3: “KVQA: Knowledge-Aware Visual Question Answering” VQA system .4: KBVQA system with BERT base and CNN.5: Generated answer about Barack Obama from the KBVQA (BERT base + CNN) system.6:Bad generated answer Isaac Newton from the KBVQA (BERT base + CNN) system.7: KBVQA system with BERT and CLIP.8: Good generated answer about Amelia Earhart from the KBVQA (BERT large + CLIP) system.9: KBVQA system with BERT, CLIP and GPT-2.10: Generated answer with more information from KBVQA system (BERT Large +CLIP +GPT-2) about Elvis Presley.1 Research Problem: Fusing multiple modalities, such as image and text, to retrieve relevant information is a long-standing problem in the field of artificial intelligence. In recent years, significant progress has been made in many subdomains of machine learning. Neural networks are now capable of solving Computer Vision and Natural Language Processing tasks in a much different way, with greater speed and accuracy.
One of the most challenging and promising tasks in this area is Visual Question Answering (VQA), which aims to automatically generate an accurate answer to a natural language question about an image. It was once thought that developing a computer vision system capable of answering arbitrary natural language questions about images was an ambitious but intractable goal. However, since 2016, there has been tremendous progress in developing systems with these capabilities. VQA systems aim to correctly answer natural language questions about an image input and comprehend the contents of an image in the same way that humans do, while also communicating effectively about that image in natural language.
With the increasing availability of large-scale annotated datasets, deep learning-based approaches have recently achieved remarkable progress in VQA. However, most existing VQA models only focus on answering questions about objects or actions in images, one of the main challenges is the lack of knowledge representation in VQA systems, which limits their ability to handle complex questions that require a deeper understanding of the image content. There was an idea about using the visual question answering system to answer more extensive questions about the information of the entities in the image, and to do that, the idea of retrieving the information contained in the knowledge base that has been applied to VQA. Combining all of this, a new task called Knowledge-based Visual Question Answering (KBVQA) has been proposed, which requires a model to retrieve relevant information about a entity in the image from a knowledge base and use it to answer questions.
2 To address the limitation of traditional Visual Question Answering that we mention before, Knowledge-based Visual Question Answering (KBVQA) has emerged as a another more specialized direction for VQA field. KBVQA requires external knowledge (knowledge bases - KBs) beyond the image to answer the question.