VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY HA MINH DUC SUPPORTING VOICE COMMUNICATION IN CHATBOT Major: COMPUTER SCIENCE Major code: 8480101 MASTER’S THESIS HO CHI MINH CITY, January 2024 THIS THESIS IS COMPLETED AT HO CHI MINH UNIVERSITY OF TECHNOLOGY – VNU-HCM Supervisor: Le Thanh Van, Ph. D Examiner 1: Ton Long Phuoc, Ph. D Examiner 2: Vo Dang Khoa, Ph. D This master’s thesis is defended at Ho Chi Minh City University of Technology (HCMUT) – VNU-HCM on 23rd Jan 2024 Master’s Thesis Committee: 1.
Tran Van Hoai, Ph. Ton Long Phuoc, Ph. Vo Dang Khoa, Ph. Le Thanh Van, Ph.
Tran Ngoc Thinh, Ph. D Secretary Approval of the Chairman of the Master’s Thesis Committee and Dean of Faculty of Computer Science and Engineering after the thesis being corrected (If any). CHAIRMAN OF THESIS COMMITTEE DEAN OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING i VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness THE TASK SHEET OF MASTER’S THESIS Full name: Ha Minh Duc Student code: 2270348 Date of birth: 20/03/1985 Place of birth: Kien Giang Major: Computer Science Major code: 8480101 I. THESIS TITLE: Supporting voice communication in chatbot.
TASKS AND CONTENTS: Task 1: Research and Experimentation for Sequence-to-Sequence Model Development. The primary objective is to create a powerful Sequence-to-Sequence model tailored for chatbot applications. This is a neural network architecture known for its success in natural language processing tasks, and this task focuses on exploring, experimenting, and optimizing the Sequence-to-Sequence model to enhance its performance in the context of chatbot interactions. Task 2: Research and Experimentation for Automatic Speech Recognition Model Development.
During this phase, the primary emphasis is on thorough research and experimentation with diverse methods to craft high-performance automatic speech recognition models. Exploring various techniques is essential to achieving precise conversion from audio to text. The objective is to pinpoint the most effective model that aligns with the project's requirement. Task 3: Sequence-to-Sequence Model and Automatic Speech Recognition Evaluation and Future Work.
After developing the Sequence-to-Sequence model for Automatic Speech Recognition, a comprehensive evaluation process will be conducted. The achieved results will be analyzed in detail using appropriate metrics and techniques to assess accuracy performance. The strengths and weaknesses of ii each Sequence-to-Sequence model will be identified and assessed meticulously. Based on this analysis, recommendations for future work will be provided, addressing potential improvements and further developments in Automatic Speech Recognition technology.
THESIS START DAY: Feb-06-2023 IV. THESIS COMPLETION DAY: Dec-10-2023 V. SUPERVISOR: Le Thanh Van, Ph. D Ho Chi Minh City, January 22, 2024 SUPERVISOR CHAIR OF PROGRAM COMMITTEE (Full name and signature) (Full name and signature) DEAN OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING (Full name and signature) iii ACKNOWLEDGEMENT First of all, I would like to appreciate the dedicated guidance and support of my lecturers during my work.
They not only enthusiastically shared with my business plan but also suggested me many great ideas that helped my thesis much more fruitful and interesting. Especially to Dr. Le Thanh Van who has been always willing to assist me in any way he could during my thesis. Furthermore, I would also like to acknowledge the Ho Chi Minh City University of Technology for its engagement and valuable learning experiences that have given to me as well as my classmates in Vietnam.
Thanks to HCMUT’s personalized programs, I have had an ideal Master of Computer Science course during my busy working schedule. The next, I am eternally grateful for the network in which my friends Pham Thanh Huu, Nguyen Thi Ty, Vo Thi Kim Nguyet, Le Duc Huy, Nguyen Tan Sang and Pham Dien Khoa. They are not only experiencing the course with me, but also sharing life and working tips that are meaning for a person like me. I am so lucky to have them in my entire life.
The last, words cannot describe how thankful I am for IMP Academic Team’s kindly support. Without their accompany, I could not complete my Master of Computer Science course. Sincerely, Ha Minh Duc Ho Chi Minh City, Jan 2024 iv ABSTRACT This master's thesis delves into the improvement of voice-based communication in healthcare chatbots through the integration of cutting-edge natural language processing and automatic speech recognition technologies. The research centers on leveraging the GPT-3-based sequence-to-sequence architecture for enhancing natural language understanding and generation.
Additionally, it incorporates the innovative Way2vec 2.0 model to empower robust Automatic Speech Recognition capabilities. The GPT-3 architecture is chosen for its adeptness in comprehending medical contexts, generating contextually relevant responses, and handling dynamic healthcare-related conversational flows. The integration of Way2vec 2.0 ensures precise and context-aware transcription of voice inputs, enhancing the accuracy of healthcare-related information capture. This research contributes to the field of healthcare technology by presenting a novel approach to improving patient engagement and satisfaction through voice interactions.
The combination of GPT-3 and Way2vec 2.0 not only strengthens the chatbot's ability to understand and generate natural language responses but also extends this proficiency to healthcare-focused voice interactions, thereby widening the applicability and accessibility of chatbot system in the aspect of speech recognition. v TÓM TẮT LUẬN VĂN THẠC SĨ Luận văn thạc sĩ này tập trung vào việc cải thiện giao tiếp dựa trên giọng nói trong chatbot chăm sóc sức khỏe thông qua sự kết hợp của các công nghệ xử lý ngôn ngữ tự nhiên và nhận dạng giọng nói tự động tiên tiến. Nghiên cứu tập trung vào việc sử dụng kiến trúc dựa trên GPT-3 cho quá trình nâng cao hiểu biết và tạo ra ngôn ngữ tự nhiên. Ngoài ra, nó kết hợp mô hình Way2vec 2.0 sáng tạo để cung cấp khả năng nhận dạng giọng nói tự động mạnh mẽ.
Kiến trúc GPT-3 được chọn vì khả năng hiểu biết về ngữ cảnh y tế, tạo ra các phản ứng liên quan đến ngữ cảnh và xử lý các luồng trò chuyện y tế động. Sự tích hợp của Way2vec 2.0 đảm bảo việc chuyển đổi chính xác và nhận thức ngữ cảnh của đầu vào giọng nói, từ đó nâng cao độ chính xác của việc thu thập thông tin liên quan đến sức khỏe. Nghiên cứu này đóng góp cho lĩnh vực công nghệ chăm sóc sức khỏe bằng cách trình bày một cách tiếp cận mới để cải thiện sự tương tác và sự hài lòng của bệnh nhân thông qua giao tiếp giọng nói. Sự kết hợp giữa GPT-3 và Way2vec 2.0 không chỉ củng cố khả năng của chatbot trong việc hiểu và tạo ra phản ứng tự nhiên bằng ngôn ngữ, mà còn mở rộng khả năng áp dụng và tiếp cận của hệ thống chatbot trong phương diện nhận dạng tiếng nói.
vi DECLARATION OF AUTHORSHIP I hereby declare that this thesis was carried out by myself under the guidance and supervision of Le Thanh Van, Ph.D; and that the work contained and the results in it are true by author and have not violated research ethics. The data and figures presented in this thesis are for analysis, comments, and evaluations from various resources by my own work and have been duly acknowledged in the reference part. In addition, other comments, reviews and data used by other authors, and organizations have been acknowledged, and explicitly cited. I will take full responsibility for any fraud detected in my thesis.
Ho Chi Minh City University of Technology (HCMUT) – VNU-HCM is unrelated to any copyright infringement caused on my work (if any). Ho Chi Minh City, Jan 2024 Author Ha Minh Duc vii TABLE OF CONTENTS LIST OF FIGURES. ix LIST OF TABLES. Target of the Thesis.
Scope of the Thesis. Hidden Markov Model (HMM). Deep Neural Networks. Artificial Neural Networks.
Convolutional Neural Network. The Vanishing Gradient Problem. Recurrent Neural Networks. Long Short-Term Memory.
The Long-Term Dependency Problem. Word Embedding Model. Frequency-based Embedding. Prediction-based Embedding.
The History of Chatbots. Using Luong’s Attention for Sequence 2 Sequence Model. Sequence to Sequence Model. Automatic Speech Recognition.
Speech Recognition Using Recurrent Neural Networks. Speech-to-Text Using Deep Learning. GPT-3 Language Model .87 ix LIST OF FIGURES Figure 1. The History of IVR [17].
HMM-based phone model [1]. A Deep Neural Network. A Simple Example of the Structure of a Neural Network. The McCulloch-Pitts Neuron [4].
The architecture of CNN [57]. Sigmoid Function and its Derivative by [43]. Underfitting, Optimal and Overfitting. Underfitting, Optimal weight decay and Overfitting.
The Recurrent Neural Network [5]. LSTM Network Architecture [6]. RNN and Short-Term Dependencies [7]. RNN and Long-Term Dependencies [7].
The Repeating Modules in an RNN Contains One Layer [7]. The Repeating Modules of an LSTM Contain Four Layers [7]. The Architecture of GRU [9]. The CBOW Model with One Input [14].
The Skip-gram Model [14]. ELIZA – The First Chatbot in the World at MIT by Joseph Weizenbaum. The Conversational between Elize and Parry [19]. Seq2Seq Model with GRUs [15].
Attention Mechanism in Seq2Seq Model [20]. Luong’s Global Attention [21]. Basic Voicebot Architecture [56]. Graph showing how the loss function change depending on the size of the trained network [50].
Graph showing how the loss function change depending on the size of the training set[50]. Taxonomy of Sequence to Sequence Models. The Transformer and GPT Architecture [22][23]. Multi-headed Attention.
Dot Product of Query and Key. Scaling Down the Attention Scores. SoftMax of the Scaled Scores. Multiply SoftMax Output with Value Vector.
Computing Multi-headed Attention. Multi-headed Attention Output. Residual Connection of the Input and Output. Decoder First Multi-Headed Attention.
Adding Mask to Scaled Matrix. Applying SoftMax function to Attention Score. The Process Flow of Multi-headed Attention. Final Stage of Transformer’s Decoder.
Transformer Architecture and Training Objectives [24]. Taxonomy of Speech Recognition .0 Latent Feature Encoder. Example of datasets for Seq2seq model. Example of datasets for ASR model.76 xii LIST OF TABLES Table 2.
Example of Co-occurrence Matrix. Intent Recognition Results. Entity Recognition Results. Chat Handoff and Fallback.
Conversation Log with Chatbot. Result of Scenarios. Result of Way2vec 2.0 on Vietnamese Audio Files .81 xiii ACRONYMS Abbreviation Standard Word AI Artificial intelligence ASR Automatic speech recognition ANN Artificial Neural Network BLEU Bilingual Evaluation Understudy CBOW Continuous Bag Of Words CNN Convolutional Neural Network DNN Deep Neural Network EOS end-of-sentence token GloVe Global Vectors GPT Generative Pretrained Transformer GRU Gated Recurrent Unit HMM Hidden Markov Model IVR Interactive voice response LSTM Long Short-Term Memory NLP Natural Language Processing NN Neural Networks PER Perplexity ReLU Rectified Linear Unit RNN Recurrent Neural Network Seq2Seq Sequence-to-Sequence SGD Stochastic Gradient Descent SVD Singular Value Decomposition WER Word Error Rate 1 CHAPTER 1. Overview The idea of human-computer interaction through natural language has been created in Hollywood movies.
3-CPO is one of the legends of the Revolutionary Army in the world of Star Wars movie. This robot guy has served through many generations of Skywalkers and is one of the top personality robots in the universe. In the series of this movie, we can see that 3-CPO not only has very similar gestures and communication to humans, but sometimes has great instructions for its owner. This is a cinematic product that is ahead of another era when it comes to predicting the future of Artificial intelligence (AI).
The Star Wars fictional movie universe is set in a galaxy where humans and alien creatures live in harmony with droids. These robots are capable of assisting people in daily life or traveling across other planets.