DEVELOP A MULTIMODAL CHATBOT FOR A FASHION STORE

Báo cáo dự án: Phát triển chatbot đa phương thức cho cửa hàng thời trang. Ứng dụng AI, học sâu tăng trải nghiệm mua sắm, hỗ trợ tìm kiếm sản phẩm.

Chuyên ngành

Computer Science

Người đăng

Ẩn danh

Thể loại

Specialized Project

Academic Year

105
1
0

Phí lưu trữ

35 Point

Mục lục chi tiết

Declaration Of Authenticity

Acknowledgement

Abstract

Table of Glossary

List of Tables

List of Figures

1. Chương 1: Introduction

1.1. Motivation

1.2. Goals

1.3. Scope

1.4. Thesis Structure

2. Chương 2: Related Works

2.1. Vision Transformer

2.2. Residual Neural Network

2.3. Language Models

2.3.1. Traditional Language Models

2.3.2. PhoBERT: Pre-trained language models for Vietnamese

2.3.3. Other Large Language Models

2.4. Retrieval-Augmented Generation

2.4.1. Combining Retriever with Generative Models

2.4.2. Combining Multimodal Model with Vector Database

2.5. Multimodal Image-Text Contrastive Learning

2.6. Limitations and Challenges

3. Chương 3: Theoretical Background

3.1. Multi-Layer Perceptron

3.2. Binary Cross-Entropy Loss

3.3. Convolutional Neural Network

3.4. Recurrent Neural Network

3.4.1. Back-propagation through Time

3.5. Long-Short Term Memory

3.6. Large Language Models

5. Chương 5: Conclusion

References

Appendix A: Inference on PhoCLIP

A.1. Text-only queries

A.2. Image-only queries

A.3. Image-text combined queries

Appendix B: Table of Workload

Tóm tắt

I. Tổng Quan Chatbot Đa Phương Thức Cho Thời Trang AI

Chatbot đa phương thức đang nổi lên như một công cụ mạnh mẽ trong ngành thời trang, mang đến trải nghiệm tương tác khách hàng hoàn toàn mới. Sự kết hợp giữa học sâutrí tuệ nhân tạo (AI) cho phép chatbot không chỉ hiểu ngôn ngữ tự nhiên mà còn xử lý hình ảnh, giọng nói, từ đó cung cấp tư vấn và hỗ trợ chính xác, cá nhân hóa. Ứng dụng chatbot thời trang giúp các cửa hàng tăng cường tương tác, cải thiện trải nghiệm mua sắm và tối ưu hóa quy trình vận hành. Chìa khóa thành công nằm ở khả năng xây dựng chatbot thông minh, am hiểu sâu sắc về sản phẩm, xu hướng thời trang và nhu cầu của từng khách hàng. Theo báo cáo, việc triển khai chatbot AI có thể giúp tăng doanh số bán hàng lên đến 25% và giảm chi phí hỗ trợ khách hàng đáng kể. Hãy cùng tìm hiểu cách ứng dụng công nghệ này vào thực tế.

1.1. Định Nghĩa và Ưu Điểm Chatbot Đa Phương Thức

Chatbot đa phương thức là một hệ thống tương tác có khả năng xử lý nhiều loại dữ liệu đầu vào khác nhau như văn bản, hình ảnh, giọng nói. Điều này cho phép chatbot hiểu ngữ cảnh một cách toàn diện hơn và cung cấp phản hồi phù hợp hơn cho người dùng. Ưu điểm vượt trội bao gồm khả năng tư vấn thời trang dựa trên hình ảnh, hỗ trợ tìm kiếm sản phẩm bằng giọng nói và cá nhân hóa trải nghiệm mua sắm dựa trên phân tích dữ liệu khách hàng. Sự linh hoạt này mang lại trải nghiệm người dùng vượt trội so với các chatbot truyền thống.

1.2. Vai Trò Của AI và Học Sâu Trong Chatbot Thời Trang

Học sâutrí tuệ nhân tạo là nền tảng cốt lõi của chatbot đa phương thức thông minh. Các mô hình mạng nơ-ron sâu được sử dụng để huấn luyện chatbot hiểu ngôn ngữ tự nhiên, nhận diện hình ảnh và dự đoán hành vi người dùng. Khả năng tự học và cải thiện theo thời gian giúp chatbot ngày càng trở nên thông minh hơn, đáp ứng tốt hơn nhu cầu của khách hàng. Ứng dụng các thuật toán xử lý ngôn ngữ tự nhiên (NLP)Computer Vision là yếu tố then chốt để xây dựng chatbot hiệu quả.

II. Thách Thức Phát Triển Chatbot Thời Trang Hiệu Quả Cá Nhân Hóa

Mặc dù tiềm năng to lớn, việc phát triển chatbot thời trang hiệu quả vẫn đối mặt với nhiều thách thức. Khả năng hiểu và phản hồi chính xác ngôn ngữ tự nhiên phức tạp, đặc biệt là các thuật ngữ chuyên ngành thời trang, vẫn còn hạn chế. Bên cạnh đó, việc thu thập và xử lý dữ liệu đủ lớn để huấn luyện các mô hình học sâu đòi hỏi nguồn lực đáng kể. Vấn đề cá nhân hóa trải nghiệm mua sắm cũng đặt ra yêu cầu cao về khả năng phân tích dữ liệu khách hàng và đưa ra gợi ý phù hợp. Quan trọng hơn, đảm bảo tính bảo mật và quyền riêng tư cho dữ liệu cá nhân của khách hàng là yếu tố then chốt để xây dựng lòng tin.

2.1. Khó Khăn Trong Xử Lý Ngôn Ngữ Tự Nhiên NLP Thời Trang

Ngôn ngữ trong ngành thời trang rất đa dạng và phức tạp, bao gồm nhiều thuật ngữ chuyên ngành, biệt ngữ và cách diễn đạt mang tính cá nhân. Việc huấn luyện chatbot hiểu và phản hồi chính xác các câu hỏi liên quan đến phong cách, chất liệu, kiểu dáng và xu hướng đòi hỏi mô hình NLP phải được đào tạo trên một lượng lớn dữ liệu chuyên biệt. Thêm vào đó, khả năng xử lý các câu hỏi mơ hồ, không rõ ràng cũng là một thách thức lớn.

2.2. Vấn Đề Thu Thập và Xử Lý Dữ Liệu Huấn Luyện Chatbot

Để xây dựng một chatbot AI thông minh, cần thu thập một lượng lớn dữ liệu huấn luyện, bao gồm các cuộc hội thoại thực tế giữa khách hàng và nhân viên tư vấn, hình ảnh sản phẩm, thông tin về xu hướng thời trang và dữ liệu cá nhân của khách hàng. Quá trình này tốn kém và đòi hỏi sự tuân thủ nghiêm ngặt các quy định về bảo mật dữ liệu. Việc xử lý và làm sạch dữ liệu cũng là một công đoạn quan trọng để đảm bảo chất lượng của mô hình học sâu.

III. Phương Pháp Học Sâu và Computer Vision Cho Chatbot Fashion AI

Để vượt qua các thách thức, việc ứng dụng các phương pháp học sâuComputer Vision tiên tiến là rất quan trọng. Các mô hình mạng nơ-ron tích chập (CNN) có thể được sử dụng để phân tích hình ảnh sản phẩm, nhận diện kiểu dáng và màu sắc. Mô hình ngôn ngữ lớn (LLM) như GPT-3 có khả năng tạo ra các phản hồi tự nhiên và phù hợp với ngữ cảnh. Việc kết hợp các phương pháp này cho phép chatbot không chỉ hiểu văn bản mà còn "nhìn" và "suy nghĩ", từ đó cung cấp tư vấn chính xác và cá nhân hóa.

3.1. Sử Dụng Mạng Nơ Ron Tích Chập CNN Cho Nhận Diện Hình Ảnh

Các mô hình CNN có khả năng phân tích hình ảnh sản phẩm để nhận diện các đặc điểm như kiểu dáng, màu sắc, chất liệu và hoa văn. Thông tin này có thể được sử dụng để cung cấp các gợi ý trang phục phù hợp, tìm kiếm sản phẩm tương tự và hỗ trợ khách hàng lựa chọn sản phẩm ưng ý. Computer Vision là yếu tố then chốt để chatbot có thể "nhìn" và hiểu thế giới xung quanh.

3.2. Ứng Dụng Mô Hình Ngôn Ngữ Lớn LLM Như GPT 3 Tạo Phản Hồi

Mô hình ngôn ngữ lớn (LLM) như GPT-3 có khả năng tạo ra các phản hồi tự nhiên, mạch lạc và phù hợp với ngữ cảnh. Điều này giúp chatbot tương tác với khách hàng một cách tự nhiên và thân thiện hơn. Việc huấn luyện LLM trên dữ liệu chuyên ngành thời trang cho phép chatbot hiểu và phản hồi chính xác các câu hỏi liên quan đến sản phẩm, xu hướng và phong cách.

IV. Ứng Dụng Chatbot Đa Phương Thức Tăng Trưởng Doanh Số Bán Lẻ

Chatbot đa phương thức không chỉ là một công cụ hỗ trợ khách hàng mà còn là một kênh bán hàng hiệu quả. Chatbot có thể được sử dụng để quảng bá sản phẩm mới, cung cấp thông tin khuyến mãi, gợi ý trang phục phù hợp và hỗ trợ khách hàng hoàn tất quá trình mua hàng. Khả năng tương tác 24/7 giúp cửa hàng không bỏ lỡ bất kỳ cơ hội bán hàng nào. Theo nghiên cứu, việc triển khai chatbot có thể giúp tăng doanh số bán hàng lên đến 30% và giảm chi phí marketing đáng kể. Tương tác khách hàng tự động giúp tăng hiệu quả hoạt động.

4.1. Chatbot Tư Vấn Thời Trang Ảo và Gợi Ý Trang Phục Thông Minh

Tư vấn thời trang ảo là một trong những ứng dụng phổ biến nhất của chatbot trong ngành thời trang. Chatbot có thể thu thập thông tin về phong cách, sở thích và dáng người của khách hàng, sau đó đưa ra các gợi ý trang phục phù hợp. Khả năng phân tích hình ảnh và dữ liệu khách hàng giúp chatbot cung cấp các gợi ý cá nhân hóa, giúp khách hàng tìm được những bộ trang phục ưng ý một cách dễ dàng.

4.2. Hỗ Trợ Khách Hàng 24 7 và Xử Lý Yêu Cầu Nhanh Chóng

Hỗ trợ khách hàng 24/7 là một lợi ích lớn của việc triển khai chatbot. Chatbot có thể trả lời các câu hỏi thường gặp, cung cấp thông tin về sản phẩm và xử lý các yêu cầu của khách hàng một cách nhanh chóng và hiệu quả. Điều này giúp cải thiện trải nghiệm khách hàng và giảm tải cho đội ngũ hỗ trợ.

V. Giải Pháp Tích Hợp Chatbot Fashion AI với Các Nền Tảng

Để triển khai thành công chatbot đa phương thức, việc tích hợp với các nền tảng thương mại điện tử, mạng xã hội và ứng dụng nhắn tin là rất quan trọng. Các nền tảng như Dialogflow, Rasa, PyTorchTensorFlow cung cấp các công cụ và thư viện mạnh mẽ để xây dựng và triển khai chatbot AI. Việc lựa chọn nền tảng phù hợp phụ thuộc vào yêu cầu cụ thể của dự án và nguồn lực hiện có. Quan trọng hơn, cần xây dựng một quy trình quản lý và bảo trì chatbot hiệu quả để đảm bảo tính ổn định và hiệu suất.

5.1. Tích Hợp Chatbot Với Các Nền Tảng Thương Mại Điện Tử

Việc tích hợp chatbot với các nền tảng thương mại điện tử như Shopify, WooCommerce và Magento cho phép chatbot truy cập vào thông tin sản phẩm, dữ liệu khách hàng và lịch sử giao dịch. Điều này giúp chatbot cung cấp tư vấn chính xác và cá nhân hóa, hỗ trợ khách hàng hoàn tất quá trình mua hàng và tăng doanh số bán hàng. Chatbot cho thương mại điện tử là một xu hướng tất yếu.

5.2. Sử Dụng Dialogflow Rasa PyTorch TensorFlow Để Phát Triển

DialogflowRasa là các nền tảng phát triển chatbot phổ biến, cung cấp các công cụ trực quan và dễ sử dụng. PyTorchTensorFlow là các thư viện học sâu mạnh mẽ, cho phép xây dựng các mô hình AI phức tạp. Việc lựa chọn nền tảng phù hợp phụ thuộc vào yêu cầu cụ thể của dự án và kỹ năng của đội ngũ phát triển.

VI. Tương Lai Chatbot Thời Trang AI Trợ Lý Cá Nhân Ảo

Tương lai của chatbot thời trang hứa hẹn nhiều tiềm năng phát triển. Chatbot có thể trở thành một trợ lý cá nhân ảo, cung cấp tư vấn thời trang chuyên nghiệp, gợi ý trang phục phù hợp với từng dịp và giúp khách hàng quản lý tủ quần áo thông minh. Việc tích hợp với các công nghệ mới như thực tế ảo (VR)thực tế tăng cường (AR) sẽ mang đến trải nghiệm mua sắm chân thực và sống động hơn. Quan trọng hơn, sự phát triển của AI sẽ cho phép chatbot tự học và cải thiện liên tục, đáp ứng tốt hơn nhu cầu của khách hàng.

6.1. Chatbot AI Trở Thành Trợ Lý Cá Nhân Ảo Cho Khách Hàng

Trong tương lai, chatbot AI có thể đóng vai trò là một trợ lý cá nhân ảo, giúp khách hàng lựa chọn trang phục, quản lý tủ quần áo và cập nhật các xu hướng thời trang mới nhất. Chatbot có thể học hỏi từ lịch sử mua hàng, phong cách cá nhân và các sự kiện quan trọng trong cuộc sống của khách hàng, từ đó cung cấp các gợi ý phù hợp và cá nhân hóa.

6.2. Ứng Dụng Thực Tế Ảo VR và Tăng Cường AR Trong Chatbot

Việc tích hợp thực tế ảo (VR)thực tế tăng cường (AR) sẽ mang đến trải nghiệm mua sắm chân thực và sống động hơn cho khách hàng. Chatbot có thể giúp khách hàng thử đồ ảo, xem các bộ trang phục trên người mình và trải nghiệm không gian mua sắm ảo một cách dễ dàng. Điều này giúp tăng tính tương tác và hấp dẫn cho trải nghiệm mua sắm trực tuyến.

Tóm tắt và mô tả trên trang này được tạo với sự hỗ trợ của AI. Nếu bạn thấy nội dung không chính xác hoặc có vấn đề, vui lòng Báo lỗi nội dung.

24/04/2025
Report specialized project semester 232 academic year 2023 2024 develop a multimodal chatbot for a fashion store

Trích đoạn nội dung tài liệu

VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY FACULTY OF COMPUTER SCIENCE AND ENGINEERING đà TP.HCM REPORT SPECIALIZED PROJECT SEMESTER 232 ACADEMIC YEAR 2023-2024 DEVELOP A MULTIMODAL CHATBOT FOR A FASHION STORE MAJOR: COMPUTER SCIENCE COUNCIL: COMPUTER SCIENCE - 01 CLC SUPERVISOR(s): QUAN THANH THO, Ph. SECRETARY: TRAN HUY, MEng. o0o STUDENT 1: VOHOANG NHAT KHANG - 2152646 STUDENT 2: NGUYEN PHAN TRI DUC - 2152528 HO CHI MINH CITY, May 2024 Instructor’s Signature Date: Assoc. Quan Thanh Tho, Ph.

(Instructor) Associate Professor Faculty of Computer Science and Engineering Declaration Of Authenticity We declare that we solely conducted this specialized project, under the supervision of Assoc. Quan Thanh Tho at the Faculty of Computer Science and Engineering, Vietnam National University - Ho Chi Minh City University Of Technology. We have taken care to properly acknowledge and document all external sources and references used in the project. If there is any instance of plagiarism, we are ready to accept the consequences.

Ho Chi Minh City University of Technology - Vietnam National University HCMC will not be held responsible for any copyright violations that may have occurred during my research. Ho Chi Minh City, May 2024 Authors, Vo Hoang Nhat Khang, Nguyen Phan Tri Duc Acknowledgement We would like to express my appreciation to Assoc. Quan Thanh Tho for his invaluable guidance. Our research has greatly benefited from his deep knowledge, per- ceptive criticism, and constant support.

Besides, we are grateful for the help from Mr. Nguyen Hieu Nghĩa from University of Information Technology - Vietnam National University (UIT - VNU) for valuable comments for research aspect and our proposed pipeline, including pointing out strengths as well as weaknesses of previous works on multimodal models so that we can elaborate on our ideas more effectively. In addition, we would like to extend our gratitude to my family. Their unwavering faith in our abilities and constant encouragement have been my pillars of strength.

Their belief in my potential has been a constant source of motivation and resilience. We are forever grateful for their love and support. 1H Abstract This thesis explores the integration of multimodal deep learning models into search for fashion accessories, enhancing the user shopping experience through interactive capa- bilities. Our research focuses on developing a robust retrieval system that leverages both image and text inputs to provide accurate product suggestions.

The system architecture allows users to input images of desired products along with additional textual descrip- tions to refine their search queries. Utilizing state-of-the-art multimodal language mod- els, the system retrieves relevant product recommendations based on the similarity of visual and semantic features. Key functionalities include providing detailed information about product availability, pricing, and suggesting alternative products based on specific customer preferences. The interactive nature of the system ensures a tailored shopping experience, enabling users to refine their search criteria and explore a diverse range of fashion accessories.

Preceded by comprehensive research on multimodal deep learning techniques frame- works, this thesis contributes to advancing the field of information retrieval. The ex- perimental evaluation demonstrates the efficacy of the proposed approach in delivering accurate and relevant items, ultimately enhancing customer satisfaction and engagement in online shopping environments. IV Table of Glossary Term Definition Note Multimodal Relating to multiple modes of in- put/output, such as text and images, used in communication or computing systems. Chatbot A computer program designed to sim- ulate conversation with human users, typically over the internet.

RAG System Retrieval-Augmented Generation sys- tem, a framework that combines re- trieval and generation models for en- hanced natural language understand- ing and generation. CLIP Encoder Contrastive Language-Image Pretrain- ing Encoder, a deep learning model that learns joint representations of images and text through contrastive learning. Term Definition Note End-to-end A design process or system designed to operate without intermediate stages or interactions from separate compo- nents. In other words, it refers to the end-to-end nature of a process from start to finish without intermedi- ate steps.

Phase A stage or step. For example, to solve a large problem, we would address smaller tasks (in phases). In the scope of the topic, there are phase | and phase 2. Transformer A machine learning model architecture Reference [49] proposed in 2017, achieving high ef- ficiency in Natural Language Process- ing (NLP).

Pretraining The process of training a machine learning model before it is fine-tuned for a specific task. In this stage, the model learns from large and diverse datasets to understand general patterns and representations, forming a strong semantic understanding. Pretraining 1s an important step in building high- performance models for various tasks. vi Term Definition Note Fine-tuning In machine learning, fine-tuning is the process of adjusting a model that has been pre-trained on a small or spe- cific dataset.

The goal is to make the model understand more and apply spe- cific knowledge from new data or spe- cific tasks. Vil Contents 1 Introduction LL Motivation. Quy va cố.ee WwW 2 Related Works Œứ 2.3 PhoBERIT: Pre-trained language models for Vietnamese 2.4 Other Large Language Models.3 Retrleval-Augmented Generalion. Combining Retrlever with Generatve Models.4 Combining Multimodal Model with Vector Database.4 Multimodal Image-Text Contrasttve Learning .2 Limttatons and Challenges.

vill 3 Theoretical Background 3.2 Multi-Layer Pereeptron.2 Bimary Cross-EntropyLoss.4 Convolutional Neural Network.5 Recurrent Neural Network .3 Back-propagation through Time. 36 Long-Short Term Memory .10 Large Language Models. 67 5 Conclusion 69 References 70 A Inference on PhoCLIP 77 A.I Text-only queries .2 Image-only queries .3 Image-textcombined queles. 86 B Table of Workload 90 B.0 0000000000 eee ee 90 List of Tables 4.4 Average cosine similarity.

65 x1 List of Figures 2.2 In a regular block (left), the portion within the dotted-line box must di- rectly learn the mapping f(x). In a residual block (right), the portion within the dotted-line box needs to learn the residual mapping g(x) = f(x) — x, making the identity mapping f(x) = x easier to learn’.3 Retrieval-Augmented Œeneration pIpeline.4 CLIP jointly trains an image encoder and a text encoder to predict the correct pairings of a batch of (image, text) training examples. At test time the learned text encoder synthesizes a zero-shot linear classifier by embedding the names or descriptions of the target dataset’s classes.1 MLP architecture with 2 hidden layers.2 Graph ofthe sigmoid functlon.3 Graph ofthe tanh funcion.4 Graph ofthe ReLU function.5 Filter and stridein CNN 2.6 Pooling demonstrations with Max Pooling and Average Pooling 3.7 RNN General Architecture 2.8 One-to-One RNN archtectre .9 One-to-Many RNNarchtecue .10 Many-toe-OneRNNarchtecue .11 Many-to-Many RNN archttectire.12 LSTM architecture and íts cell structure'.13 Anexample of Tokenizer.14 Transformer architecture from paper ’’Attention Is All You Need” [49] .15 Multi-head Attention architecture from paper Attention Is All You Need” [49] 2 v2 53 3.16 Early Fusion pipeline.17 Late Fusion pipeline. 00 57 41 Architecture of PhoCLIP, adopt from OpenAI CLIP.

67 xill Chapter 1 Introduction In chapter I, the overview, objectives, and goals of the research project are illustrated. The outline of the report is also presented.1 Motivation As the fashion market expands to cater to diverse demographics and preferences, the con- ventional methods of searching for products become increasingly inefficient. Consumers often face challenges in finding fashion items that align with their specific preferences, budget constraints, and other unique requirements. Traditional approaches rely heavily on human-assisted interactions through websites, where customers engage in conversa- tions with store managers to seek guidance or recommendations.

However, this manual process is time-consuming, prone to errors, and fails to leverage the full potential of technology in enhancing customer experiences. The inefficiency of this process raises a significant challenge: How can this whole process be automated to enhance customer experience and streamline operations for fashion stores? By introducing a multimodal searching method, we aim to revolutionize the way cus- tomers interact with fashion stores online. The integration of image and text capabili- ties allows for a more intuitive and efficient communication channel between the con- sumer and the store. With this technology, customers can simply provide an image or description of the desired product, and the system can quickly analyze their preferences and recommend suitable options.

Through fine-tuning a Vietnamese multimodal model and implementing an efficient information retrieval pipeline, the system can understand customer preferences, recommend appropriate products, and provide personalized assis- tance, thus revolutionizing the way customers interact with fashion stores online. Moreover, the development of such a system addresses the pressing need for automation in the fashion retail sector. As the number of fashion stores continues to rise, streamlin- Ing operations and improving customer experiences become paramount. A well-trained multimodal search system can effectively assist customers in navigating through the vast array of products, providing personalized recommendations, and answering queries re- garding price, availability, and discounts.2 Goals The goals of this project encompass both technical advancements and practical outcomes aimed at revolutionizing the fashion retail experience.

Firstly, the primary objective is to develop and deploy a sophisticated multimodal chatbot model capable of efficiently assisting customers in navigating the diverse offerings of fashion stores. This entails in- tegrating image and text processing capabilities to accurately interpret customer requests and provide relevant product recommendations. Additionally, a key goal is to enhance the overall customer experience by reducing the time and effort required for shopping online. By leveraging the chatbot’s capabilities in understanding customer preferences and responding promptly to inquiries, the project aims to streamline the shopping process, making it more intuitive and convenient for users.

Another crucial aspect of the project is to optimize the efficiency of fashion store operations through automation. This efficiency gain not only improves the productiv- ity of fashion stores but also reduces operational costs and enhances overall business performance. Moreover, the project seeks to foster innovation in the field of natural language pro- cessing and multimodal technology. By pushing the boundaries of what is achievable in terms of understanding and responding to customer queries, the development of the chatbot contributes to advancements in Al-driven conversational systems.

This not only benefits the fashion industry but also has broader implications for other sectors seeking to enhance customer interactions through automated solutions.3 Scope Firstly, our project involves designing and implementing the core chatbot architec- ture that integrates both image and text processing capabilities. The chatbot should be able to accurately interpret customer inputs, whether they are in the form of im- ages depicting desired products or text descriptions. Secondly, we will develop a model that can give most similar items to user based on user’s input image and/or text prompt. Our baseline model to start this project will be CLIP [36], and we also propose a new model for encoding Vietnamese language along with the image, namely PhoCLIP.

Thirdly, the project also involves implementing efficient information retrieval mech- anisms to fetch product details such as price, availability, discounts, and other rele- vant information from the fashion store’s database or external sources by Retrieval- Augmented Generation (RAG). This ensures that the chatbot can provide accurate and up-to-date information to customers in real-time. Finally, is our effort to design an intuitive and user-friendly interface for interacting with the chatbot is crucial for enhancing the overall user experience, by creating conversational flows, interactive elements, and visual representations to facilitate seamless communication between the chatbot and the user.4 Thesis Structure There are five chapters in our project: Chapter 1 provides a comprehensive introduction to the motivation behind the problem, its goals, objectives, and the scope of the project. Chapter 2 highlights previous research and related works pertaining to the task at hand, aiming to gain insights into existing methodologies and their limitations.

* Chapter 3 delves into the foundational concepts of Natural Language Processing (NLP), Vision models and Multimodal models, which are essential for the project’s implementation. ¢ Chapter 4 is dedicated to our approach about the implementation details of our project, including the development of the multimodal chatbot and associated sys- tems. ¢ Chapter 5 concludes our project by summarizing the key findings and outcomes, while also outlining our plans for future development and improvements.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ