VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY HO CHI MINH UNIVERSITY OF TECHNOLOGY NGUYỄN ĐỨC PHÚ PREDICTING QUALITY OF HOME CARE LIQUID PRODUCTS Major: Computer Science Major code: 8480101 MASTER’S THESIS HO CHI MINH CITY, June 2024 THIS THESIS IS COMPLETED AT HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY – VNU-HCM Supervisor: 1. Lê Thành Sách Examiner 1: Dr. Hà Việt Uyên Sinh - Ho Chi Minh International University Examiner 2: Dr. Nguyễn Quang Hùng This master’s thesis is defended at HCM City University of Technology, VNU- HCM City on June 18th, 2024.
Master’s Thesis Committee: 1. Nguyễn Lê Duy Lai 2. Lê Thanh Vân 3. Hà Việt Uyên Sinh 4.
Nguyễn Quang Hùng 5. Huỳnh Tường Nguyên Approval of the Chair of Master’s Thesis Committee and Dean of Faculty of Computer Science and Engineering after the thesis being corrected (If any). CHAIR OF THESIS COMMITTEE DEAN OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING VIETNAM NATIONAL UNIVERSITY - HO CHI MINH CITY SOCIALIST REPUBLIC OF VIETNAM HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY Independence – Freedom - Happiness THE TASK SHEET OF MASTER’S THESIS Full name: Nguyễn Đức Phú Student ID: 2171014 Date of birth: 19/06/1999 Place of birth: Hồ Chí Minh Major: Computer Science Major ID: 8480101 I. THESIS TITLE (In Vietnamese): Giải pháp dự đoán chất lượng sản phẩm nước vệ sinh nhà cửa II.
THESIS TITLE (In English): Predicting quality of home care liquid products III. TASKS AND CONTENTS: Task Content Timeline Literature Review Review and analyze key theories, models and 01/03/2024 studies related to quality prediction and explanable AI. Methodology Describe the data sets, data preprocessing and 15/03/2024 feature engineering in the thesis. Result and Analytics Run analysis and interpret the results.
15/04/2024 Conclusion Summarize key contributions and limitations. THESIS START DAY: 15/01/2024 V. THESIS COMPLETION DAY: 20/05/2024 VI. SUPERVISOR (Please fill in the supervisor’s full name and academic rank) Associate Professor Thoại Nam Doctor Lê Thành Sách Ho Chi Minh City, June 2024 SUPERVISOR SUPERVISOR DEPARTMENTAL BOARD DEAN OF FACULTY OF COMPUTER SCIENCE AND ENGINEERING Note: Student must pin this task sheet as the first page of the Master’s Thesis booklet i Acknowledgements I would like to express my sincere gratefulness to my supervisor, Assoc.
Thoại Nam, for his enthusiastic guidance and continuous support during my research. Thanks to his broad knowledge and deep experience, his feedback has greatly increased my knowledge and scientific research skills. The inspiring weekly catchups with him have encouraged me to push myself further and complete such challenging work. I would like to show my all-hearted love for my family, Mom, Dad and my younger brother, for always being by my side along the journey.
Their warmness and cheerfulness significantly boosted my confidence and energy to continue working on the study. There is no time that they did not show me the pride they take in. Special thanks to my love, Nhàn, for the empathy and emotional support she shows me all the time. They are so prestigious and invaluable to me and I heartfully appreciate them.
I also want to thank my supervisors and teammates, Mr. Nam, Huy, Quân, Trân, Mr. Tâm, Mr Thọ and Khôi, in the companies, both previous and current, for providing me the opportunity to balance the time spent on daily work and the research, especially, Mr. Quý and Mr.
Điền to help me with the industrial expertise and experience, Hiếu for suggesting I try out the SHAP, and Quỳnh for helping me organize my thesis and providing proofreading. Finally, I want to thank all my family members and friends who are looking forward to my com- pletion of this study and show me all the love. ii VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY Abstract Faculty of Computer Science and Engineering HO CHI MINH UNIVERSITY OF TECHNOLOGY Master of Science PREDICTING QUALITY OF HOME CARE LIQUID PRODUCTS by Nguyễn Đức Phú The industrial applications for machine learning are the main concern for the research community, with interpretability increasingly gaining attention. We contribute this study to propose an ap- proach for a traditional problem in the industry of chemical engineering, predicting the quality of home care liquid products but only focusing on the explainability of the intelligence system.
By conducting domain-based feature engineering, we ensure the inputs extracted from industrial in- strument time series data are understandable by users. Two different approaches, which are using a transparent architecture like Linear Regression and conducting post hoc analysis with SHAP value for an ensemble model Random Forest Regression, are experimented with in the study. The results are promising in that the ensemble models achieved over 70% accuracy while the influencing features are aligned with domain expertise. In conclusion, the study provides a practical example of deploying an explainable artificial intelligence solution in a traditional industry such as chemical engineering.
iii Tóm tắt Luận văn by Nguyễn Đức Phú Ứng dụng Học máy trong công nghiệp là mối quan tâm chính của cộng đồng nghiên cứu, trong đó tăng cường khả năng diễn giải cho mô hình học máy ngày càng được chú ý. Nghiên cứu này đề xuất một phương pháp cho một vấn đề truyền thống trong ngành công nghiệp hóa chất là dự đoán chất lượng sản phẩm nước vệ sinh nhà cửa, tập trung vào khả năng giải thích của hệ thống thông minh. Bằng cách thực hiện trích xuất đặc trưng mô hình dựa trên kiến thức lĩnh vực chuyên môn, luận văn đảm bảo các dữ liệu đầu vào xử lý từ chuỗi giá trị từ máy móc thiết bị có thể hiểu được bởi người dùng. Nghiên cứu đã thử nghiệm hai phương pháp khác nhau: sử dụng kiến trúc “trong suốt” như Hồi quy tuyến tính và thực hiện phân tích hậu nghiệm bằng giá trị SHAP cho mô hình Hồi quy Rừng ngẫu nhiên.
Kết quả cho thấy các mô hình xây dựng được đạt độ chính xác trên 70%. Các đặc trưng quan trọng tìm được phù hợp với kiến thức chuyên môn. Tổng kết, nghiên cứu cung cấp một ví dụ thực tiễn về việc triển khai giải pháp trí tuệ nhân tạo có thể giải thích được trong một ngành truyền thống như hóa chất. iv Declaration of Authorship I, Nguyễn Đức Phú, declare that this thesis titled, PREDICTING QUALITY OF HOME CARE LIQUID PRODUCTSand the work presented in it are my own.
I confirm that: • This work was done wholly or mainly while in candidature for a research degree at this University. • Where any part of this thesis has previously been submitted for a degree or any other quali- fication at this University or any other institution, this has been clearly stated. • Where I have consulted the published work of others, this is always clearly attributed. • Where I have quoted from the work of others, the source is always given.
With the exception of such quotations, this thesis is entirely my own work. • I have acknowledged all main sources of help. • Where the thesis is based on work done by myself jointly with others, I have made clear exactly what was done by others and what I have contributed myself. Signed: Date: v “It’s not the destination, it’s the journey.” Ralph Waldo Emerson vi Table of contents Acknowledgements i Abstract ii Declaration of Authorship iv 1 Introduction 1 2 Literature Review 3 2.1 Machine learning in Industry .2 Predict the quality of home care product .3 Application of Explainable AI .4 SHapley Additive exPlanation .1 Build batches time windows .2 Weight of the main mixer .3 Temperature of the main mixer .4 Pressure of main mixer circulation system .5 Speed of main mixer circulation pump .6 The amount of chlorinated water .7 The amount of liquid materials .8 The amount of Dehydol .9 The amount of hot water .10 The amount of Plantacare .11 The amount of enzyme .12 The amount of Glycerol .13 The agitator speed of the main mixer .14 The flow of chlorinated water, LAS, and NaOH .15 The amount of TEA .16 The temperature of chlorinated and hot water .1 Baseline model for viscosity .2 Baseline model for pH .3 The features for liquid materials .4 The feature of physical signals .4 SHAP for feature selection and model explanation.
37 4 Result and analytics 42 4.1 Viscosity baseline model .2 pH baseline model .2 Model predictions explanation .3 Models with SHAP-based feature selection. 49 5 Conclusion 54 References 56 Appendices 59 A Experiment with different architecture 59 A. 60 viii List of Figures 3.1 Distribution of products pH .2 Distribution of products viscosity .3 Distribution of batches duration .4 An example of filtered main mixer volume .5 An example of filtered main mixer temperature .6 An example of filtered main mixer pipe pressure of circulation system .7 Comparision two time-series of circulation pump speed .8 The amount of chlorinated water used in a batch .9 More of dosing materials .10 The amount of dehydol used in a batch .11 The amount of Hot water used in a batch .12 The amount of reworked water used in a batch .13 The amount of Plantacare used in a batch .14 The amount of Enzyme used in a batch .15 The amount of Glycerol used in a batch .16 The speed of the main mixer agitator .17 The speed of flushing chlorinated water to the main mixer .18 The amount of TEA .19 The temperature of chlorinated water and hot water .20 An example of detecting a stable time window .21 Distribution of first stable index found .22 Distributions of baseline model features .23 Boxplot of flushed water .24 Boxplot of major liquid material .25 Boxplot features of liquid materials .26 Boxplot features of liquid materials, outlier removed .27 Boxplot physical features .28 Boxplot physical features .29 Boxplot physical features .30 Boxplot physical features for dosing phases .1 Predicted value and truth value for baseline viscosity models .2 Predicted value and truth value for baseline ph models .3 SHAP Summary plot for baseline viscosity .4 SHAP Summary plot for baseline viscosity, accurately predicted points .5 SHAP Interaction values of Press and Temp .6 SHAP Summary plot for baseline pH .7 SHAP Summary plot for baseline pH, accurately predicted points .8 Feature importances over iterations, above 0.9 Feature importances over iterations, above 50 only .1 Comparison in performance for pH .2 Comparision in performance for viscosity. 60 x List of Tables 3.1 Descriptive analysis for pH .2 Descriptive analysis for Vis .2 Descriptive analysis for Vis .3 Data type of time series batch_name .4 Descriptive analysis of batches duration .4 Descriptive analysis of batches duration .5 Inputs of Viscosity baseline model .6 Descriptive analysis for detecting stable indexes with threshold 5% .7 Summary of the features for the baseline model .8 Inputs of pH baseline model .9 The amount of water found in batches (raw) .10 Summary amount of liquid materials .11 Summary duration of liquid materials .11 Summary duration of liquid materials .1 Performance of Viscosity baseline models .2 Performance of pH baseline models .3 Performance of Random Forest pH models .4 Feature importance of the final pH model .5 Performance of Random Forest pH models .5 Performance of Random Forest pH models .6 Feature importance of the final viscosity model .1 Selected features of the final iteration for pH models .2 Selected features of the final iteration for viscosity models.
61 1 Chapter 1 Introduction In recent years, the raising of machine learning and deep learning has been kept at the highest level. The same is true for the attention they attracted from the research community. The introduction of Large Language Models is still one of the hottest topics around the world. However, alongside the influence deep learning has brought, its application in manufacturing and heavy industries is still questionable.