VIETNAM NATIONAL UNIVERSITY, HANOI INTERNATIONAL SCHOOL GRADUATION PROJECT FEATURE ENGINEERING AND MACHINE LEARNING MODELS FOR STUDENTS’ LEARNING TIME PREDICTION IN ENGLISH COURSE AT VNU-IS Student’s name Nguyễn Thùy Anh Hanoi - Year 2023 VIETNAM NATIONAL UNIVERSITY, HANOI INTERNATIONAL SCHOOL GRADUATION PROJECT FEATURE ENGINEERING AND MACHINE LEARNING MODELS FOR STUDENTS’ LEARNING TIME PREDICTION IN ENGLISH COURSE AT VNU-IS SUPERVISOR: Dr. Nguyen Quang Thuan STUDENT: Nguyen Thuy Anh STUDENT ID: 20070896 COHORT: SUBJECT CODE: MAJOR: Business Data Analysis Hanoi - Year 2023 ACKNOWLEDGEMENTS I would like to express my gratitude to my thesis advisor Dr. Nguyen Quang Thuan for his enthusiastic and thoughtful help throughout the research process of this graduation thesis. Thanks to that guidance, I received a lot of helpful knowledge and became more self-employed with my chosen major.
I would like to thank the Faculty of Applied Science for providing resources and the learning environment to help carry out this thesis research. I am greatly appreciative for the knowledge that I received from all the teachers during my undergraduate learning under International School. I would like to thank my family and friends for their help and support during this long journey. The graduation thesis is a long and difficult work; without anyone's companionship, I would not have been able to go through this journey completely.
My gratitude to you all for being one part of my academic milestone. DECLARATION I, Nguyen Thuy Anh, hereby declare that this thesis, titled “Feature engineering and machine learning models for students’ learning time prediction in English course at VNU-IS”, submitted partially to fulfill the requirements for the degree of Business Data Analytics at International School, Vietnam National University, Hanoi, is my original work. This thesis's ideas, data, and information are sourced and referenced appropriately. This thesis represents my independent thinking and understanding of the subject matter without falsifying, fabricating, or manipulating data or results.
All images, figures, tables, and other materials used in this thesis are either created by myself or sourced from the public domain or with proper permissions and acknowledgments from the original creators. I understand that any violation of the above declarations may result in consequences as per the academic policies and integrity guidelines of International School, Vietnam National University, Hanoi. Nguyen Thuy Anh 2023 List of acronyms Acronym Meaning VNU-IS International School – Vietnam National University DOB Date of birth POB Place of birth EDA Explore data analysis RF Random Forest SVM Support Vector Machine KNN K-Nearest Neighbors PCA Principal Component Analysis List of tables Table no.1 VNU-IS dataset 1.2 Data Mining processes 1.1 Original data after cleaning 3.2 Datasets used for model training List of figures Figure no.1 VNU-IS pre-English course 2.1 Feature selection techniques 2.1 Data labelling example 3.2 Starting levels with Birth Region 3.3 Major and Major type number 3.4 Starting levels with subject groups 3.5 Data features distribution 3.6 English and no-English 3.7 Uni exam and placement test scores 3.8 Intake’s placement test scores 3.9 Output classes student number 3.10 5 classes and 3 classes 3.11 Random Forest 5 and 3 classes 3.12 Feature engineering and RF 3.13 Feature engineering and SVM 3.14 Feature engineering and KNN 3.16 PCA in RF 3.17 PCA in SVM 3.18 PCA scatter plot with LE TABLE OF CONTENTS CHAPTER 1. PROBLEM OF PREDICTING STUDENTS’ LEARNING TIME IN PRE-ENGLISH COURSE.
Introduction to data mining. FEATURE ENGINEERING METHODS AND MACHINE LEARNING MODELS. Categorical data encoder. Support Vector Machine.
Implementation tools: Python language using Google Colab. New features – grouping data. Exploratory data analysis (EDA). Model construction and evaluation.
New output label. Categorical feature encode. Principal Components Analysis (PCA). Train-test split.
Machine learning models implementation. Compare between two types of output labels. Models and data engineering result. PCA or no PCA?.
41 ABSTRACT Data mining brings insight and benefits to many aspects of life, including education. Analyzing trends, special signs, and characteristics of learning and teaching and then performing predictions about student performance has brought many benefits to the education system. As a university with special training programs such as training in English, dual degree education, or modern data and technology programs, International School – Vietnam National University is on the rise and will benefit from the application of data mining. In this thesis, the pre-English course of the International School is the problem used for analysis.
The process has gained many insights and positive predictive results by properly engineering the data, performing exploratory data analysis, and building predictive models of students’ pre-English course study time. International School can consider those as the reference for future program adjustment. PROBLEM OF PREDICTING STUDENTS’ LEARNING TIME IN PRE-ENGLISH COURSE This chapter will describe the problem of predicting students’ learning time in the International School - Vietnam National University Hanoi, and the application of data mining in education for problem-solving. Problem description The International School belongs to Vietnam National University Hanoi, with training programs in English for all majors.
The International School is at the forefront of new subjects in the fields of information technology and data. Also, it organizes double degree programs for students, creating opportunities to connect students with dynamic foreign environments. In recent years, the International School has received a large number of students in each enrollment period, increasing the teaching scale of the school rapidly, which results in many opportunities as well as challenges. To respond to the new situation, the International School adapts to the modern helping tool, such as the online learning software LMS tested on the preparatory program or means to help analyze strengths and weaknesses of the current training program.
The application of information technology is helping the education industry a lot, in which education data analysis has many practical proofs to help universities solve many significant problems. Applying it to learning and teaching data analysis will give the school a stronger foundation to continue taking care of students while welcoming students of the following courses more thoughtfully. International School is very diverse in student nature. To conduct a good analysis of the school students, the school needs to start the long journey of building the system now.
It will begin with using the existing tools and material (school labs, the internet system, and student data) for analyzing small problems such as the preparatory English program. This program is the starting point for many students when enrolling, therefore, adjustment based on the actual situation of new students is necessary. By analyzing this problem, the return insight can help the International School adapt to the new student 2 course each year, creating a more suitable program for the new generation's English ability. As an International School, its curriculum is taught in English, broadening the scope of the curriculum and making it relevant to international students as well.
Therefore, the International School will require students to have English proficiency before attending the main program, a B2 in English entry is required, and all types of certificates such as IELTS, TOEIC or TOEFL, etc. If not available at enrollment, students will have to participate in the preparatory training program that the school organizes to help students with the exam. An entrance test with English listening, reading, and writing skills will be held to assign students to classes of different level: Foundation, 1, 2, 3 and 4. The pre-English course is described in the following figure: Figure 1.
VNU-IS pre-English course The school needs to know the real effects of the course on students’ English ability, to make adjustment if needed, which means every aspect of the pre-English course needed to be put into consideration, including the placement test importance, level division, student differentiation, etc. To perform this problem meaningfully, students’ learning learning time will be the main factor for analysis and mining, the problem is defined as “Students’ Learning Time Prediction”, using data that are related to students’ initial English ability and the pre-English course information to predict the learning time of student in the course. The Academic department has listed out the related features and provides the dataset of student school year 2020 for research purposes. The dataset categories and features are listed as follows: 3 Table 1.
VNU-IS student dataset Category No Data features Features description Data type Student 1 Date of birth Date personal (dd/mm/yyyy) information 2 Gender Categorical Demographic 3 Citizen Identification Categorical 4 Major Categorical 5 Phone Numerical 6 Place of birth Categorical Certificate 7 Certificate exam date Date information (dd/mm/yyyy) 8 Certificate date Date (dd/mm/yyyy) 9 Certificate status Categorical 10 Enroll certificate Categorical 11 Score of enroll Numerical certificate 12 Study status Categorical 13 Note Categorical Pre-English 14 Writing score /30 Placement test score Numerical course 15 Reading score /35 Numerical information 16 Listening score /30 Numerical 17- Intake 1 level, class, Intake level, class Intake level: 19 score arrangement and final Categorical 20- Intake 2 level, class, score of intakes. Intake class: 22 score Categorical 23- Intake 3 level, class, Intake score: 25 score Numerical 4 26- Intake 4 level, class, 28 score 29- Intake 5 level, class, 31 score Enrollment 32 Group subject code Code of subject group Categorical information for university enrollment exam 33 VNU-IS order International school Numerical order in the student’s university wish list 34 Enroll score Combination of the 3 Numerical subjects score 35 Subject 1 First subject in subject Categorical group 36 Subject 1 score Numerical 37 Subject 2 Second subject in Categorical subject group 38 Subject 2 score Numerical 39 Subject 3 Third subject in Categorical subject group 40 Subject 3 score Numerical 41 Admission method Student admission Categorical method: direct entry, university exam Highschool 41 Conduct Categorical performance 42 Performance Categorical information 43 Grade 12 average Numerical score 5 The dataset is divided into 5 categories: students demographic, certificate information, pre-English course class arrangement, enrollment information and high school performance. It is shown that the features are pretty selective and have some correlation with students’ initial English ability. In addition, some features that are assumed to relate to the ability to learn English, such as place of birth, subject group, and gender, are also included in the data set to be able to analyze and solve these conjectures.
For the course information, there are five intakes, equivalent to 10 months maximum from the enrollment for students to submit the certificate. The timeline for five intakes is shown below: Figure 1. Intakes' timeline After thoroughly observing the dataset, there are several questions are made: - What are the factors related to students' English learning? - Are stereotypes about the place of birth, gender, and majors related to English ability true? - What features will assess students' English learning ability and speed in the above data set? - Can the school differentiate students’ learning ability and speed using the given features? For the questions to be answered, a process of data mining needs to be done. The concept of data mining, specifically educational data mining, and related work will be mentioned in the next section for the reference of the thesis.
Introduction to data mining For the last 20 years, data mining has become one of the most highlighted topics in the data science community. Data mining is the process of “digging” into the data to find patterns and relationships and any possible trends that can happen in the near future. The return results from data mining can provide insights for development in many fields like business [1], health inspection [2], education [3], etc., and help with decision making. The process of data mining is not strictly defined, but usually proceeds in the orders of data collecting, data preprocessing, data transformation, pattern analyzing, model building and evaluation.
The evaluation process is based not only on the assessment of the data analyst but also the criteria of the organization. Data Mining processes Even though data mining is a powerful tool, it still requires data preprocessing, which results in a clean, concise and understandable dataset to utilize data mining maximum capability.