VIET NAM NATIONAL UNIVERSITY HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY ADVANCED PROGRAM IN INFORMATION SYSTEMS LAM HA TUAN CANH THESIS GRADUATION APPLYING PREDICTION MODELS TO FORECAST REAL ESTATE PRICES BANCHELOR OF ENGINEERING IN INFORMATION SYSTEMS HO CHI MINH CITY, 2021 LAM HA TUAN CANH -15520056 THESIS GRADUATION APPLYING PREDICTION MODELS TO FORECAST REAL ESTATE PRICES IN HO CHI MINH CITY BANCHELOR OF ENGINEERING IN INFORMATION SYSTEMS THESIS ADVISOR Dr. CAO THI NHAN ASSESSMENT COMMITTEE The Assessment Committee is established under the Decision. by Rector of the University of Information Technology. - Member ACKNOWLEDGEMENTS First of all, I would like to show my appreciation Dr.
Cao Thi Nhan for being my thesis advisor not only during the time we work on graduation thesis but since I joined the University of Information Technology as a consultant. With patience, motivation, and immense knowledge, she helped us to keep track of the direction of the research and gave us lots of advice to have the thesis completed. I also express our sincere thanks to Dr. Do Trong Hop for the very careful review of my thesis, and for all the insightful comments, suggestions, and corrections.
TABLE OF CONTENTS css TABLE OF CONTENT S.cccsssssssssssssesesecssscccsesessescseseseseseesesesessncseseseeseeeseseeeaees 2 LIST OF FIGURES .ccceccssssssssesesessseseessssesseeesessssssesesesasssessesnsssscsesessssssssesnsesaeseens 4 LIST OF TABLES .-- << 5-5-5555 Es+SS4 3 3E1ES.E13 3 8031104010001811 010g 6 LIST OF ABBREYVIA TIONS. 5c HS 003030108013001080808403010808004040401004040101010404010896 8 Chapter 1 : INTRODUCTION .2 Objective and scope. nh xe etseereresese 1. Thesis structure Chapter 2 : DATA SET AND METHODOLOGY 2.2 Data Collection and Data Set Generating 2.
¿St sseeeeeerrrrierriee LT 2.3 Exploratorry Data Analysis ResuÏL.4 Data pre-processing oo.3 Regression Model and Evaluation Metrics Used.1 Linear Regression Models — Stochastic Dual Coordinate “Ascent G000.2 Decision Tree — Fast Forest Regression .3 Gradient Boosting Regression — LightGBM and Fast Tree G000.-- - ¿+ ¿5S Sky 31 Chapter 3 : IMPLEMBENTA TION.net Machine Learning Model Builder. 1xx 5 Chapter 4 : CONCLUSIONS. 5-5-5 S*HỲHnHxgg H11 ung rsee 60 4.2 Limitations and challenges. Ăn 101 010101403030108080040401010004040101000196 LIST OF FIGURES css Figure 2-1.
Data Preparation Process. Process Flow of Prediction Model [12].e- LỘ Figure 2-3 Website UI [3] .- - - + + St kg g1 1tr gret 16 Figure 2-4 A raw data set example [3]. Figure 2-5 Processed Data Set. SellPrice Scatterplot in HCMC.
SellPrice Distribution of HCMC. Average price for types of real ©SA(€. Relationship between SellPrice and Legal Document. Average Price in each District .-- ¿5-55-5255 c+ccseseeererseeeeeee 27, Figure 3-1 Options to Add ML Model.
Figure 3-2 Scenarios Choosing UL Figure 3-3 Information of training environment Figure 3-4 Data PT€VI€W. HH Hi OD Figure 3-5 Data types settings of the variables. Set time for training da(a. Recommended time for training data .----¿-‹--5--<e-<+<--c---- OD) Figure 3-8 I“ experiment’s R-squared.
Figure 3-9 1* experiment sample no. Figure 3-10 Actual price of alike property of I* experiment sample no. Figure 3-11 Housing prices Reference of Lac Long Quan Street from Mogi.vn Figure 3-12 1“ experiment sample 0. -- - + + e++£+++x££keEkeErkrkerkekkrrerkee Figure 3-13 Actual alike property of 1 experiment sample no.2 [13] Figure 3-14 2"4 experiment’s R-squared.
Figure 3-15 2" experiment sample. Figure 3-16 Actual alike property of 2"4 experiment sample [13].---- 49 Figure 3-17 Housing prices Reference of Linh Dong from Mogi. 49 Figure 3-18 3 experiment’s R-squared .cscsssesssssssseesseesstecssesssseessecsseesssecsseesseeessees SO) Figure 3-19 3" experiment sample no. 2 Figure 3-20 Actual alike property of 3 experiment sample no.
53 Figure 3-21 Housing prices Reference of My Hue Street from Mogi. 53 Figure 3-22 3TM experiment sample 10.ssssccssssssessssecsseesssecsseccssecssecssscsssecsseesseeessees 4 Figure 3-23 UI of the website Figure 3-24 Filters available Figure 3-25. UI when choosing a property. “ Figure 3-26 Displayed reSuÏtS.-- -- 55-555 S++£sketkerrkrrkerkerrrerkrerrerrercev OD Figure 3-27.
Actual average price of the property on location.- 2Ø LIST OF TABLES css Tab le 2-1 Raw Data Summary. cecceesesecseessseeeeseeeeenessseesessseensaeeeseeesaeseeseeeessseeeaees 2 Tal le 2-2 Data Description .ccccecscsceseescesesesesseseseesesesssssessssesssssssssesssesssssssessesees 2 Tab le 2-3 Processed Data SUMMALY.ce sce sseesseetesesesessessseeneneassesessseeteneatenenesees 20) Tal le 3-1 1“ experiment dataset. Tab le 3-2 2"4 experiment dataset Tab! le 3-3 3 experiment dataset. Tab le 3-4 1“ Experimental results.
AL Tab! le 3-5 Testing examples of 1S experiment .-- ©5555 +xes++£vzverxerererxrre 42 Tab! le 3-6 2" Experimental Metrics .ccscsessssesssseesssseecssseesssneessnneesssneessnneeesnneeesnees 46 Tab! e 3-7 Testing Examples of 2° experiment.:-cccccsccecrerrrrrrrrrrrrrrrreere 47, Tab! e 3-8 3 Experimental results. Tab! e 3-9 Testing Examples of 3" experiment. LIST OF ABBREVIATIONS EDA Exploratory Data Analysis LASSO Least absolute shrinkage and selection operator VAR Vector autoregressive ADL Autoregressive distributed lag XGBoost Extreme Gradient Boosting SVR Support Vector Regression SGD Stochastic Gradient Descent GBR Gradient Boosting Regression SDCA Stochastic Dual Coordinate Ascent HCMC Ho Chi Minh City MAE Mean absolute error MSE Mean squared error RMSD Root-mean-square deviation FF Fast Forest UI User Interface VND Vietnam Dong RAM Random Access Memory SSD Solid-state Drive ABSTRACT Different models used in house price forecasting are tested on their prediction accuracy. Using data from detailed house price indices to Ho Chi Minh City of Viet Nam in the third and fourth quarter of 2021.
Some regression techniques such as Stochastic Dual Coordinate Ascent (SDCA) method, Fast Forest for Decision Tree models and Gradient Boosting algorithms, namely Fast Tree Regression and Light Gradient Boosting Machine are selected to forecast house price index changes. Such models are used to build a predictive model, and to pick the best performing model by performing a comparative analysis on the predictive errors obtained between these models. The data set used for this report was downloaded from the website www. The data set consisted of nearly 2000 observations and 18 variables.
The target variable from the given data set was Price. The results of the experiments are illustrated through a website built with .Net Core and Angular to obtain a demonstration of the used algorithms’ performance. KEYWORDS: House Price Prediction, Linear Regression, Gradient Boosting Regression, Decision Tree Chapter 1: INTRODUCTION 1.1 Background Vietnam is developing into a rapidly growing and prosperous real estate market in Southeast Asia. It is considered one of the hotspots of the most developed real estate market in Asia, with a growing economy, some laws have made it easier for foreigners to buy the property.
As of 2017, the increase in the number of investors, including national and international players in the respective sub-markets, has led to the development of new housing companies, green buildings, etc. of several mega-projects in major cities are fundamental to the growth of the residential real estate market, both in the basic and in the luxury segment. In the context that there has been a significant increase in the number of people demanding of owning a house or land so the real estate market in Viet Nam has become more and more appealing to the investors along with the price countinuously fluctuates. This is also synonymous with the confusion of whether they have purchased a property at a proper price among the customers and a massive number of scammers taking advantage of this situation is inevitable.
So this is the major consideration why should we need predictive models. In short, predictive modeling is an applied mathematics technique exploitation machine learning and data processing to predict and forecast possible future outcomes with the help of historical and existing information. It works by analyzing current and historical data and sticking out what it learns on a model generated to forecast likely outcomes. In this thesis, Forcast Model is used because of its popularity in working with numerical values based on training data [1].
The data used in this thesis is available from the websites laydulieu.com under the link of [3] by Nguyễn Đức Nam and mainly focus on Ho Chi Minh City territory. In particular, 4 typical areas of the city are investigated, namely District 1 — the city’s heart, Thu Duc District — the newly emerged as the most potential area, Tan Binh District — the place for laborers from other provinces to come and settle down and Hoc Mon District — the outskirts of this metropolis where the infrastructure is still not worth-concerned. The full list of data variables is given in Section 2. There are various considerations influencing the price of properties.
According to [6],[7], price of real estate is influenced by several factors like: e Property-related factors ¢ Locational Factors ¢ Environmental Factors The purpose of this thesis is first to examine the influence of various variables on the real esate prices in Ho Chi Minh City by using EDA. Secondly, researching the linear regression algorithm for their theories and operations to indicate the highest predictor through experiments is the main target. Three models based on Linear Regression, Decision Tree and Gradient Boosting respectively are proposed and demonstrated through a website.2 Objective and scope 1.1 Objectives e Understand the implementation of business data analysis and machine learning on providing results. ¢ Covering real estates in 4 typical areas in HCMC territory, which are: o District 1: The downtown o Hoc Mon District: The suburb o Tan Binh District: The stable area o Thu Duc District: The developing area e There will be clear explanations for the reasons why these 4 areas are selected in Section 2.2 Scope Using Linear Regression Algorithms and Decision Tree Models for the value predictor of real estate prices.
Besides, R-squared is the main mectric used for a evaluation in terms of the efficiency. A demonstration website is built to illustrate the result of the experiments and performance of the used algorithms.3 Thesis structure This thesis is divided into five chapters, as follows: e Chapter 1: Introduction e Chapter 2: Data set and Methodology e Chapter 3: Implementation e Chapter 4: Conclusions 11 Chapter 2: DATA SET AND METHODOLOGY 2.1 Related works In recent decades, there has been a demand to extend the house price prediction services which help investors and settlers to take a correct decision. This section describes the previous work done by several researchers in the selected domain of housing price prediction. Following are the contributions of various researcher done in this domain: In 2016, Martijn Duijster [4] used ARIMA/ADL/VAR to forecast the Dutch house index changes.
The experiemental analysis was based on the the period of 1995 — 2016 Netherlands house price data. This paper shows that ADL had the best performance among the algorithms used and claims that 1-to-6-quarter forward prediction is available but because of the out-dated data and a broad spectrum are the major limitations. In 2019, Nebojša Dubošanini, Jan Eric Biihlmann2 and Pauline Offeringa [2] applied Linear Regression, RIDGE Regression, LASSO Regression, Random Forrest on predicting the index of dataset of Melbourne, Australia. This project indicated that the Decision Tree model can be simple and still had the best performance compared to RIDGE and LASSO Regression.
Moreover, this model also provide an explicit look at the scheme and how the target variable was computed. However, the drawbacks is the inability of above threshold prediction and unspecified data time so it would be difficult to get a precise forecast because of the fluctuation in prices of this area. In 2020, Yichen Zhou [8] put Linear Regression, LASSO Regression, Random Forrest and XGBoost into practice to predict house price in Ames, Iowa, USA with the 79-variable data of 5-year period from 2005 to 2010. The experiemtal results expressed high accuracy of XGBoost model’s predicted values in comparison with actual figures, with more than 94% accurate.
Nevertheless, the considerable variables data would take loads of time to do the EDA for feature extraction and the recency is also worth-concerned. 12 Even in 2021, LASSOLARS Regression, Bayesian Ridge Regression, SVR, SGD and GBR are still used for price prediction of house dataset of Islamabad — Capital of Parkistan [9] by Imran , Umar Zaman ,Muhammad Wagar and Atif Zaman. The results show that SVR performs best than the rest of the machine learning algorithms.