MINISTRY OF EDUCATION AND TRAINING HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY AND EDUCATION GRADUATION THESIS MAJOR: ELECTRONICS AND COMMUNICATION ENGINEERING PREDICTION OF RISK AND RETURN BY USING MACHINE LEARNING INSTRUCTOR: PHẠM NGỌC SƠN, PhD. STUDENT: LÂM MINH NHẬT SKL011197 Ho Chi Minh City, July 2023 HO CHI MINH CITY UNIVERSITY OF TECHNOLOGY AND EDUCATION FACULTY FOR HIGH QUALITY TRAINING DEPARTMENT OF COMPUTER AND COMMUNICATIONS GRADUATION PROJECT PREDICTION OF RISK AND RETURN BY USING MACHINE LEARNING Student’s name: LÂM MINH NHẬT Student ID: 19161037 Major: ELECTRONICS AND COMMUNICATION ENGINEERING Advisor: PHẠM NGỌC SƠN, PhD. Ho Chi Minh City, July 2023 THE SOCIALIST REPUBLIC OF VIETNAM Independence – Freedom– Happiness -------- Ho Chi Minh City, July 5th, 2023 GRADUATION PROJECT ASSIGNMENT Student name: LÂM MINH NHẬT Student ID: 19161037 Major: ELECTRONICS AND Class: 19161CLA COMMUNICATION ENGINEERING Advisor: PHẠM NGỌC SƠN, PhD. Phone number: Date of assignment: June 22th 2023 Date of submission: June 23th 2023 1.
Project title: RISK AND RETURN BY USING MACHINE LEARNING 2. Initial materials provided by the advisor: ___________________________________ 3. Content of the project: _________________________________________________ 4. Final product: ________________________________________________________ CHAIR OF THE PROGRAM ADVISOR (Sign with full name) (Sign with full name) THE SOCIALIST REPUBLIC OF VIETNAM Independence – Freedom– Happiness -------- Ho Chi Minh City, July , 2023 ADVISOR’S EVALUATION SHEET Student name:.
Content of the project:. Approval for oral defense? (Approved or denied). Ho Chi Minh City, (month day, year) ADVISOR (Sign with full name) THE SOCIALIST REPUBLIC OF VIETNAM Independence – Freedom– Happiness -------- Ho Chi Minh City, July, 2023 PRE-DEFENSE EVALUATION SHEET Student name:. Name of Reviewer:.
Content and workload of the project. Approval for oral defense? (Approved or denied). Reviewer questions for project valuation .) Ho Chi Minh City, (month day, year) REVIEWER (Sign with full name) Faculty for High Quality Training – HCMC University of Technology and Education PREAMBLE During the implementation of the graduation project, our team has received a lot of help, suggestions and indicators from teachers and friends. To complete the project “RISK AND RETURN BY USING MACHINE LEARNING.” We sincerely thank PhD.
Phạm Ngọc Sơn - Lecturer of the Department of Computer Engineering - Telecommunications, Faculty of Electrical - Electronics, University of Technology and Education of Ho Chi Minh City. With dedicated guidance, guidance, facilitating and supporting the team to successfully complete this project. Individual would also like to thank the authors of the reference sources who helped the group to have more knowledge and choices in the process of implementing the topic. Although the group has tried to complete this topic in the most complete way, certain errors in research work, practical approach, as well as limitations in knowledge and time cannot be avoided implementation time.
Looking forward to receiving your comments so that the group can supplement and correct the topic to be more complete. Faculty for High Quality Training – HCMC University of Technology and Education ABBREVIATION Short writing Meaning NaN Not a Number 2D Two-direction I/O Input / Output LSTMs Long-Short Term Memory networks RNN Recurrent Neural Network Stdv Standard deviation Rmse Root-mean-square deviation S&P500 Standard & Poor's 500 Index CAGR Compound Annual Growth Rate rfr Risk-free-rate CHP Central Hydropower JSC SAB Saigon Beer Alcohol Beverage Corp FPT FPT Corp VND Vietnam dong Faculty for High Quality Training – HCMC University of Technology and Education MATHEMATICAL SYMBOL Symbol Meaning Tanh The function Tanh is the ratio of Sinh and Cosh. – third layer state Xt Data input ht Data in epochs Ct Cell state 𝜎 Sigmoid ft First layer state it Second layer state ot Fourth layer state rfr Risk-free-rate stdv Standard deviation PercentGrowth Daily growth percentage 𝑆ℎ𝑎𝑟𝑝𝑒𝑅𝑎𝑡𝑖𝑜 The amount of return received per unit of risk Faculty for High Quality Training – HCMC University of Technology and Education Table of Contents CHAPTER 1: OVERVIEW OF PROJECT. OBJECTIVES OF THE PROJECT.
LIMITATION OF THE PROJECT. 2 CHAPTER 2: THEORETICAL BASIS. PYTHON PROGRAMING LANGUAGE. THE CORE IDEA OF LSTM.
SHARPE’S RATIO CONCEPT. REVIEW ANOTHER PROJECT .12 CHAPTER 3: IMPLEMENTATION PROCESS. MAIN FLOWCHART AND SIMULATION PARAMETERS.23 CHAPTER 4: RESULT OF SIMULATION ANALYSIS AND ASSESSMENT .30 CHAPTER 5: CONCLUSION AND DEVELOPMENT. DEVELOPMENT DIRECTION OF THE PROJECT.
34 Faculty for High Quality Training – HCMC University of Technology and Education List of figure Figure 1) The repeating module in a standard RNN contains a single layer. 7 Figure 2) The repeating module in an LSTM contains four interacting layers. 7 Figure 3) The cell state runs straight down the entire chain with only some minor linear interactions. 8 Figure 4) They are combined by a sigmoid lattice layer and a multiplication.
8 Figure 5) Decision of what information is going to throw away from the cell state. 9 Figure 6) Decision of what new information is going to store in the cell state. 9 Figure 7) Updating the old cell state. 10 Figure 8) Decision of what is going to output.
10 Figure 9) William Forsyth Sharpe. 11 Figure 10) Summarize daily prices for Amazon and Facebook. 12 Figure 11) Plot daily prices for Amazon and Facebook. 12 Figure 12) Summarize daily values for the S&P 500.
13 Figure 13) Plot daily values for the S&P 500. 13 Figure 14) Visualize the daily return of Amazon and Facebook. 14 Figure 15) Visualize the daily return of S&P 500. 14 Figure 16) Calculating Excess Returns for Amazon and Facebook vs S&P 500.
15 Figure 17) The Average Difference in Daily Returns Stocks vs S&P 500. 15 Figure 18) Standard Deviation of the Return Difference. 16 Figure 19) Comparation the Sharpe ratio between Amzazon and Facebook. 16 Figure 20) Model system.
23 Figure 22) Development chart of CHP. 24 Figure 23) The predicted price compares with valid price. 24 Figure 24) Comparasion of Sharpe Ratio. 26 Figure 25) Development chart of SAB.
26 Figure 26) The predicted price compares with valid price. 27 Figure 27) Comparasion of Sharpe Ratio. 28 Figure 28) Development chart of FPT. 28 Figure 29) The predicted price compares with valid price.
29 Figure 30) Comparasion of Sharpe Ratio. 30 Figure 31) CHP’s Open Price in 23-06-09. 31 Figure 32) SAB’s Open Price in 23-06-09. 31 Figure 33) FPT’s Open Price in 23-06-09.
32 Faculty for High Quality Training – HCMC University of Technology and Education Chapter 1: Overview of project 1. Introduction People often say that "The stock market is the barometer of the economy", or in other words, the stock market is a future indicator of the economic movement. In recent years, securities are an effective investment and capital allocation channel for many people besides bank accounts; especially in the context that bank interest rates tend to decrease gradually, the securities investment channel becomes a bright spot. From that, it can be seen that the Stock Market is becoming an important aspect of the market economy and the investment choice of many people[1].
The application of technology can help investors analyze stock market indices and indicators to make accurate and timely investment decisions. By using “Machine Learning” is the "door" for investors to understand the methods and methods of applying machine learning achievements to understand about stocks – about a field that is constantly changing every day. Focusing on technical application, technical application in quantitative finance will support Investors to have a realistic, intuitive and easy-to-understand view of the stock market. Objectives of the project Based on Sharpe’s ratio formula; Using machine learning model to simulate predicted stock values through the GoogleColab compilation environment.
From the predicted values of model’s result, we will draw conclusion, analyze the parameters. Moreover, evaluating the comparation between the “Predicted Sharpe’s ratio”- taken from parameters in the model with “Valid Sharpe’s ratio”- taken from the data collected in reality. Finally, giving the advices and observe the stock evolutions that they grow similar with the trending. Limitation of the project This project focuses on building machine learning models to simulate and analyze that will predict the collection of data through regression methods.
In addition, the above results are only simulations and can be applied to actual models. 1 Faculty for High Quality Training – HCMC University of Technology and Education 1. Project’s layout Chapter 1: Overview- Learn about the research situation, objectives, limitations and project layout Chapter 2: Theoretical foundations. An overview of the theories used in the project.
Chapter 3: Implementation content. Provide a system model and analyze the result chart and parameters. Chapter 4: Simulation analysis and evaluation. Provide simulation scenarios, flowcharts, simulation results and comment on results obtained.
Chapter 5: Conclusion and development direction. Provide conclusions and directions for the development of the topic. 2 Faculty for High Quality Training – HCMC University of Technology and Education Chapter 2: Theoretical basis 2. Python programing language Python is a high-level general-purpose programming language that was developed by Guido van Rossum and originally made available in 1991.
Python was created with the significant benefit of being simple to read, understand, and remember. Python, which is frequently used in the creation of artificial intelligence, has a very brilliant look, a clear structure, is handy for novices, and is simple to master. The structure of Python also enables users to write code with few keystrokes. Python consistently ranks as one of the most popular programming languages[2].
Numpy The fundamental Python module for scientific computing is called NumPy. A multidimensional array object, various derived objects (like masked arrays and matrices), and a variety of routines for quick operations on arrays are provided by this Python library. These operations include discrete Fourier transforms, basic linear algebra, basic statistical operations, random simulation, and much more. The ndarray object is the base of the NumPy package.
This contains homogenous n-dimensional arrays of data kinds, with many operations carried out in compiled code for speed. NumPy arrays and regular Python sequences have a number of significant distinctions.: _ Unlike Python lists (which can grow dynamically), NumPy arrays have a fixed size when they are created. An ndarray's size change results in the creation of a new array and the deletion of the old one. _ A NumPy array's elements must all be the same data type in order to share the same amount of memory.
The exception: arrays of (Python, including NumPy) objects are possible, allowing for arrays with various element sizes. _NumPy arrays make it easier to do complex mathematical and other operations on enormous amounts of data. The majority of the time, these actions can 3 Faculty for High Quality Training – HCMC University of Technology and Education be carried out more quickly and with less code than is achievable when utilizing Python's built-in sequences. _NumPy arrays are used by an increasing number of Python-based scientific and mathematical tools; while they normally allow Python-sequence input, they convert it to NumPy arrays before processing and frequently output NumPy arrays as well.
In other words, merely being able to utilize Python's built-in sequence types is not enough to effectively use the majority (perhaps even all) of the scientific and mathematical Python-based applications available today. Pandas Pandas is an open-source library designed primarily for working quickly and logically with relational or labeled data. It offers a range of data structures and procedures for working with time series and numerical data. The NumPy library serves as the foundation for this library.
Pandas is quick and provides users with excellent performance & productivity. Advantages: - Quick and effective data manipulation and analysis. - It is possible to load data from different file objects. - Simple handling of missing data in both floating point and non-floating point data (expressed as NaN).
- Size mutability: columns in DataFrame and higher-dimensional objects can be added and removed. - Merging and connecting data sets. - Flexible data set reshaping and pivoting - Time-series functionality is provided.