VIETNAM NATIONAL UNVIERSITY HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY COMPUTER SCIENCE FACULTY HO SY TUYEN UNDERGRADUATE DISSERTATION CARTOONIZING PORTRAIT IMAGES USING GENERATIVE ADVERSARIAL NETWOK BACHELOR OF SCIENCE IN COMPUTER SCIENCE HONORS PROGRAM HO CHI MINH CITY, 2021 VIETNAM NATIONAL UNVIERSITY HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY COMPUTER SCIENCE FACULTY HO SY TUYEN- 17521244 UNDERGRADUATE DISSERTATION CARTOONIZING PORTRAIT IMAGES USING GENERATIVE ADVERSARIAL NETWOK BACHELOR OF SCIENCE IN COMPUTER SCIENCE HONORS PROGRAM SUPERVISOR Dr. NGUYEN VINH TIEP HO CHI MINH CITY, 2021 DANH SÁCH HỌI ĐÒNG BẢO VỆ KHÓA LUẬN Hội đồng chấm khóa luận tốt nghiệp, thành lập theo Quyết định số. của Hiệu trưởng Trường Đại học Công nghệ Thông tin. ACKNOWLEDGEMENT This dissertation would not be done without the great support and assis- tance from a lot of people.
First of all, I would like to express gratitude to my supervisor, Dr. Nguyen Vinh Tiep, who not only taught me fundamental knowledge but also gave me valuable guidance in intensive research. Next, I would like to acknowledge Nguyen Thanh Danh, Phan Nguyen, Nguyen Vu Anh Khoa, Ngo Huu Manh Khanh, and other members from Multimedia Lab (University of Information and Technology). They pro- vided stimulating discussions, where my idea had been improved day by day.
I would also like to thank my colleagues Huu-Trong Le, Vinh-Loi Ly, Xuan- Phung Pham, and other members from Viettel Cyberspace Center (Viettel Group) for their patient support and all the opportunities that I was given to complete my dissertation. Last but not least, I would like to say thank my family for considering my education and being always with me in my tough time. ABSTRACT Besides doing traditional tasks such as classification, object detection, seg- mentation, and so on, computer vision’s techniques are more creative. The style transfer problem gives the machine the ability to generate an artistic image from a realistic one.
More specifically, since cartoon is a popular art form that widely presents in our daily life, the cartoon style transfer is also well-studied recently. However, despite the great progress in im- age stylization, few methods focus on cartoonizing the image of a human face, which is very challenging due to the complicated requirements of a cartoon face, such as smooth color, clear edges, unique facial features, etc. The cartoonizing portrait images problem has a wide range of applications, for example in the cartoon film industry or developing filters for Tiktok, Instagram, etc. In this dissertation, we aim to study the cartoonizing portrait images prob- lem, including the overview, challenges, and difficulties.
Because the paired dataset between the cartoon and the human face is unavailable. Conse- quently, we consider our problem as an unsupervised image-to-image one. We have three main contributions. Firstly, to our knowledge, we public the largest cartoon face dataset for related researches.
Secondly, we proposed a method for cartoonizing the human face fast and robustly. Last but not least, base on Frechet Inception Distance (FID) metrics, we proposed a rea- sonable way to quantitatively evaluate our problem. Experimental results show that our method is capable of generating high-quality cartoon face anime and outperforms other state-of-the-art models. Contents 1 Introduction 1 11 Mofivalion.
HQ va 1 1117 Inappleation. ch xa 1 112 Inscience. mm co 2 12 Overvev.2 Challenges and Dilicullies. eeen 4 14 Main Contributions.
2ee 5 2 Related work 7 P Overview.2 Traditional image processing approaches.3 Learning based approaches.1 Methods based on Neural Style Transfr .2 Methods based on Generative Adversarial Network .3 Two-stage fine-tuning approach.2 Synthesising Cartoon Face phase. eee 40 vii 4 Experiments A2 4.1 Frechet Inception Distance 42 4. Aor ag ee 52 References 54 viii List of Figures 1.1 Applying cartoonization model to cartoon and comic production. An cartoon filter in famous social networks.
Illustration of the cartoonizing portrait image problem. From left to right are the image of human face and its cartoon version.1 An example of style transfer problem.2 The overview in training phase of FastÑST.3 The general architecture of Generative Adversarial Networks .4 An example for generator and discriminator in DOGAN.5 Training the Pix-to-Pix network for translating edges to photo.6 The comparison of Encoder-Decoder and UNet architectures. The comparison of SingleGAN, PatchGAN, and ImageGAN .8 Some applications of Pix-to-Pix in real-world problems .10 Overview of the CycleGAN training.11 High level structure of Generator in CycleGAN .12 CUT: Patchwise Contrastive Learning allows one-side translation 2.13 Describe the process of calculating PatchNCE loss.14 The one-side architecture in U-GAT-IT.15 An example of CartoonGAN application. From left to right are real- world scene and its cartoon Version.
cà cà so 2.16 The description of CartoonGAN traming .17 The new cartoonization system is proposed by Wang etal.2 Common features of cartoon face image. Unique facial features; 2. Sharp and Clear edges; 3. Smooth and parse blocks of color.3 Our proposed Two-stage Fine-tuning approach.4 Resnet-based architecture for image abstraction algorithm .5 Computing perceptual loss bases on ResNet-18 classifier.
39 Architecture of the discriminator in proposed method .7 Architecture of the generator in proposed method. In this case, FID equals to 57.34, which is the smallest FID score in our experiment. However, we find that the results well capture cartoon features but ignore the global structure of the input.2 Let A,B’, and B are the source, target, and generated images. We consider the EIDpp to take FIDag as the upper bound.4 Result of removing and changing each components in our proposed ap- proach.
50 List of Tables 21 Inference time (in seconds) of NST and FastNST. All benchmarks use a Titan X GPU. ee 41 Performance comparison based on FID, ASSIM and Fscoregp .2 Quantitative analysis of each component in our proposed approach.3 Summary table of existing methods and our proposed approach. Three last rows are the comparison when ablating each components.
(*) Be- sides three mentioned criteria, the result of ablating smooth stage cause messy noise. xi Chapter 1 Introduction 1.1 In application Cartoon is a popular art form that widely presents in our daily life. The cartoon film and comic production grow dramatically fast recently. According to Statista [1], the global animation market was worth 256 billions US dollars and is expected to reach 270 billions in 2020.
As a result, the relevant industries also develop to meet a large number of fans around the world. In this dissertation, we aim to build a method that automatically cartoonizes a given portrait image. It could spend hours with a professional artist. Hence, our method has a wide range of real-world applications, especially in: e Cartoon film and comic production.
Nowadays, a lot of animation and comics are inspired from reality. Existing tools and research works allow to transfer the general scenes. However, cartoonizing portrait images are very com- plex due to some requirements of cartoon face, such as smooth color, clear edges, unique facial features, etc. Cartoon portraits use in customizing characters of the online game.
Commom social networks, such as Tiktok, Snapchat, or Prisma release filters for cartoonizing human challenges that have become trending recently. However, those filters are too slow to perform. Due to the need of many resources in the inference time so they need to serve as a client-server system.2: An cartoon filter in famous social networks [2] 1.2 In science Art and science are always a constant companion of our life. Previous techniques of creating masterpieces in computer graphic require meticulous system design and devel- oper’s expertise in art.
Thanks to the break of Neural Style Transfer [19], in computer vision research, the studies in creative problems have attracted many scientists. Cartoonizing portrait images is a new and challenging problem. The results not only demand the need for a smoother texture but also synthesis cartoon facial features, which are unable to exist methods. To address this problem, we focus on Generative Adversarial Networks [23] (GAN) based algorithms.
However, instability training and time-consuming are two serious problems in GAN. These above issues prompt us to do this dissertation, where we introduced a robust and fast (both in training and inference time) for the cartoonizing an image of human face.1 Definition Cartoonizing portrait images problem is a sub-branch of the unsupervised image-to- image translation problem.3 illustrates the process of cartoonizing a human face image. e Input: an image of human face e@ Output: a cartoonized image of human face 1.2 Challenges and Difficulties Besides common challenges of computer vision tasks, the cartoonizing portrait im- ages problem has its own difficulties, which almost come from Generative Adversarial Networks and cartoon domain. e Lack of public large-scale dataset.
A paired dataset is unable to our problem, so we consider looking for an un-paired dataset. To our survey, the only suitable public dataset is selfie2anime from [31]. However, selfie2anime just contains female images, which may affect the performance of learning-based methods. e Time-consuming in training phase.
The consequences of using an un-paired dataset let us go to cycle consistency based models, which integration two inverse deep learning model. Hence, the training time is double. e The instability training. To synthesis cartoon face features, we intend to apply mechanism of Generative Adversarial Network.
During the adversarial training, the adversarial loss does not easily converge like other deep learning models. e Lack of reliable evaluation metrics. Evaluation metrics in style transfer problems are always the largest barrier to scientists. The FID score [28] or previous metrics do not satisfy our problem due to some reason, which is discussed in section 4.3 Objectives In this dissertation, we aim to solve the problem of portrait cartoonization.
We define a so-called good model, which satisfies the following three conditions: (1) having smooth and parse color blocks with clears edges, (2) synthesising cartoon facial feature, and (3) the resulting image is identifiable. Besides, we find that the existing metrics [28] [41] [11] do not suitable for un- supervised image-to-image translation problems in general and cartoonizing portrait images in particular. Therefore, we want to build a proper evaluation for this group of problems.4 Main Contributions In this dissertation, the main contributions of this paper are as follows: e Investigating overview of related research. Creating artwork has a long history of research.
We conduct a survey from traditional image processing to learning based methods. Specifically, we focus on Neural Style Transfer based and Generative Adversarial Networks based models. e Introducing the CartoonFace10k dataset. To our knowledge, this is the largest and most qualified dataset for cartoon face.
e Developing a two-stage training process for cartoonizing portrait im- age. This training scheme is adopted by two remarkable changes in a cartoonized portrait image compared to the real-world one. e Proposing the Escoresg. We propose a reasonable evaluation metrics called F sc ør by combining SSIM and FID metrics for the cartoonizing portrait im- ages problem in particular and unsupervised image-to-image problems in general.
e Setting up experiments. Extensive experiments including quantitative, qual- itative, user study, and ablation study have been conducted to prove proper of Fscoregr metric. Then, we find that our two-stage training establishes new state-of-the-art with Fscoregp.5 Dissertation Outline The structure of this dissertation includes: e Chapter 1: Introduction. We present an overview of the cartoonizing portrait images problems, which consists of research motivation, definition, challenges, and our main contributions.
e Chapter 2: Related work. We describe our survey on related research works, which directly or in-directly resolve our problem. e Chapter 3: Proposed methods. We introduce a new dataset CartoonFace10K and two-stage training scheme approach, which concentrate in cartoonizing im- ages of human face.
We first revise the evaluation protocol for our problem. Then, we show and explain some experiments to compare our new methods with other state-of-the-art models. We summarize our main contribution and discuss future research work. Chapter 2 Related work 2.1 Overview In this chapter, we introduce some approaches to the cartoonizing human face prob- lem.
Firstly, we briefly discuss traditional image processing methods including Non- Photorealistic Rendering [21] (NPR) and Image Analogies [27], which are not involved in deep learning.