Chuyển Đổi Ảnh Chân Dung Thành Hoạt Hình Sử Dụng Mạng Đối Kháng Sinh Generative

Luận văn tốt nghiệp nghiên cứu tốt nghiệp hoạt hình hóa ảnh chân dung sử dụng mạng đối nghịch tạo sinh, điều tra thực trạng, phân tích số liệu, đề xuất biện pháp cải tiến thực tế.

Chuyên ngành

Computer Science

Người đăng

Ẩn danh

Thể loại

Undergraduate Dissertation

2021

69
3
0

Phí lưu trữ

30 Point

Mục lục chi tiết

LỜI MỞ ĐẦU

1. CHƯƠNG 1: INTRODUCTION

1.1. In application

1.2. In science

1.3. Definition

1.4. Challenges and Difficulties

1.5. Objectives

1.6. Main Contributions

1.7. Dissertation Outline

2. CHƯƠNG 2: RELATED WORK

2.1. Overview

2.2. Traditional image processing approaches

2.3. Learning based approaches

2.3.1. Methods based on Neural Style Transfer

4. CHƯƠNG 4: EXPERIMENTS

4.1. Frechet Inception Distance

REFERENCES

Tóm tắt

I. Tổng Quan Về Chuyển Đổi Ảnh Chân Dung Thành Hoạt Hình

Chuyển đổi ảnh chân dung thành hoạt hình là một lĩnh vực đang phát triển mạnh mẽ trong công nghệ hình ảnh. Sự kết hợp giữa nghệ thuật và công nghệ đã tạo ra những sản phẩm độc đáo, thu hút sự quan tâm của nhiều người. Công nghệ này không chỉ giúp tạo ra những bức tranh hoạt hình từ ảnh chân dung mà còn mở ra nhiều cơ hội ứng dụng trong các ngành công nghiệp như điện ảnh, truyền thông và mạng xã hội.

1.1. Định Nghĩa Về Chuyển Đổi Ảnh Chân Dung

Chuyển đổi ảnh chân dung là quá trình biến đổi hình ảnh của con người thành các hình ảnh hoạt hình. Điều này đòi hỏi sự kết hợp giữa các kỹ thuật xử lý hình ảnh và nghệ thuật để tạo ra những bức tranh sống động và hấp dẫn.

1.2. Lịch Sử Phát Triển Công Nghệ

Công nghệ chuyển đổi ảnh chân dung đã có từ lâu, nhưng sự phát triển của các mạng đối kháng sinh Generative (GAN) đã mang lại những bước tiến vượt bậc. Các nghiên cứu gần đây đã chỉ ra rằng GAN có khả năng tạo ra những hình ảnh hoạt hình chất lượng cao từ ảnh chân dung thực tế.

II. Thách Thức Trong Chuyển Đổi Ảnh Chân Dung Thành Hoạt Hình

Mặc dù công nghệ chuyển đổi ảnh chân dung thành hoạt hình đã đạt được nhiều thành tựu, nhưng vẫn còn nhiều thách thức cần phải vượt qua. Những thách thức này bao gồm việc thiếu dữ liệu lớn, thời gian huấn luyện lâu và độ ổn định trong quá trình đào tạo.

2.1. Thiếu Dữ Liệu Đào Tạo Chất Lượng

Một trong những thách thức lớn nhất là thiếu các bộ dữ liệu lớn và chất lượng cao cho việc huấn luyện. Việc không có bộ dữ liệu phù hợp có thể ảnh hưởng đến hiệu suất của các mô hình học máy.

2.2. Thời Gian Huấn Luyện Lâu

Quá trình huấn luyện các mô hình GAN thường mất nhiều thời gian, đặc biệt khi sử dụng các bộ dữ liệu không được ghép cặp. Điều này có thể làm chậm tiến độ nghiên cứu và phát triển.

III. Phương Pháp Chuyển Đổi Ảnh Chân Dung Thành Hoạt Hình

Để giải quyết các thách thức trong việc chuyển đổi ảnh chân dung thành hoạt hình, nhiều phương pháp đã được phát triển. Các phương pháp này chủ yếu dựa trên các mạng đối kháng sinh Generative (GAN) và các kỹ thuật học sâu khác.

3.1. Sử Dụng Mạng Đối Kháng Sinh Generative

Mạng đối kháng sinh Generative (GAN) đã trở thành công cụ chính trong việc chuyển đổi ảnh chân dung thành hoạt hình. GAN cho phép tạo ra các hình ảnh mới bằng cách học từ các mẫu có sẵn.

3.2. Kỹ Thuật Huấn Luyện Hai Giai Đoạn

Phương pháp huấn luyện hai giai đoạn giúp cải thiện chất lượng hình ảnh hoạt hình bằng cách tối ưu hóa các đặc điểm của ảnh chân dung và hình ảnh hoạt hình.

IV. Ứng Dụng Thực Tiễn Của Chuyển Đổi Ảnh Chân Dung Thành Hoạt Hình

Công nghệ chuyển đổi ảnh chân dung thành hoạt hình có nhiều ứng dụng thực tiễn trong các lĩnh vực khác nhau. Từ sản xuất phim hoạt hình đến việc phát triển các bộ lọc cho mạng xã hội, công nghệ này đang ngày càng trở nên phổ biến.

4.1. Ngành Công Nghiệp Phim Hoạt Hình

Công nghệ này được sử dụng rộng rãi trong ngành công nghiệp phim hoạt hình, giúp tạo ra các nhân vật hoạt hình từ các diễn viên thực tế.

4.2. Phát Triển Bộ Lọc Cho Mạng Xã Hội

Nhiều ứng dụng mạng xã hội như Instagram và TikTok đã tích hợp các bộ lọc hoạt hình, cho phép người dùng tạo ra các bức ảnh chân dung độc đáo và thú vị.

V. Kết Luận Về Tương Lai Của Chuyển Đổi Ảnh Chân Dung Thành Hoạt Hình

Tương lai của công nghệ chuyển đổi ảnh chân dung thành hoạt hình rất hứa hẹn. Với sự phát triển không ngừng của công nghệ AI và các phương pháp học sâu, chất lượng và tốc độ của quá trình chuyển đổi sẽ ngày càng được cải thiện.

5.1. Tiềm Năng Phát Triển Công Nghệ

Công nghệ chuyển đổi ảnh chân dung thành hoạt hình có tiềm năng lớn trong việc cải thiện trải nghiệm người dùng và tạo ra các sản phẩm sáng tạo mới.

5.2. Hướng Nghiên Cứu Tương Lai

Các nghiên cứu trong tương lai có thể tập trung vào việc phát triển các mô hình GAN mạnh mẽ hơn và cải thiện độ chính xác của các hình ảnh hoạt hình được tạo ra.

10/07/2025
Khóa luận tốt nghiệp hoạt hình hóa ảnh chân dung sử dụng mạng đối nghịch tạo sinh

Trích đoạn nội dung tài liệu

VIETNAM NATIONAL UNVIERSITY HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY COMPUTER SCIENCE FACULTY HO SY TUYEN UNDERGRADUATE DISSERTATION CARTOONIZING PORTRAIT IMAGES USING GENERATIVE ADVERSARIAL NETWOK BACHELOR OF SCIENCE IN COMPUTER SCIENCE HONORS PROGRAM HO CHI MINH CITY, 2021 VIETNAM NATIONAL UNVIERSITY HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY COMPUTER SCIENCE FACULTY HO SY TUYEN- 17521244 UNDERGRADUATE DISSERTATION CARTOONIZING PORTRAIT IMAGES USING GENERATIVE ADVERSARIAL NETWOK BACHELOR OF SCIENCE IN COMPUTER SCIENCE HONORS PROGRAM SUPERVISOR Dr. NGUYEN VINH TIEP HO CHI MINH CITY, 2021 DANH SÁCH HỌI ĐÒNG BẢO VỆ KHÓA LUẬN Hội đồng chấm khóa luận tốt nghiệp, thành lập theo Quyết định số. của Hiệu trưởng Trường Đại học Công nghệ Thông tin. ACKNOWLEDGEMENT This dissertation would not be done without the great support and assis- tance from a lot of people.

First of all, I would like to express gratitude to my supervisor, Dr. Nguyen Vinh Tiep, who not only taught me fundamental knowledge but also gave me valuable guidance in intensive research. Next, I would like to acknowledge Nguyen Thanh Danh, Phan Nguyen, Nguyen Vu Anh Khoa, Ngo Huu Manh Khanh, and other members from Multimedia Lab (University of Information and Technology). They pro- vided stimulating discussions, where my idea had been improved day by day.

I would also like to thank my colleagues Huu-Trong Le, Vinh-Loi Ly, Xuan- Phung Pham, and other members from Viettel Cyberspace Center (Viettel Group) for their patient support and all the opportunities that I was given to complete my dissertation. Last but not least, I would like to say thank my family for considering my education and being always with me in my tough time. ABSTRACT Besides doing traditional tasks such as classification, object detection, seg- mentation, and so on, computer vision’s techniques are more creative. The style transfer problem gives the machine the ability to generate an artistic image from a realistic one.

More specifically, since cartoon is a popular art form that widely presents in our daily life, the cartoon style transfer is also well-studied recently. However, despite the great progress in im- age stylization, few methods focus on cartoonizing the image of a human face, which is very challenging due to the complicated requirements of a cartoon face, such as smooth color, clear edges, unique facial features, etc. The cartoonizing portrait images problem has a wide range of applications, for example in the cartoon film industry or developing filters for Tiktok, Instagram, etc. In this dissertation, we aim to study the cartoonizing portrait images prob- lem, including the overview, challenges, and difficulties.

Because the paired dataset between the cartoon and the human face is unavailable. Conse- quently, we consider our problem as an unsupervised image-to-image one. We have three main contributions. Firstly, to our knowledge, we public the largest cartoon face dataset for related researches.

Secondly, we proposed a method for cartoonizing the human face fast and robustly. Last but not least, base on Frechet Inception Distance (FID) metrics, we proposed a rea- sonable way to quantitatively evaluate our problem. Experimental results show that our method is capable of generating high-quality cartoon face anime and outperforms other state-of-the-art models. Contents 1 Introduction 1 11 Mofivalion.

HQ va 1 1117 Inappleation. ch xa 1 112 Inscience. mm co 2 12 Overvev.2 Challenges and Dilicullies. eeen 4 14 Main Contributions.

2ee 5 2 Related work 7 P Overview.2 Traditional image processing approaches.3 Learning based approaches.1 Methods based on Neural Style Transfr .2 Methods based on Generative Adversarial Network .3 Two-stage fine-tuning approach.2 Synthesising Cartoon Face phase. eee 40 vii 4 Experiments A2 4.1 Frechet Inception Distance 42 4. Aor ag ee 52 References 54 viii List of Figures 1.1 Applying cartoonization model to cartoon and comic production. An cartoon filter in famous social networks.

Illustration of the cartoonizing portrait image problem. From left to right are the image of human face and its cartoon version.1 An example of style transfer problem.2 The overview in training phase of FastÑST.3 The general architecture of Generative Adversarial Networks .4 An example for generator and discriminator in DOGAN.5 Training the Pix-to-Pix network for translating edges to photo.6 The comparison of Encoder-Decoder and UNet architectures. The comparison of SingleGAN, PatchGAN, and ImageGAN .8 Some applications of Pix-to-Pix in real-world problems .10 Overview of the CycleGAN training.11 High level structure of Generator in CycleGAN .12 CUT: Patchwise Contrastive Learning allows one-side translation 2.13 Describe the process of calculating PatchNCE loss.14 The one-side architecture in U-GAT-IT.15 An example of CartoonGAN application. From left to right are real- world scene and its cartoon Version.

cà cà so 2.16 The description of CartoonGAN traming .17 The new cartoonization system is proposed by Wang etal.2 Common features of cartoon face image. Unique facial features; 2. Sharp and Clear edges; 3. Smooth and parse blocks of color.3 Our proposed Two-stage Fine-tuning approach.4 Resnet-based architecture for image abstraction algorithm .5 Computing perceptual loss bases on ResNet-18 classifier.

39 Architecture of the discriminator in proposed method .7 Architecture of the generator in proposed method. In this case, FID equals to 57.34, which is the smallest FID score in our experiment. However, we find that the results well capture cartoon features but ignore the global structure of the input.2 Let A,B’, and B are the source, target, and generated images. We consider the EIDpp to take FIDag as the upper bound.4 Result of removing and changing each components in our proposed ap- proach.

50 List of Tables 21 Inference time (in seconds) of NST and FastNST. All benchmarks use a Titan X GPU. ee 41 Performance comparison based on FID, ASSIM and Fscoregp .2 Quantitative analysis of each component in our proposed approach.3 Summary table of existing methods and our proposed approach. Three last rows are the comparison when ablating each components.

(*) Be- sides three mentioned criteria, the result of ablating smooth stage cause messy noise. xi Chapter 1 Introduction 1.1 In application Cartoon is a popular art form that widely presents in our daily life. The cartoon film and comic production grow dramatically fast recently. According to Statista [1], the global animation market was worth 256 billions US dollars and is expected to reach 270 billions in 2020.

As a result, the relevant industries also develop to meet a large number of fans around the world. In this dissertation, we aim to build a method that automatically cartoonizes a given portrait image. It could spend hours with a professional artist. Hence, our method has a wide range of real-world applications, especially in: e Cartoon film and comic production.

Nowadays, a lot of animation and comics are inspired from reality. Existing tools and research works allow to transfer the general scenes. However, cartoonizing portrait images are very com- plex due to some requirements of cartoon face, such as smooth color, clear edges, unique facial features, etc. Cartoon portraits use in customizing characters of the online game.

Commom social networks, such as Tiktok, Snapchat, or Prisma release filters for cartoonizing human challenges that have become trending recently. However, those filters are too slow to perform. Due to the need of many resources in the inference time so they need to serve as a client-server system.2: An cartoon filter in famous social networks [2] 1.2 In science Art and science are always a constant companion of our life. Previous techniques of creating masterpieces in computer graphic require meticulous system design and devel- oper’s expertise in art.

Thanks to the break of Neural Style Transfer [19], in computer vision research, the studies in creative problems have attracted many scientists. Cartoonizing portrait images is a new and challenging problem. The results not only demand the need for a smoother texture but also synthesis cartoon facial features, which are unable to exist methods. To address this problem, we focus on Generative Adversarial Networks [23] (GAN) based algorithms.

However, instability training and time-consuming are two serious problems in GAN. These above issues prompt us to do this dissertation, where we introduced a robust and fast (both in training and inference time) for the cartoonizing an image of human face.1 Definition Cartoonizing portrait images problem is a sub-branch of the unsupervised image-to- image translation problem.3 illustrates the process of cartoonizing a human face image. e Input: an image of human face e@ Output: a cartoonized image of human face 1.2 Challenges and Difficulties Besides common challenges of computer vision tasks, the cartoonizing portrait im- ages problem has its own difficulties, which almost come from Generative Adversarial Networks and cartoon domain. e Lack of public large-scale dataset.

A paired dataset is unable to our problem, so we consider looking for an un-paired dataset. To our survey, the only suitable public dataset is selfie2anime from [31]. However, selfie2anime just contains female images, which may affect the performance of learning-based methods. e Time-consuming in training phase.

The consequences of using an un-paired dataset let us go to cycle consistency based models, which integration two inverse deep learning model. Hence, the training time is double. e The instability training. To synthesis cartoon face features, we intend to apply mechanism of Generative Adversarial Network.

During the adversarial training, the adversarial loss does not easily converge like other deep learning models. e Lack of reliable evaluation metrics. Evaluation metrics in style transfer problems are always the largest barrier to scientists. The FID score [28] or previous metrics do not satisfy our problem due to some reason, which is discussed in section 4.3 Objectives In this dissertation, we aim to solve the problem of portrait cartoonization.

We define a so-called good model, which satisfies the following three conditions: (1) having smooth and parse color blocks with clears edges, (2) synthesising cartoon facial feature, and (3) the resulting image is identifiable. Besides, we find that the existing metrics [28] [41] [11] do not suitable for un- supervised image-to-image translation problems in general and cartoonizing portrait images in particular. Therefore, we want to build a proper evaluation for this group of problems.4 Main Contributions In this dissertation, the main contributions of this paper are as follows: e Investigating overview of related research. Creating artwork has a long history of research.

We conduct a survey from traditional image processing to learning based methods. Specifically, we focus on Neural Style Transfer based and Generative Adversarial Networks based models. e Introducing the CartoonFace10k dataset. To our knowledge, this is the largest and most qualified dataset for cartoon face.

e Developing a two-stage training process for cartoonizing portrait im- age. This training scheme is adopted by two remarkable changes in a cartoonized portrait image compared to the real-world one. e Proposing the Escoresg. We propose a reasonable evaluation metrics called F sc ør by combining SSIM and FID metrics for the cartoonizing portrait im- ages problem in particular and unsupervised image-to-image problems in general.

e Setting up experiments. Extensive experiments including quantitative, qual- itative, user study, and ablation study have been conducted to prove proper of Fscoregr metric. Then, we find that our two-stage training establishes new state-of-the-art with Fscoregp.5 Dissertation Outline The structure of this dissertation includes: e Chapter 1: Introduction. We present an overview of the cartoonizing portrait images problems, which consists of research motivation, definition, challenges, and our main contributions.

e Chapter 2: Related work. We describe our survey on related research works, which directly or in-directly resolve our problem. e Chapter 3: Proposed methods. We introduce a new dataset CartoonFace10K and two-stage training scheme approach, which concentrate in cartoonizing im- ages of human face.

We first revise the evaluation protocol for our problem. Then, we show and explain some experiments to compare our new methods with other state-of-the-art models. We summarize our main contribution and discuss future research work. Chapter 2 Related work 2.1 Overview In this chapter, we introduce some approaches to the cartoonizing human face prob- lem.

Firstly, we briefly discuss traditional image processing methods including Non- Photorealistic Rendering [21] (NPR) and Image Analogies [27], which are not involved in deep learning.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ