Phân đoạn hình ảnh ngữ nghĩa trong điều kiện tối với phương pháp thích ứng miền

Khóa luận trình bày phương pháp phân đoạn ngữ nghĩa ảnh trong điều kiện thiếu sáng, sử dụng kỹ thuật tương thích miền dữ liệu hiệu quả.

Chuyên ngành

Computer Science

Người đăng

Ẩn danh

Thể loại

Thesis

2021

101
2
0

Phí lưu trữ

35 Point

Mục lục chi tiết

Acknowledgement

ABSTRACT

1. CHƯƠNG 1: INTRODUCTION

1.1. Practical Context

1.2. Problem Definition

2. CHƯƠNG 2: RELATED WORK

2.1. Fundamental Knowledge in Convolutional Neural Network

2.2. Semantic Image Segmentation

2.2.1. Overview of Semantic Image Segmentation

2.2.2. Fully Convolutional Networks

2.2.3. Pyramid Scene Parsing Network

2.2.4. Google DeepLab Family

2.3. Generative Adversarial Network

2.3.1. Overview of Generative Adversarial Network

2.3.2. Conditional Generative Adversarial Network

2.3.3. Pix2Pix

2.3.4. CycleGAN

2.3.5. Image Domain Adaptation

3. CHƯƠNG 3: PROPOSED FRAMEWORK

3.1. GAN-based Image Translation Component

3.1.1. Variational Autoencoders-GAN

3.1.2. Perceptual loss maintains the semantic features

3.2. Semantic Image Segmentation with Self-training Strategy

3.2.1. Panoptic Feature Pyramid Networks

3.2.2. Proposed Loss Function

3.2.3. Our Self-training Method

4. CHƯƠNG 4: EXPERIMENTS

4.1. Datasets for Image Domain Translation

4.2. Datasets for Semantic Image Segmentation

4.2.1. Nighttime Driving Testset

4.2.2. Extra Unlabeled Data Selection

4.3. Day2Night Image Domain Translation Component

4.3.1. With Perceptual Loss Refinement Results

4.4. Semantic Image Segmentation Component

4.4.1. Daytime Cityscapes Images Training

4.4.2. Daytime and Nighttime Images Training

4.4.3. Daytime and Nighttime Images Training with Perceptual Loss for Image Translation Module

4.4.4. Only-night Images Training with Perceptual Loss for Image Translation Module

4.4.5. Daytime and Nighttime Images Training with Focal Loss

4.4.6. Daytime and Nighttime Images Training with FID-based Method for Extra Unlabeled Data

4.5. Lessons from Series of Experiments

4.5.1. Improving GAN-based Method

4.5.2. Improving Semantic Image Segmentation Component

Publication

References

Tóm tắt

I. Tổng quan về phân đoạn hình ảnh ngữ nghĩa trong điều kiện tối

Phân đoạn hình ảnh ngữ nghĩa là một nhiệm vụ quan trọng trong lĩnh vực thị giác máy tính. Nhiệm vụ này giúp máy tính hiểu và phân loại các đối tượng trong hình ảnh. Trong điều kiện tối, việc thực hiện phân đoạn này trở nên khó khăn hơn do thiếu ánh sáng và độ tương phản thấp. Nghiên cứu gần đây đã chỉ ra rằng các phương pháp truyền thống không đạt hiệu suất cao trong điều kiện này. Do đó, việc áp dụng các phương pháp mới như thích ứng miền là cần thiết để cải thiện kết quả.

1.1. Định nghĩa phân đoạn hình ảnh ngữ nghĩa

Phân đoạn hình ảnh ngữ nghĩa là quá trình phân chia hình ảnh thành các vùng có ý nghĩa khác nhau. Mỗi pixel trong hình ảnh được gán một nhãn tương ứng với đối tượng mà nó thuộc về. Điều này rất quan trọng trong nhiều ứng dụng như xe tự lái và y tế.

1.2. Tầm quan trọng của phân đoạn hình ảnh trong điều kiện tối

Trong điều kiện tối, việc phân đoạn hình ảnh trở nên khó khăn hơn do ánh sáng yếu. Điều này ảnh hưởng đến khả năng nhận diện và phân loại đối tượng. Các nghiên cứu đã chỉ ra rằng hiệu suất của các mô hình phân đoạn hình ảnh giảm đáng kể trong điều kiện này.

II. Thách thức trong phân đoạn hình ảnh ngữ nghĩa vào ban đêm

Một trong những thách thức lớn nhất trong phân đoạn hình ảnh ngữ nghĩa vào ban đêm là sự thiếu hụt dữ liệu hình ảnh đã được chú thích. Hầu hết các bộ dữ liệu hiện có chủ yếu tập trung vào hình ảnh ban ngày. Điều này dẫn đến việc các mô hình học được không thể hoạt động hiệu quả trong điều kiện tối. Ngoài ra, sự khác biệt về ánh sáng và màu sắc giữa hai miền cũng gây khó khăn cho việc áp dụng các mô hình đã được huấn luyện.

2.1. Thiếu hụt dữ liệu hình ảnh đã chú thích

Việc thiếu hụt dữ liệu hình ảnh đã chú thích cho điều kiện tối là một vấn đề lớn. Các bộ dữ liệu hiện có thường không đủ để huấn luyện các mô hình phân đoạn hình ảnh hiệu quả. Điều này dẫn đến việc các mô hình không thể nhận diện chính xác các đối tượng trong hình ảnh tối.

2.2. Sự khác biệt giữa hình ảnh ban ngày và ban đêm

Hình ảnh ban đêm thường có độ tương phản thấp và màu sắc không rõ ràng. Điều này làm cho các mô hình phân đoạn hình ảnh gặp khó khăn trong việc phân loại chính xác các đối tượng. Sự khác biệt này cần được xem xét khi phát triển các phương pháp mới.

III. Phương pháp thích ứng miền cho phân đoạn hình ảnh ngữ nghĩa

Phương pháp thích ứng miền là một giải pháp tiềm năng để cải thiện hiệu suất phân đoạn hình ảnh trong điều kiện tối. Bằng cách sử dụng mạng đối kháng sinh điều kiện (GAN), có thể chuyển đổi hình ảnh từ miền ban ngày sang miền ban đêm. Điều này giúp tạo ra dữ liệu hình ảnh đã chú thích cho điều kiện tối, từ đó cải thiện khả năng phân đoạn của mô hình.

3.1. Giới thiệu về mạng đối kháng sinh điều kiện GAN

Mạng đối kháng sinh điều kiện (GAN) là một phương pháp mạnh mẽ trong việc tạo ra hình ảnh mới từ dữ liệu đã có. GAN có thể được sử dụng để chuyển đổi hình ảnh ban ngày thành hình ảnh ban đêm, giúp tạo ra dữ liệu cho việc huấn luyện mô hình phân đoạn.

3.2. Ứng dụng của GAN trong phân đoạn hình ảnh

GAN có thể được áp dụng để tạo ra hình ảnh ban đêm từ hình ảnh ban ngày, từ đó cung cấp dữ liệu cho mô hình phân đoạn. Việc này giúp cải thiện độ chính xác của mô hình trong điều kiện tối, nơi mà dữ liệu thực tế rất hạn chế.

IV. Kết quả nghiên cứu và ứng dụng thực tiễn

Nghiên cứu đã chỉ ra rằng việc áp dụng phương pháp thích ứng miền có thể cải thiện đáng kể hiệu suất phân đoạn hình ảnh trong điều kiện tối. Các thử nghiệm cho thấy mô hình có thể phân loại chính xác hơn các đối tượng trong hình ảnh ban đêm. Điều này mở ra nhiều cơ hội ứng dụng trong các lĩnh vực như xe tự lái và giám sát an ninh.

4.1. Kết quả thử nghiệm với mô hình phân đoạn

Các thử nghiệm cho thấy mô hình phân đoạn sử dụng phương pháp thích ứng miền đạt được độ chính xác cao hơn so với các mô hình truyền thống. Điều này chứng tỏ rằng phương pháp này có thể giải quyết hiệu quả vấn đề phân đoạn hình ảnh trong điều kiện tối.

4.2. Ứng dụng trong xe tự lái

Phân đoạn hình ảnh ngữ nghĩa trong điều kiện tối có thể cải thiện khả năng nhận diện của xe tự lái. Điều này giúp xe tự lái hoạt động an toàn hơn trong môi trường tối, giảm thiểu rủi ro tai nạn.

V. Kết luận và triển vọng tương lai

Phân đoạn hình ảnh ngữ nghĩa trong điều kiện tối với phương pháp thích ứng miền là một lĩnh vực nghiên cứu đầy tiềm năng. Các kết quả đạt được cho thấy rằng có thể cải thiện đáng kể hiệu suất phân đoạn trong điều kiện tối. Tương lai, cần tiếp tục nghiên cứu để phát triển các phương pháp mới và tối ưu hóa các mô hình hiện có.

5.1. Tương lai của phân đoạn hình ảnh ngữ nghĩa

Nghiên cứu trong lĩnh vực phân đoạn hình ảnh ngữ nghĩa sẽ tiếp tục phát triển, đặc biệt là trong điều kiện tối. Các phương pháp mới sẽ được phát triển để cải thiện độ chính xác và hiệu suất của các mô hình.

5.2. Hướng nghiên cứu tiếp theo

Hướng nghiên cứu tiếp theo có thể tập trung vào việc cải thiện các phương pháp thích ứng miền và phát triển các bộ dữ liệu hình ảnh đã chú thích cho điều kiện tối. Điều này sẽ giúp nâng cao khả năng phân đoạn hình ảnh trong các ứng dụng thực tiễn.

10/07/2025
Khóa luận tốt nghiệp phân đoạn ngữ nghĩa ảnh trong điều kiện thiếu sáng với phương pháp tương thích miền dữ liệu

Trích đoạn nội dung tài liệu

VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY FACULTY OF COMPUTER SCIENCE NGUYEN THANH DANH - PHAN NGUYEN THESIS SEMANTIC IMAGE SEGMENTATION IN THE DARK WITH DOMAIN ADAPTATION METHOD HONORS BACHELOR IN COMPUTER SCIENCE HO CHI MINH CITY, 2021 VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY FACULTY OF COMPUTER SCIENCE NGUYEN THANH DANH - 17520324 PHAN NGUYEN - 17520828 THESIS SEMANTIC IMAGE SEGMENTATION IN THE DARK WITH DOMAIN ADAPTATION METHOD HONORS BACHELOR IN COMPUTER SCIENCE THESIS ADVISOR NGUYEN VINH TIEP, Ph. HO CHI MINH CITY, 2021 ASSESSMENT COMMITTEE The Assessment Committee is established under the Decision., date HE 118141111 1e by Rector of the University of Information Technology. Ly Quoc Ngoc — Chairman. Nguyen Thanh Son — Secretary.

Nguyen Hong Phuc — Member. Acknowledgement It is not possible to prepare a dissertation without the assistance and en- couragement of other people. This one is certainly no exeption. Thus, this dissertation is dedicated to those who make it possible: First of all, we would like the express our deep gratitude towards our advisor PhD.

Nguyen Vinh Tiep, who inspired us the passion of research, for his instructions and supports during the process of our work. Besides, we would like to thank the Faculty of Computer Science - University of Information Technology where we have been prepared enough knowledge to pursuing the field of Computer Science research. We also thank Multi Media Laboratory (MMLab-UIT) for supporting us the research environment and cutting-edge devices throughout the university journey. Specially, we want to say thank you to our fellow friends who support us with precious ideas and are always willing to help us when we are in need.

Last but not least, we genuinely thank our families without whose love and support we would not have been here to complete this dissertation at the university. ABSTRACT Semantic image segmentation is a fundamental computer vision task that brings many conveniences to human life. Its applications vary in many aspects such as autonomous vehicles, medical systems, photo editors and so on. For instance in the field of autonomous vehicles, the vehicles take images from cameras along with other sensors to understand the surroundings.

The taken images are automatically handled by computer vision algorithms including semantic image segmentation. In the normal condition of daytime, many previous work were finished with yielding high performance. However, in nighttime condition which we define as in the dark, there are several work done with poor results. To be specific, we discover that the specifications of images in daytime is different from those in nighttime mainly due to light condition.

Furthermore, there is a limitation of annotated nighttime images for semantic image segmentation purpose. Focusing on the application of autonomous vehicles, we do the task of semantic image segmentation on nighttime cityscapes images. In this work, we propose a complete framework to solve the problem of semantic image segmentation in the dark with the help of domain adaptation method. To address the problem of lacking suitable annotated dataset, we leverage Generative Adversarial Network - GAN to transfer images from daytime to nighttime domain.

In segmentation task, we apply self-training method and propose a loss function to enhance the performance of the model. Finally, we show the improvement and efficiency of our hypothesis in a series of experiments and lessons for future researches. ee xa 5 12 Objectives.4 Dissertation Structure 7 2 Related Work 2.2 Fundamental Knowledge in Convolutional Neural Network .3 Semantic Image Segmentation.1 Overview of Semantic Image Segmentation 10 2. Fully Convolutional Networks .5 Pyramid Scene Parsing NeUwork.7 Google DeepLab Family .4 Generative Adversarial Network.1 Overview of Generative Adversarial Network.2 Conditional Generative Adversarial Network .3 Pix2Pix Ha ee 33 244 CycleGAN.5 Image Domain Adaptation.1 Overview Proposed Framework .2 GAN-based Image Translation Component.1 Variational Autoencoders- GAN .4 Perceptual loss maintains the semantic features .3 Semantic Image Segmenation with Self-training Strategy .1 Panoptic Feature Pyramid Networks .2 Proposed Loss Function .3 Our Self-training Method .1 Datasets for Image Domain Translation.

Datasets for Semantic Image Segmentation.2 Nighttime Driving Testset.3 Extra Unlabeled Data Selecion. Day2Night Image Domain Translation Component.2 With Perceptual Loss Refinement Results .4 Semantic Image Segmentation Component .1 Daytime Cityscapes Images Training .2 Daytime and Nighttime Images Training .3 Daytime and Nighttime Images Training with Percep- tual Loss for Image Translation Module .4 Only-night Images Training with Perceptual Loss for Image Translation Module.5 Daytime and Nighttime Images Training with Focal Loss 73 4.6 Daytime and Nighttime Images Training with FID-based Method for Extra Unlabeled Data .3 Lessons from Series of Öxperiments.1 Improving GAN-based Method 80 5.2 Improving Semantic Image Segmentation Component. 80 A Publication 82 References 83 iv List of Figures 1.1 Illustration of the virtual vision of an autonomous car. Autonomous vehicles understand the surroundings with the help of cameras, sensors orradars, 60 .2 Artificial intelligence can well help automatically segment brain tumors.3 Satellite image segmentation helps automatically draw a map from an image.4 Definition of nighttime cityscapes image segmentation as a toy example.

Image segmentation model take a nighttime cityscapes image as input and process to output a segmentation map that assigns each pixel a semantic labgly. 4 21 Illustration of semantic segmentation problem. The Image Segmentation Model identifies that there exists a cat in the input image and which pixels belong to the cat.2 Results of semantic image segmentation and instance segmentation. Se- mantic segmentation on the left segments all the cats as one semantic class; instance segmentation on the right outputs each cat as a separate semantic segment.3 Fully Convolutional Networks [31] proposed by J.

Long et al in 2015.4 Upsampling example using transposed convolution. The input and kernel are 2 x 2 matrices, transposed convolution is performed with stride = 1. Multiplying the kernel with each component of the input, we get four 4x 4 matrices. The output is the sum of those four matrices.5 Fusing output to establish three variances of FCNs [31].6 Differences among the three variances of FCNs with an exemplary image from PASCAL VOC 2012 [31].7 Instances of the drawbacks of FCNs model [35].

(a) Too large objects face over segmentation issue; (b) Too small objects is missed.8 Deconvolutional Network [35] Architecture.9 Illustration of pooling and unpooling in DeconvNet [35].10 Comparison of the results of FCNs and DeconvNet [35].11 Visualization of decoder stage represented as heat maps. The information of the bicycle is gradually recovered in the last layers of DeconvNet [35].13 Extracted features are the results of a wider region |39].14 Architecture of Pyramid Scene Parsing Network |53]|.15 Results of Pyramid Scene Parsing Network on Cityscapes benchmark [53].16 Overall Architecture of RefineNet Model [2].17 Illustration of standard and atrous convolution. Atrous convolution gives a wider field of view with the same number of parameters compared with standard convolution.18 Illustration the differences when using standard and atrous convolution (r = 2). When applying standard convolution to the input image, we ownsize the image to H/2 x W/2 then apply filter 7 x 7.

In contrast, atrous convolution is directly applied to H x W input image. Basically, when using standard convolution, we process on 1/4 of the input image then upsize the feature map in output. Atrous convolution is applied to the whole image, therefore results in more dense feature maps.20 Atrous spatial pyramid pooling module in DeepLabv3 [4].21 Visualization of normal convolution.22 Visualization of separable convolution.23 Visualization of atrous separable convolution [ð].24 The application of GAN to generate datasets. New example images in datasets are generated by GAN.

(a) MNIST handwritten digit dataset, (b) CIFAR-10 small object photograph dataset, (c) Toronto Face database.25 Applications of image to image translation [16]. Image translation based on paired dataset such as day-to-night, sketch-to-image, segment map- to-photo and soon, 6. HH kg và 29 2.26 Applications of image to image translation [54]. Various domains are translated using unpaired image-to-image translation methods such as Monet-to-photo, zebras-to-horses, summer-to-winter and soon.27 An overview framework of GAN which contains the generator model G and the discriminator model D.28 Comparison result between traditional GAN and Conditional GAN [32].

(a) is randomly generated images, (b) is conditional generated images.1 Our Semantic Segmentation with Domain Adaptation Method Framework. The shared latent space assumption j30].3 Image-to-Image translation results in preliminary stage.4 Comparison of the result when the Perceptual Loss is applied and not.5 Overview of perceptual loss.6 General Panoptic Feature Pyramid Networks Architecture [20].7 Panoptic FPN Architecture for Semantic Segmentation.8 Illustration of self-training process in semantic segmentation module.1 Modified dataset contains two different domains which are daytime and nighttime distribufion.2 Exemplary images of Cityscapes Dataset.3 Exemplary images of Nighttime Driving Test.4 Results at 280k iterations. These images seem to be wrong located traffic and vehicle lights.5 Image-to-Image translation results with additional Perceptual loss.6 Image-to-Image translation results. It is too dark even for human vision.7 Image-to-Image translation results.

The results look too bright in com- parison with nighttime distribution.8 Examples of shiny sparkling of vehicle light in dataset.9 Examples of dark images in dataset.10 Examples of twilight images in đataset.11 Visualization of Segmentation Experiment-1 results.1 is the results of FPN-resnet101 trained on daytime images of Cityscapes; ID 1.2 is the results of self-training with the model from scratch; ID 1.3 is the results of self-training with checkpoint of ID 1. The unlabeled data in this experiment is 701 daytime cityscapes images of CamVid.12 Visualization of Segmentation Experiment-2 results.1 is the results of FPN-resnet101 trained on day and night images of Cityscapes; ID 2.2 is the results of self-training with checkpoint of ID 2. The unlabeled data in this experiment is around 15000 unlabeled nighttime cityscapes images of NEXĐT.13 Visualization of Segmentation Experiment-3 results.1 is the results of FPN-resnet 101 trained on day and images of Cityscapes; ID 3.2 is the results of self-training with the model with checkpoint of ID 3. The image translation module has been added perceptual loss to maintain objects features.14 Visualization of Segmentation Experiment-4 results.1 is the previous results trained on day and night images; ID 4.1 is the results of self- training with the model with checkpoint of ID 3.

only tests the effect of unlabeled data. Here we pick 1600 images from 15000 unlabeled nighttime images of NEXET dataset by histogram-based method.15 Visualization of Segmentation Experiment-5 results.1 is the results of FPN-resnet101 trained on only nighttime cityscapes images (generated by GAN); ID 5.2 is the model in ID 3.1 trained with more nighttime images; ID 5.4 are the results of self-training on 1600 unla- beled nighttime images selected by histogram-based method with the checkpoints of ID 5.16 Visualization of Segmentation Experiment-6 results.1 is the results of FPN-resnet101 trained on day and night cityscapes images with focal loss; ID 6.2 is the results of self-training on 1600 unlabeled nighttime images and also trained with focal loss. This experiment was held to compare the performance of focal loss and cross entropy loss among models.17 Visualization of Segmentation Experiment-7 results.1 is the results of self-training on around 1600 unlabeled nighttime images chosen by FID-based method and trained with cross entropy loss on checkpoint ID 3.2 is similar to ID 7.1 but with our proposed combined loss; ID 7.3 is to compared histogram-based with FID-based methods. FID-based method together with our proposed loss function yields the finest score.18 Visualization of Segmentation Experiment-8 results.1 is the results of self-training on around 1600 unlabeled nighttime images chosen by histogram-based method from checkpoint ID 5.2 is similar but unlabeled data is chosen with FID-based method, both 8.2 use cross entropy loss; ID 8.3 is the results of self-training on around 1600 unlabeled nighttime images from FID-based method with the checkpoint of ID 8.1 and the help of our proposed combined loss function.3 is our finest performance model.

TT viii List of 'Tables 21 A brief comparison of tasks in computer vision. 41 Cityscapes Dataset Class Definitions.2 Quantitative Result of Day2Night Translation Component.3 Results of Segmentation Experiment-1. Verifying self-training perfor- mance on daytime cityscapes dataset.4 Results of Segmentation Experiment-2. Narrowing down the distance between trainset and testset by adding generated nighttime cityscapes images together with self-training on true nighttime images.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ