VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY FACULTY OF COMPUTER SCIENCE LE THANH PHUOC HIỂU THESIS SINGLE IMAGE HDR RECONSTRUCTION HONORS BACHELOR IN COMPUTER SCIENCE HO CHI MINH CITY, 2021 VIETNAM NATIONAL UNIVERSITY HO CHI MINH CITY UNIVERSITY OF INFORMATION TECHNOLOGY FACULTY OF COMPUTER SCIENCE LE THANH PHUOC HIỂU - 17520474 THESIS SINGLE IMAGE HDR RECONSTRUCTION HONORS BACHELOR IN COMPUTER SCIENCE THESIS ADVISOR Dr. NGO DUC THANH Dr. HUA BINH SON HO CHI MINH CITY, 2021 DANH SÁCH HỘI ĐÒNG BẢO VỆ KHÓA LUẬN Hội đồng chấm khóa luận tốt nghiệp, thành lập theo Quyết định số. của Hiệu trưởng Trường Đại học Công nghệ Thông tin.
1V Acknowledgements This thesis, a remarkable milestone that closed those years of me as an under- graduate student, would not have been possible without the support from people around me. First and foremost, I would like to thank my supervisor Dr. Binh-Son Hua. It has been a great privilege for me to work under your supervision.
With your knowl- edge and skills, you have guided me since I was still very new to research and have strengthened my research skill over time. Through your expertise, you have shown me possible directions as well as provided insightful feedback that has had a signif- icant impact on this thesis. I always feel grateful and lucky because I have worked under your supervision. Time sure flies when we look back at those years we have had in university.
I would also like to thank all my classmates in KHTN2017, especially Phat, Hoang- An, Trung, Nghiem, Bao, and a friend of mine - Ngoc-Khanh, for all the time we have been through that helped me become a lot better than who I was. This work has been done while I was doing my internship at VinAI Research. In this time, great supports from my colleagues also had inspired me a lot on this journey. This is the most memorable time I have ever had.
Once again, many thanks to all of you! Ho Chi Minh City, February 2021. Contents Acknowledgements iv Abstract xi 1 Introduction -waor, M £7".1 Dynamic Scenes HDR ReconstrucHon|.2 Single-Image HDR ReconstrucHon|.2 Indirect approach| 14 3 Our Method ¬ ee Le 3. ee eee eee 3.2 ‘Training process vi 3. Structure of the proposed network|.35 Up/Down-ExposureNetl.4 Loss function 29 4 Experimental Results 33 4.1 Peak Signal-to-Noise Ratio (PSNR).2 Structural Similarity Index Measure (SSIM)|.
eee eee 41 ¬¬ 41 Qualitative comparisons on tone-mapped images|. 44 5_ Conclusion and Future Work 50 5.3 Future work 51 References 52 vil List of Figures 1.1 A typical scene contains a high dynamic range|. Anexample dynamic scene that contains large foreground motion 7 2.2 The overall pipeline of|5][.3 The proposed CNN architecture used in [5]J|.4 The end-to-end framework proposed by [20]} .5 The Attention-Guided NetworkÌl.8 The suggested framework by [10] that learns to reverse the camera Cee 12 2.9 The proposed feature masking mechanism by [18]].10 An overview of the proposed method by [3].11 The proposed Deep Chain HDRI modell.12 The proposed network’s structure relationship by [8]] .13 The process to generate bracketed images of [8||.1 The typical LDR image formation pipelnel.2 The overview approach for our framework|.4 Training pipeline of our proposed framework|. 24 "¬ 25 27 the proposed framework|.1 Samples of bracketed images from the dataset synthesized by [3]} .2 Anexample of five different CRFs contains in the dataset].4 Multi-exposure images inferred using our framework}.
43 | Recursive Deep HDRI [8], and the corresponding ground truth] .6 Tone-mapped HDR images compare between different methods}. 46 VS eee 47 ¬ 48 1X List of Tables 4.1 Quantitative comparisons on HDRimagesl.2 Quantitative comparisons on LDR images|.3 Quantitative results on inferred bracketed images|. List of Abbreviations CRF Camera Response Function CNN Convolutional Neural Network EV Exposure Value HDR High Dynamic Range HVS Human Visual System LDR Low/Limited Dynamic Range MSE Mean-Squared Error PSNR Peak Signal-to-Noise Ratio SSIM Structural Similarity Index Measure TMO Tone Mapping Operator x1 Abstract Reconstructing high dynamic range (HDR) images from multiple exposures is a classical problem in computational photography. With an increased range of lu- minances and colors, HDR images can convey much more useful information than what could be achieved with a conventional camera in scenes with extreme illumi- nations.
However, the main challenge with HDR imaging using multiple exposures is that the exposures need to be aligned, merged, and the dynamic range needs to be compressed for display, which could result in visual artifacts such as blurring and ghosting. Recent advances in machine learning, especially deep learning, have attracted the research community to utilize deep models for HDR imaging. In this thesis, we investigate the problem of reconstructing an HDR image from a single photograph. This is an ill-posed problem as opposed to using multi-exposure im- ages in conventional methods.
To tackle this problem, we propose a new solution based on a neural network that synthesizes multiple exposures from a single pho- tograph and then reconstructs the HDR image from the estimated exposures. Our network reconstructs image irradiance to synthesize exposures by learning to in- vert the image formation pipeline, hallucinating details in under- and over-exposed regions. Our experiments show that our proposed model can achieve comparable performance to the state-of-the-art methods. Keywords: computational photography, high dynamic range imaging, HDR image reconstruction, inverse tone-mapping, machine learning, deep learning Chapter 1 Introduction In this chapter, we first introduce the problem of HDR imaging, describing the motivation and the challenge of single-image HDR reconstruction.
Then we briefly describe our solution and the thesis contributions. Finally, we outline the overall thesis structure.1 Motivation We are familiar with cameras in our daily life designed to mimic the human visual system (HVS). Such cameras can capture the surrounding environment as close as what our eyes can observe in general conditions. Unfortunately, this is hardly the case in extreme conditions.1] the first and second picture from the left shows a scenario that can be challenging to a conventional camera.
In this scene, if we set the exposure to capture all details in the blue sky with sunlight and white clouds, this will make the ground underexposed and not observable. Nevertheless, if we set the exposure to show the ground’s details, the sky be- comes saturated, leaving no details in the final photograph. By contrast, our eyes can observe the third image with all the details in the sky and the ground. Such differences are due to an essential factor in imaging: the dynamic ranges.
Remark- ably, the dynamic ranges captured by a camera and by our eyes are not the same. —A stop 4 stop HDR FIGURE 1.1: A typical scene contains a high dynamic range that consumer-grade cameras usually fail to capture all details. The first two images from the right are an LDR image with -4 and 4 stop, re- spectively. The last is an HDR image with all details that can be better observed.
Note that the HDR image has been tone-mapped to display using [16 tone-mapping technique. Source: camera captures images with relatively low dynamic ranges (LDR images) while our eyes can perceive very high dynamic ranges (HDR). HDR imaging in computational photography is a field of study that focuses on reconstructing HDR images from a set of LDR images. Images captured using our consumer-grade cameras often result in overexposed regions with many saturated details, less vibrate than our eyes can perceive.
The proportion between the largest and smallest value of luminance in the scene de- fined as the dynamic range. These cameras can only capture a limited dynamic range due to hardware limitations compared to the high dynamic range in nature and the dynamic range the HVS can detect. The limited dynamic range means more visual information is available in the scene than what can be captured and repro- duced by the camera. With the above limitations, the need for a higher dynamic range image is necessary.
HDR images can present a more excellent dynamic range of luminosity than a standard low dynamic range image, with the details of the Chapter 1. Introduction 3 brightest and dimmest object can be better observed. Apart from getting closer to HVS than LDR images can do, HDR images also have a wide range of applications such as image-based lighting, HDR display, computer vision. As the need for HDR imaging becomes evident, techniques for HDR imaging reconstruction are requisite.2 Challenges Despite these clear advantages of HDR image, reconstructing HDR image from low dynamic one is an ill-posed problem.
To reconstruct an HDR image, typically, one has to capture many LDR images with different exposure times or own an HDR camera. Owning an HDR camera is still not affordable widely currently and only available in specific industries. Therefore, the former approach seems more plau- sible at the moment. However, when we capture multi-exposure LDR images, the scene has to be static or with very little object motion contains within.
In reality, capturing bracketed images often include scene motion. These motions are often caused by the person holding the camera, the object’s motion in the scene, or both. If the mentioned criteria do not satisfy, the resulting image may be distorted, con- tain artifacts. For example, in Figure the input used to restore the HDR image was captured when the object’s motion within the scene leading to ghosting artifacts emerge in the HDR image.
In this case, the artifacts are because broadly used HDR image techniques only work well with the scene with little to no motion.3 Contributions Because of the difficulties mentioned above, a solution that can effectively recon- struct an HDR image from a single input LDR image is a plausible way to resolve this problem. By utilizing the deep learning model, we could generate a sequence Chapter 1. Introduction 5 of images with different exposures, which later use in conventional methods like or [12] to generate an HDR image. This thesis proposes a new framework that generates a sequence with an ar- bitrary size of different exposure images from a single input LDR image using a learning-based method.
Our main contributions are: ® A novel end-to-end framework can generate an arbitrary number of differ- ent exposure images instead of a fixed number that the previous works have shown. ® A way to incorporate prior knowledge from the imaging formation pipeline to constrain the model when training. ¢ The proposed framework, together with trained parameters, labels of the dataset used for training/testing, are made available online, enabling reproducibility for further research. ® A comprehensive quantitative and qualitative evaluation with results show that the proposed framework is comparable to existing models.4 Disposition This introductory chapter served as an introduction that gives the motivation, encountered difficulties in the field of HDR imaging, and the contribution of this thesis.
Chapter [2] will discuss related work in the HDR image reconstruction prob- lem and their current limitations. The proposed method, as well as general back- ground and definitions, are available in Chapter {3} Next, Chapter [4| will give more information on qualitative and quantitative results, dataset, and implementation details. Lastly, Chapter[5|summarizes the work and its limitations with further dis- cussion about future work. Chapter 2 Related Work This chapter will discuss related work that used the learning-based approach to solve the HDR reconstruction problem and its current limitations that have not been addressed yet.
As HDR reconstruction is a long-studied problem, researchers have proposed many techniques in the past attempting to solve this. This problem’s typical ap- proach is to use [1|, which merges multiple images with different exposures to reconstruct the HDR image. These methods often produce high-quality results given input images to satisfy their assumptions. One of its is no moving objects contained in the captured scene.
However, when it comes to the real scenario, this is very seldom the case. Those mentioned methods often fail to reconstruct the de- sired HDR image properly, leading to artifacts, ghosting, tearing in the final HDR image when motion is introduced in the scene.