BOSTON UNIVERSITY GRADUATE SCHOOL OF ARTS AND SCIENCES Dissertation A LAMINAR CORTICAL MODEL OF 3D SURFACE PERCEPTION AND FIGURE-GROUND SEPARATION: STEREOGRAM DEPTH, LIGHTNESS, AND AMODAL COMPLETION by LIANG FANG B., University of Science and Technology of China, 1998 M., Automation Institute, Chinese Academy of Sciences, 2001 Submitted in partial fulfillment of the requirements for the degree of Doctor of Philosophy 2007 UMI Number: 3246603 Copyright 2006 by Fang, Liang All rights reserved. INFORMATION TO USERS The quality of this reproduction is dependent upon the quality of the copy submitted. Broken or indistinct print, colored or poor quality illustrations and photographs, print bleed-through, substandard margins, and improper alignment can adversely affect reproduction. In the unlikely event that the author did not send a complete manuscript and there are missing pages, these will be noted.
Also, if unauthorized copyright material had to be removed, a note will indicate the deletion. ® UMI UMI Microform 3246603 Copyright 2007 by ProQuest Information and Learning Company. All rights reserved. This microform edition is protected against unauthorized copying under Title 17, United States Code.
ProQuest Information and Learning Company 300 North Zeeb Road P. Box 1346 Ann Arbor, MI 48106-1346 © Copyright by LIANG FANG 2006 Approved by First Reader reraleso Stephen Grossberg, PA. Ồ Wang Professor of Cognitive and Neural Systems Professor of Mathematics, Psychology, and Biomedical Engineering Second Reader Emmi Thing Ennio Mingolla, Ph. Professor of Cognitive and Neural Systems and Psychology Third Reader Gaal Carpenter,đã D.
Professor of Cognitive and Neural Systems and Mathematics Acknowledgements I would first like to thank my advisor, Professor Stephen Grossberg, for his talented scientific insights and generous guidance throughout my work. I would also thank Professor Ennio Mingolla for introducing me into this fantastic vision field and much other generous assistance. I appreciate and cherish the assistance and friendship that I received from everyone in the CNS community. Special thanks to Dr.
Yongqiang Cao and Dr. Arash Yazdanbakhsh for the intuitive discussions and wonderful coffee times that we spent together. I wish to present my sincere thanks to my parents who, as always, have been constant sources of encouragement and supports. My final thanks are reserved to my wife who has accompanied me all the way through these years and made them part of lovely memories.
iv A LAMINAR CORTICAL MODEL OF 3D SURFACE PERCEPTION AND FIGURE-GROUND SEPARATION: STEREOGRAM DEPTH, LIGHTNESS, AND AMODAL COMPLETION (Order No. ) LIANG FANG Boston University Graduate School of Arts and Sciences, 2007 Major Professor: Stephen Grossberg, Wang Professor of Cognitive and Neural Systems. Professor of Mathematics, Psychology, and Biomedical Engineering ABSTRACT In viewing a 3D scene, object features are seen on 3D surfaces infused with lightness and color at correct depths. By only focusing on how left and right features are correctly matched, most 3D vision models have not explained how this happens.
A 3D LAMINART model (Grossberg and Howe, 2003; Cao and Grossberg, 2005) proposed that laminar cortical mechanisms interact to create 3D surface percepts using interactions between boundary and surface representations. Previous work using this model explained perception of relatively simple objects, like bars and blocks, in relatively simple spatial configurations that did not contain any mutual occlusions. This thesis extends the 3D LAMINART model to predict how textured images with multiple potential false binocular matches, e. dense stereograms, generate correct 3D surface representations of figures and their backgrounds.
The model also clarifies how sparse stereograms can induce the formation of continuous surfaces at correct depths across contrast-free regions. Furthermore, when textured stereograms define emergent occluding and occluded v surfaces, the model shows how these surfaces are correctly separated in depth and the partially occluded textured surfaces can be amodally completed behind the occluding textured surface. Thus, the model provides a unified explanation of stereopsis, 3D figure- ground separation, and completion of partially occluded object surfaces. The model clarifies how interactions between layers 4, 3B, and 2/3A in V1 and V2 contribute to stereopsis, and proposes how a disparity filter and 3D perceptual grouping laws in V2 interact with 3D surface filling-in operations in V1, V2, and V4 to produce appropriate figure-ground perception.
These interactions help to convert the complementary rules for boundary and surface formation (Grossberg, 1994) into a consistent, unitary visual percept. vi TABLE OF CONTENTS ACKNOWLEDGEMENTS. co on HH nọ ni in Bi n0 001860916 iv ABSTRACTT.-- co co HH HH HH BI 0.0 09609 090009019619 61999 Vv LIST OF TABLES .ccccscecccsvsvsccccccscscsccesessccescnsesessesesssssesseceesenenes ix LIST OE EIGURES.- Go HH nh 1G GB 999006 x LIST OF ABREVIATIONS. on n1 BI 00 xii 1 Introducfion.-- onnọ Họ Đ BI 0 1000600009196 1 2 Sfereopsis processing and model cons(raÌnfS.1 Reconciles contrast-specific binocular fusion with contrast-invariant boundary percepfiOn.c c ng ng nh kh ke 29 2.2 Implements the contrast magnitude constraint on binocular fusion.3 Encourages the unique-matching rule and solves the correspondence 2.4 Combines monocular and binocular information to form depth percepts.5 Forms perceptual groupings including amodal boundary completions.6 Forms 3D amodal and modal surface representations.7 Determines the correct depth for horizontal boundaries, ensures perceptual consistency, and initiates figure-ground separation with surface-to- boundary feedback.
con HH nh kh, 43 2.8 Uses small-scale texture borders to induce large-scale perceptual ĐTOUPITES. eee 48 Vil Model description: Enhanced 3D LAMINART. con Sen se 52 3.- chen nen heo 53 3.2 VI binocular boundaries. ch nen se 53 3.3 VI monocular SUTÍAC€S.
ch nh kh ha 55 3. cọ cọ SH nh nh bà nến 56 3.5 V2 monoculÌar SUTÍAC€S. cọ HH nh kh sen 58 3.6 V4 binocular SUTÍAC€S. HH nh hen 39 Model equations.o co co Co no nh n0 6/0099066 6 60 4.1 Index legend and general processes.
HH nh kh nh ve 66 Model simulations.1 Border ownership and figure-ground percepfs. chnh n nà nà kệ 94 5.4 Figure-ground percept of RDS textured surfaces with emergent OCCÏUSIOT. SH Km nh Kinh KH tà tế 95 DÌSCUSSỈON. co co O0 no B0.1 Anatomical and physiological dafa.2 Model complexity and explanatOry pOW€F.
Hypotheses and predictions of the model. -- Go HH nh K0 BI 0008 105 CURRICULUM VITAE.-G con on ni BI 060908 124 LIST OF TABLES Table 1: Matrix of inhibition coefficients in the disparity filter. Table 2: The allelotropic shift.-‹-- -c CS SH ng kh, Table 3: Anatomical and physiological data that support the model. ix LIST OF FIGURES Figure 1: Horizontal binocular dISparity.
HH nh nh nhe 1 Figure 2: Strong 3-D percept elicited by plain 2-D image. 2 Figure 3: Perceptual ørOUpIng.--c-c ccn n n nn n KH kh nà nhe 3 Figure 4: Complementary processing of visual information. 8 Figure 5: Simulation of dense random-dot-stereogram. cà ceằ 11 Figure 6: Simulation of sparse random-dot-stereogram.
13 Figure 7: Simulation of random-dot-stereogram that implicitly define occlusion. 15 Figure 8: Comparison of boundary representations of a dense RDS simulated by two mOđ€ÌS. cọ SĐT nen Kế nà hà à KH hà tà tà ni nên 17 Figure 9: Examples of amodal completion. con nhà 19 Figure 10: Role of figure-ground separation in recognition of occluded letters.
22 Figure l1: 3D figure-ground separation and amodal completion. 24 Figure 12: 3D LAMINART macrOCIrCUII.cQ n nh nh nh nhe 27 Figure 13: 3D LAMINART model circuit diagram.- cà cv se 28 Figure 14: Contrast-specific binocular fusion and contrast-invariant boundary perception ¬ EEE E EERE LEER EE EEE EE EE EE EE EE OEE EE OE EEE EEO EO EE EEE EGER SH EEEEES 30 Figure 15: The V2 disparity filter. cccccccceccccceeeeceecuvuvecueeetesceveeeeenseses 32 Figure 16: da Vinci S†€f€OðTaim. HH renee nh kh nh nh kh khu 34 Figure 17: Boundary completion is formed inwardly, instead of outwardly.
36 Figure 18: Demonstration of 3D surface capture. co nen 41 Figure 19: Disparity ambiguity of horizontal boundaries.- 45 Figure 20: Demonstration of depth-determination of horizontal and monocular boundari€§. - con SH KH K KT KT kh nh ke nh by 46 Figure 21: Simulation of border ownership and figure-ground percept (part 1). 85 Figure 22: Simulation of border ownership and figure-ground percepts (part 2).
88 Figure 23: Simulation of border ownership and figure-ground percepts (part 3). 91 Figure 24: Role of V1 surface-to-boundary feedback in multiple-scale boundary DTOC€SSITE. ee cece e rete eee Een ene nO Re EEE EEO E EEE EE EA EA EG OHO HEED SEH eH EE ER FEES 97 xi LIST OF ABREVIATIONS 2D Two Dimensional 3D Three Dimensional FACADE Form And Color And DEpth IT Inferior Temporal area LAMINART | Laminar Adaptive Resonance Theory LGN Lateral Geniculate Nucleus RDS Random-Dot-Stereogram VI Visual area 1 V2 Visual area 2 V3 Visual area 3 V4 Visual area 4 xii Chapter 1 Introduction When we view a 3D scene, the retinas of our two eyes receive two-dimensional arrays of light, but we effortlessly perceive the world in depth. The positional differences of an object’s projections on an observer’s left and right retinas, or their binocular disparity, is a strong cue for perceiving depth at sufficiently near depths (Howard & Rogers, 2002; Julesz, 1971; see Figure 1).
Horizontal binocular disparity: when two eyes are focused on object A, the projections of object B at a different distance lie on different positions, B, and B,, on the left and right retinas. Such positional difference is used in the brain in reconstructing depth from 2-D retinal inputs. Binocular disparity is most effective for sufficiently near objects (Tyler, 2004). For distant objects, monocular cues, such as T-junctions, may be used to determine relative depth when one object is nearer than another object, and occludes parts of the farther object (Howard & Rogers, 2002).
Occlusion can also elicit strong 3D percepts when we view 2D images, see Figure 2. a T-junction indicates occlusion and relative depth Figure 2. Strong 3-D percept elicited by 2-D image: where occlusion acts as a cue of relative depth, especially for distant objects, such as mountains. (Courtesy to anonymous artist) In addition to being perceived at a farther depth, the visible parts of an occluded object are often perceptually linked together behind the occluder, see Figure 3.
Perceptual grouping: in addition to be perceived at a farther depth, the visible fragments of a tiger face are perceptually linked together instead of being seen as unrelated. (Courtesy to anonymous artist) The work that is presented in this thesis further develops a neural model of how the subcortical area LGN and the visual cortical areas V1, V2, and V4 work together to give rise to correct 3D boundary and surface percepts of binocular stimuli that contain disparity and occlusion information. Object features are seen on 3D surfaces infused with lightness and color at the correct depths. Most previous models of stereopsis (Dev, 1975; Fleet, Wagner, & Heeger, 1996; Grimson, 1981; Julesz, 1971; Lehky & Sejnowski, 1990; Lippert & Wagner, 2002; Marr & Poggio, 1976, 1979; Matthews et al., 2003; Ohzawa, DeAngelis, & Freeman, 1990, 1996; Prince & Eagle, 2000; Read, 2002; Sperling, 1970; Qian, 1997) restricted their attention to how left and right eye contours could be matched, but did not explain how this matching process spontaneously gives rise to continuous percepts in depth of surface features, including lightness.
FACADE theory (Grossberg, 1987, 1994, 1997; Kelly & Grossberg, 2000; McLoughlin & Grossberg, 1998) has proposed how 3D surface features, including lightness, can be represented as a result of interactions between boundary and surface processing streams. Recently, a 3D LAMINART model has developed FACADE theory to predict how laminar circuits within the visual cortex generate 3D boundary and surface percepts and separate figures from their backgrounds (Cao & Grossberg, 2005; Grossberg & Howe, 2003; Grossberg & Swaminathan, 2004; Grossberg & Yazdanbakhsh, 2005). This thesis further develops the 3D LAMINART model to explain how 3D boundary and surface percepts occur, and partially occluded objects can be completed and correctly recognized as a whole, in response to both dense and sparse Random-Dot-Stereograms (RDS) A vigorous recent discussion on CVNet, which was initiated by Jeremy Wilmer, summarized key aspects of the rich literature on the advantages of having binocular stereopsis.