KOREAN ZERO PRONOUNS: ANALYSIS AND RESOLUTION Na-Rae Han A DISSERTATION in Linguistics Presented to the Faculties of the University of Pennsylvania in Partial Fulfillment of the Requirements for the Degree of Doctor of Philosophy 2006 f oY ` Ellen F. Prince Supervisor of Dissertation “ns Martha Palmer Co-Supervisor of Dissertation ⁄ 2 4 Eugen Buckley [) \ Graduate Group Chairperson UMI Number: 3211080 INFORMATION TO USERS The quality of this reproduction is dependent upon the quality of the copy submitted. Broken or indistinct print, colored or poor quality illustrations and photographs, print bleed-through, substandard margins, and improper alignment can adversely affect reproduction. In the unlikely event that the author did not send a complete manuscript and there are missing pages, these will be noted.
Also, if unauthorized copyright material had to be removed, a note will indicate the deletion. ® UMI UMI Microform 3211080 Copyright 2006 by ProQuest Information and Learning Company. All rights reserved. This microform edition is protected against unauthorized copying under Title 17, United States Code.
ProQuest Information and Learning Company 300 North Zeeb Road P. Box 1346 Ann Arbor, MI 48106-1346 To my dear brother Baek-Kyoung. H Acknowledgements First and foremost, I thank my two advisors, Ellen Prince and Martha Palmer. I am eternally indebted to them for their encouragement and intellectual and spiritual! guidance, which shaped me as a researcher throughout my long journey towards graduation.
Ellen, whom I have the honor of being the last student of, has been a great source of inspiration to me for her sharp intellect and wisdom, as well as her unparalleled zest for life. I owe Martha not only for her generous financial support over the years, but also for her guardianship — she truly took me under her wings. I also thank Robin Clark for his teachings during my early years of graduate school; Bill Poser, whose love and knowledge of everything linguistic has always shown me the greatness to aspire to; and last but not the least, my committee, Maribel Romero and Ar- avind Joshi, who provided valuable feedbacks on the directions of the thesis, which not unlike any other theses started out as a entangled jumble of grand ideas. XRCE and LDC are two great institutions I was fortunate enough to have opportunities to conduct research at.
I thank Lauri Karttunen, Ken Beesley and Mike Maxwell for their mentorship during my times there. I would like to thank Justin Mott, Cassie Creswell, Tom Morton, Andy Schein, Sham Kakade, and Heejong Yi, with whomI shared the most cherished memories of my graduate school years. They stayed close and bore witness to my years of battle with my thesis; without their cheering, I might well be writing it still. lil There are many other friends and colleagues at Penn, who over the years made it much more than just a school but rather like a home to me.
I had many lunches, dinners, parties, and heated discussions with them: John Bell, Alexis Dimitriadis, Kieran Snyder, Ron Kim, Sophia Malamud, Uri Horesh, Eva Banik, Tom McFadden, Elsi Kaiser, Sandhya Sundare- san, Rashmi Prasad, Eleni Miltsakaki, Seungyen Yang, Eon-Suk Ko, Chunghye Han and Jiyoon Lee. My friends and mentors from my previous school, the linguistics program at Seoul National University, have been always there for me, sometimes in the States and other times across the Pacific Ocean. I thank Jinyoung Choi, Seunghun Lee and Hyunjoo Kim for their friendship (and also for being great “juniors” to me); Yoonshin Kim for being my close friend (and a great “senior”); Sookhee Chae, Shijong Ryu and Chulwoo Park for their kind mentoring. And finally Prof.
Chungmin Lee for the very first linguistics class I took and for guiding me into the field. He taught me during my SNU linguistics years, and he encouraged me to pursue this degree in the U. Also many special thanks to Gene Buckley and Amy Forsyth for their wonderful help regarding administrative matters; as I fumbled through the maze of graduation, their knowl- edge and attention on intricate procedural details saved me more than once from crucial mistakes. I cannot forget Mrs.
Carole Lingle, whose warm presence filled the department office in earlier years of my graduate school. Staff at IRCS also has been the quiet yet essential organizing force behind that great institution, to which I owe much of my profes- sional development. Finally, my deepest gratitude goes to my parents and my sister, without whose constant support and encouragement none of my academic goals would have been realized. Their faith in me and my love for them kept me going at times of difficulty and anguish.
iv ABSTRACT KOREAN ZERO PRONOUNS: ANALYSIS AND RESOLUTION Na-Rae Han Supervisors: Ellen F. Prince and Martha Palmer Zero pronouns, or dropped arguments, are a remarkably frequent phenomenon in Korean. This single syntactic form is made up of diverse subcategories, each of which is character- ized by distinct semantic and pragmatic properties. More widely acknowledged types are those that depend on other linguistic expressions for their reference: such text-dependent types include anaphoric and discourse-deictic zero pronouns.
Other text-independent types are deictic zero pronouns, generic and specific indefinite zero pronouns and situational zero pronouns, although it is possible for some of these text-independent types to enter coreferential relations with other nominal expressions in their surroundings. Previous re- search focusing on anaphoric zero pronouns, most notably that based on Centering Theory, claims that information-theoretic notions such as saliency govern their felicitous use and interpretation. While the general insight holds true, various efforts to encapsulate it by way of precise formulation of Cf-ranking or other hierarchies fall short, largely due to the fact that the notion is encoded by heterogeneous linguistic factors whose relations cannot be expressed single-dimensionally. From a language processing point of view, the diverse nature of Korean zero pronouns presents the unique challenge of blending the tasks of cat- egorization and identification of their antecedents.
In this dissertation, using Maximum Entropy as the machine learning method of choice, various statistical models for Korean zero pronoun resolution have been successfully trained and tested on two Korean Treebank corpora. These Models serve as a valuable opportunity for empirically testing various the- oretical claims and observations made on Korean zero pronoun anaphora. Features used in constructing the models and making predictions on zero pronoun reference encode lin- V guistic properties surrounding zero pronouns and their potential antecedents. The features found to have a particularly strong contribution are indeed those that encode the linguistic aspects that are commonly cited in the linguistic literature as playing a crucial role in Ko- rean Zero pronoun usage, such as topic-hood, subject-hood and the nullness of form.
While the relative importance of such features does not directly translate to linguistic hierarchies, it nevertheless provides support to some of the specific criteria used in them. vi Contents Acknowledgements iii Abstract M Contents vii List of Tables xi Introduction 1 I Analysis 4 1 Previous Work on Zero Pronouns 5 1.1 Previous Sentence-Level Work on Zero Pronouns .1 proin Government and Binding Theory.2 Huang’s (1983, 1984, 1989) Work on Chinese, Japanese and Ko- TEAN PYO a. Optimality Theory Approaches to Zero Pronouns .2 Centering Theory: A Discourse-Oriented Approach to Pronouns .1 The Centering Theory: An Overview .2 Centering Theory Across Languages. The Zero Pronoun and Centering.
30 2 Analysis of Korean Zero Pronouns 35 2.1 Defining the Object of the Study .1 Why Zero “Pronouns” .2 Overt Pronounsin Korean. eee ee ee 38 2.3 The Problem of Identification: Where to Find the Invisible .4 Korean Zero Pronouns by Reference Types.1 General Situational Zero Pronouns.43 Indefinite Personal Zero Pronouns.1 Specific Indefinite Zero Pronouns.2 +human Semantic Restriction on Generic Zero Pronouns 60 2. Coreference in Generic Zero Pronouns .4 Discourse-Anaphoric Zero Pronouns.5 Semantic Interpretation of Anaphoric Relations .1 Discourse(Textual)-Deictic Zero Pronouns .6 Deictic and/or Anaphoric: the Fuzzy Distincion. 76 3 The Centering Theory and Korean Zero Anaphora 80 3.1 Criteria for Cf Ranking: What EncodesSalence?.2 Establishing Cf Ranking for Korean.
Topic-Marked NPs vs.4 The Centering Theory and Zero Pronoun Resolution. 115 viii II Resolution 118 Overview 119 1 Previous Work 121 1.1 Pronoun Resolution: An Overview. 0000 ee eee nae 121 1.2 The Traditional Approach .13 The Statistical Approach.4 The Knowledge-Poor Approach .5 Resolution of Zero Pronouns: The Case of Spanish .2 Centering Theory and Pronoun Resolulon. Brennan, Friedman and Pollard(1987).2 Strube’s (1998)“S-lst”Approach.3 Left-Right Centering by letreault(1999).4 Resolution of Zero Pronouns Using Centering: The Case of Thai.
Optimality Approaches to Anaphora Resolulon.2 Hong (2002) and Kim (2003): OT-Based Korean Anaphora Reso- lution ©. nu kg và ky 141 2 The Data: The Penn Korean Treebank Corpora 145 2. HQ va ko 145 2. Q HQ ung vo 151 2.1 Zero Pronouns with a Intra-sentential Antecedent.2 Determining the Categories.
157 ix 3 A Rule-Based Approach: Variations of the Hobbs Algorithm 161 3.1 The Hobbs Algorithm.2 Variations of Hobbs on Korean Zero Pronouns 164 3.3 Significance of Syntactic Environments 168 344 Summary. 170 4 A Maximum Entropy Reference Resolution System 171 41 TheDesign .1 Classification and Resolution.2 Two-Phased vs. Single-Phased System 175 4. Maximum Entropy asa Ranking Method .4 Selection of Features.
183 42 Single-Phased System.1 Performance by Type. 196 43 Two-Phased System.4 A Resolution System for Anaphoric Zero Pronouns 204 4.1 A Maximum-Entropy Model 205 4.2 Feature Opt-InModels .5 Performance Scores: A Round-Up 219 Conclusions 222 Bibliography 227 List of Tables 1.1 Dropped topic subject in Italian (cantare(x), x=lui, x=topic) .2 Overt non-topic subject in Italian (cantare(x),x=lui) .3 Overt topic subject in English (sing(x), x=he, x=topic) .4 Overt non-topic subject in English (sing(x),x=he) .5 Overt topic subject in Yiddish (sing(x), x=he, x=topic) .6 Overt topic subject in Yiddish (ranedQ).7 Thai input with overt subject ©.8 Thai input with Prosubject. LH HH KV 21 1.9 Thai input with embedded Pro subJect.10 Centering transition states.1 Pronominal system of Korean, 1st and 2ndperson.2 Pronominal system of Korean, 3rdperson.3 Classification of Japanese zero anaphor by Kameyama (1985) .4 Classification of Korean zero pronouns. 46 11 Pronoun resolution algorithms for New York Times .2 Pronoun resolution algorithms for fictional texts.3 COHERE and ALIGN produce preference ranking of 4 centering transition TYPES / .1 Zero pronoun frequencies in KTB landKTB2.2 pro and PRO frequencies in the Penn Chinese Treebank5.3 Zero pronoun frequencies by grammaticalroles.4 Zero pronoun frequencies bytype .5 Overt pronouns inKTB 1, lstand2nd person .6 Overt pronouns in KTB 1, 3rd person and other .7 Overt pronouns in KTB 2, lstand 2ndperson.8 Overt pronouns in KTB 2, 3rd personandother.1 Naive Hobbs algorithm on Korean zero pronouns .3 Hobbs algorithm on Korean zero pronouns, adverbial NPs not considered .4 Hobbs algorithm on Korean zero pronouns, argument antecedents only .5 Hobbs algorithm on Korean zero pronouns, performance by clausal groups .1 Binary classification for zero pronoun and its potential antecedent NP in 06:2182 4.2 Binary classification of coreference for zero pronounin06:2 .3 Binary classification of coreference for zero pronounin06:2.4 Feature vector of coreference events for two pro subjects, partially repre- sented 2.5 Numbers of targets/events generated per7Ør2.6 Numbers of coreferential targets/events generated perpro.7 Training and testing set sizes forthe two corpora.
2nu ng Q g kg V v ki ki k va 194 4.9 Performance of models as a binary classifer.10 Confusion matrix for binary classifiers, KTBI.11 Confusion matrix for binary classifiers, KTB2.12 Performance of coreference resolution models.13 Performance of coreference resolution models, averaging on 10 cross-fold validation 2.14 Accuracy by zero pronoun type,KTBI,.15 Accuracy by zero pronoun type,KTB2.16 Prediction pattern for anaphoric (a) zero pronouns, KTB2 .17 Prediction pattern for deictic-speaker (1) zero pronouns, KTB2 .18 Prediction pattern for generic (g) zero pronouns, KTB2 .19 Confusion matrix for category classification model (Phase 1), KTB 1.20 Combined performance of Phase 1 and Phase2, KTB1.21 Comparison of two approaches, KTB 1 .22 Confusion matrix for category classification model (Phase 1), KTB2.23 Combined performance of Phase 1 and Phase2, KTB2.24 Comparison of two approaches, KTB2 .25 Numbers of coreferential NPs per NP-anaphorlcøo.26 Performance of coreference resolution models.