THE UNIVERSITY OF CHICAGO BROAD CLASS PHONEME DETECTION A DISSERTATION SUBMITTED TO THE FACULTY OF THE DIVISION OF THE PHYSICAL SCIENCES IN CANDIDACY FOR THE DEGREE OF DOCTOR OF PHILOSOPHY DEPARTMENT OF COMPUTER SCIENCE BY ZHIMIN XIE CHICAGO, ILLINOIS DECEMBER 2006 UMI Number: 3240138 INFORMATION TO USERS The quality of this reproduction is dependent upon the quality of the copy submitted. Broken or indistinct print, colored or poor quality illustrations and photographs, print bleed-through, substandard margins, and improper alignment can adversely affect reproduction. In the unlikely event that the author did not send a complete manuscript and there are missing pages, these will be noted. Also, if unauthorized copyright material had to be removed, a note will indicate the deletion.
® UMI UMI Microform 3240138 Copyright 2007 by ProQuest Information and Learning Company. All rights reserved. This microform edition is protected against unauthorized copying under Title 17, United States Code. ProQuest Information and Learning Company 300 North Zeeb Road P.
Box 1346 Ann Arbor, MI 48106-1346 To my parents and my wife. ABSTRACT We categorize American English phonemes into several groups: vowel, semi-vowel, nasal, whisper, fricative/affricative, closure/stop, silence and some special phonemes (/q/ and /dx/), among which five main groups (vowel, semi-vowel, nasal, fricative, stop) are fur- ther examined. Thereafter, we construct several detectors based on acoustic features for each phoneme group and compare them with HMM-based systems by testing on contin- uous speech data, TIMIT, and some data in unfavorable environments, like TIMIT with additive noise, and NTIMIT. To detect vowels, a compact vowel detector based only on two acoustic features, peri- odicity and energy, is implemented.
It performs with 86.4% total error rate. Even under some adverse environments, it still works stably. To detect fricatives, several detectors based on SVMs using different acoustic features are constructed and a typical performance of one of these has 90.8% as accuracy and total error rate, respectively. Whereas for stops, features of total energy, energy above 3kHz and Wiener entropy are employed into SVMs and the detector obtains accuracy of 93.2% and total error rate of 19.
All of these results are comparable with or even better than HMM-based systems. However, detectors based on static acoustic features for nasals and semi-vowels do not perform as well as expected. By examining the details of the errors, the associated detection problems are revealed, and inspire a new approach to detection. To deal with non-static features, we propose a combination of HMMs and SVMs for detection of phoneme groups and obtain satisfactory results.
We believe that this method can also be extended for more general speech recognition applications. ili ACKNOWLEDGEMENTS I would like to thank Partha Niyogi for motivating me to think about problems in speech recognition and giving me support through my research with knowledge, advice and re- sources. I really appreciate his guidance and encouragement in designing the speech recog- nition system, which is presented in this thesis. I would also like to thank him for being flexible and understanding during my doctoral research work.
I am very grateful to John Goldsmith and Gina-Anne Levow, who gave instructive suggestions on my research and also kindly served as my committee members. Thanks to Dinoj Surendran for vital assistance on the testing framework and providing many other resources and insightful discussions. Many of my friends have given me strength through these years. I would like to thank Jing Liu, Xuehai Zhang, Yu Hu, Jing Cao, and Vikas Sindhwani.
Last and most important, I offer my deepest thanks to my parents and wife. Without their persistent and tremendous support, I could not have come all the way through this far. iv TABLE OF CONTENTS ABSTRACT ili ACKNOWLEDGEMENTS IV LIST OF FIGURES Vili LIST OF TABLES 1 INTRODUCTION 1. HH HQ kg gà kg 1.2 Statistically Based Recognition.3 Knowledge Based Recognition and Our Approach.3 TIMIT with Additive White Noise.
ee ee ee 2.2 American English Phonemes .3 HMM-based Automatic Speech RecogmzefT. Q Q Q Q HH hà na 3 VOWEL DETECTION AND CLASSIFICATION 3. n n ng k kg kg Nà kg 3. ch HQ kg gà ta 3.
HQ HH Hà ah va 3. HH ee ees 3. Q Q Q HH HQ ng kg kg k kA 3. LH HQ HQ ko 3.6 Performance by Category.1 Insertion Errors by CategOTY.2 Deletion Errors by ContextPhonemes.6 Q Q Q Q HH HH ga kia 3.
Q Q Q Q HQ HH HH gà kg kg ki vo 3. co HH na 3. ch gà kh hà ki ha 3.1 Application to Mandarin Chinese. ee ee ki ha FRICATIVE DETECTION 4.
0 c c Q Q Q LH ng và hà k kg VN ki na 443 FricativeFeatires. LH ha kg ki kg ng 4.4 Landmark Detection with Distinctive Featires.5 SVMs with Distinctive Features. Q Q HQ HH HH hà ở 46. - CO Q ky Q2 va 46.
c c Q Q c LH ng gu và kg kg ki xa STOP DETECTION — TT TŨẶỮẶ. c c c Q LH Q Q ng cv Q k vn V k kg kg 53 StopFeatres. c Q Q Q Q HH HH hà kh kh kg 5.4 Stop Detection with Distinctive Featres. 1 ee kg na 5.
và ke Ha 5. DD NASAL, SEMIVOWEL AND OTHER PHONEMES 6. kh HaHgHaa na.3 Nasal Detection with Distinctive Features .5 Problems with Nasal Detection. và và kg k kh kia 62.22 Performance Based on Acoustic Features.
HH HH hà gà kh kg 6. HQ HH He kh ha 6. Q Q Q Q HH HH gà kg gà ky 7 HYBRID OF HMMS AND SVMS 100 7.1 Probabilistic Outputs forSVMs. ee ee ee 100 7.
Front-ends for HMM-basedSystems. es 106 8 DISCUSSION AND CONCLUSIONS 109 8. va 111 A THE PHONEMES USED IN TIMIT/NTIMIT AND SPHINX AUTOMATIC SPEECH SYSTEM 113 REFERENCES 116 LIST OF FIGURES 2.1 White noise example. The sentence is sa2, “Don’t ask me to carry an oily rag like that.” by a female speaker,faks0.2 Example of spectrogram.
The sentence is sa2, “Don’t ask me to carry an oily rag like that” by a female speaker,faks0.3 Sphinx2 HMM topology. Q QQ HH HH so 13 2.4 3-state no-skip HMM topology.1 Examples of vowel sounds (top: /iy/, /ae/; bottom: /uw/, /er/) .2 Waveform and periodicity of /eh/and/z/. ee ee ee ee 22 3.3 Example of convex hull algorithm .4 Vowel landmark detecionmethod.5 Vowel detection example. The sentence is si2016, “Heave on those ropes; the boat’s come unstuck.” The segments with ’P’ are resultant periodic seg- ments from the first step.
The final detected vowel landmarks are labeled by ’L’, while asterisks indicate possible false landmarks without the first step of periodicity segmentation.6 Another vowel detection example. The sentence is sx49, “At twilight on the twelfth day we’ll have chablis.” The segments with ’P’ are resultant periodic segments from the first step. The final detected vowel landmarks are labeled by ’L’, while asterisks indicate possible false landmarks without the first step of periodicity segmentation.7 Histogram and CDF of vowel pitchperiod.8 Histogram and CDF of vowel duraion.9 Histogram and CDF of vowel periodiclty.10 Histogram and CDF of vowelenergy. ee ee ee 32 3.11 ROC curves for the vowel detector using different energy and periodicity peak-to-dip threshold values.
The asterisk denotes the reference vowel de- tectOl, 5.12 Degradation of totaÌ errOrTafl@S.13 Example of spectrogram. The sentence is sa2, “Don’t ask me to carry an oily rag like that.” by a female speaker, faks0.14 Periodicity of an utterance under different noise environments. The sen- tence is si2016, ”Heave on those ropes; the boat’s come unstuck.1 Spectrograms of some fricatives (top: /s/, /sh/; bottom: /dh/,/v/) .2 Example of features. The utterance is ”She spouted a mouthful of water into the ait” ©.3 Performance of landmark detection.
Stars are performances of Sphinx sys- tems. Solid line with crosses is the performance of the first method. Dashed line with circles is the performance of the second combination method.4 Performance of SVMs. Circle and diamond are performances of Sphinx2 and Sphinx3.
Plus signs with solid line are SVMs using 5 features before and after post-processing. Crosses with dashed line are SVMs using multi- frames of the five features before and after post-processing. Asterisks with dotted lines are SVMs using Mel scale bands before and after post-processing.5 Performance of SVMs based on multiple frames of Mel scale band ratios and other distinctive features before post-processing.6 Spectrograms of /v/ and /dh/. The left graph is /eh/-/v/-/er/ in “never”, and the right is /n/-/dh/-/ix/in “in the’, 2.7 Degradation of deletion and insertion error rates in fricative detection .1 Waveforms and spectrograms of some stops (/k/,/d/) .3 ROC curves of stop deteCHOn.4 Degradation of deletion and insertion error rates in stop detection.1 Spectrograms of some nasals.
From left to right, they are /m/ in /iy/-/m/- /del/ , /n/ in /ow/-/n/-/ae/, and /ng/ in /ih/-/ng/-/gcl/, 6 ww .2 Spectrograms of some semivowels. The top are /r/ in /eh/-/r/-/iy/, and /1/ in /g/-/M-/ay/. The bottom are /w/ in /axr/-/w/-/ao/, and /y/ in /b/-//-/ux/.3 Spectrograms of some whispers. The top left is /hh/ in “her”, the top right is /hh/ in “his”, the bottom left is /hv/ in “she had”, and the bottom right is /hv/in “and haggard”.4 Spectrograms of some flapped stops.
The top-left is /uw/-/dx/-/ih/ in “suit in”, the top-right is /ux/-/dx/-/ux/ in “beautiful”, the bottom-left is /ay/-/dx/- /ax/ in “coincided”, and the bottom-right is /ay/-/dx/-/ax/ in “idly”.5 Spectrograms of some glottal stops. The top-left is /er/-/q/-/ao/, the top- right is /iy/-/q/-/ae/, the bottom-left is /ix/-/q/-/kcl/, and the bottom-right is Isilence/-/q/-/eh/. ee 98 LIST OF TABLES 2.1 The number of speakers, utterances, and phonemes in TIMIT.2 The number of speakers, utterances, and phonemes in NTIMIT .1 Category of vowels by articulatory description.2 Category of vowels by acoustic features. ee ee ee ee 16 3.3 Duration, periodicity and energy of all the phoneme groups from TIMIT training dataset ©.4 Duration, periodicity and energy for each category of vowels in TIMIT training dataset 26.5 Parameter setting for the baseline detector.6 Performance of detection .7 Detection accuracy in each category ofvowels.8 Category of insertion errors.9 Deletion errors by left context categorles.10 Deletion errors by right context categories.
ee eee ees 39 3.11 Periodicity and segmentation of vowels under different environments.12 Performance of the baseline vowel detector after adaptation.13 Performance on switchboard .14 Performance on Mandarin .15 Performance of syllable detection .16 Variations of second formants of/ae/andñy/(.17 Performance of classification at different locations.18 Performance of vowel classification .1 Some acoustic features of different phoneme groups.2 Performance example of landmark detection.3 Performance of our detectors and Sphinx systems.4 Deletion errors by individual fricaives.5 Insertion errors of fricative detection by category .1 Performance of stop detection ©. ee ee es 78 5.2 Deletion errors by individual stops. ee ee ee es 78 5.3 Duration of individual stops ©. ee ee ee 79 5.4 Insertion errors of stop detection by category .1 Performance of nasal detection on distinctivefeatues.2 Insertion errors of nasal detection.3 Duration of all the phoneme groups from TIMIT training dataset .4 Performance of semivowel detection .5 Insertion errors of semivowel đetecion.
ee ee 93 X 71 Performance of the hybrid model.2 Performance improvement of the hybrid model on vowels, nasals, and semivowels 2. 1 HQ Q HH gà gà gà k K kg 105 7.3 Confusion matrix of the hybrdmodel.1 Motivation We will address the problem of pure speech recognition [41]. Pure speech recognition is the task of obtaining a complete or adequate phonological representation directly from the speech signal based purely on the acoustics and additional knowledge of the phonologi- cal aspects of the language, with no linguistic cues from any higher level (i. syntactic, semantic, or pragmatic) modules.
This problem is not artificial and rather at the heart of spoken language processing. We are also motivated by three considerations on the way of pursuing approaches for pure speech recognition.