PREDICTING GENE STRUCTURE IN EUKARYOTIC GENOMES by Jonathan Edward Allen A dissertation submitted to The Johns Hopkins University in conformity with the requirements for the degree of Doctor of Philosophy. Baltimore, Maryland September, 2006 © Jonathan Edward Allen 2006 All rights reserved UMI Number: 3240661 INFORMATION TO USERS The quality of this reproduction is dependent upon the quality of the copy submitted. Broken or indistinct print, colored or poor quality illustrations and photographs, print bleed-through, substandard margins, and improper alignment can adversely affect reproduction. In the unlikely event that the author did not send a complete manuscript and there are missing pages, these will be noted.
Also, if unauthorized copyright material had to be removed, a note will indicate the deletion. ® UMI UMI Microform 3240661 Copyright 2007 by ProQuest Information and Learning Company. All rights reserved. This microform edition is protected against unauthorized copying under Title 17, United States Code.
ProQuest Information and Learning Company 300 North Zeeb Road P. Box 1346 Ann Arbor, MI 48106-1346 Abstract Obtaining the complete set of proteins for each eukaryotic organism is an important step in the quest to understand how life evolves and functions. The complex physiology of eu- karyotic cells, however, makes direct observation of proteins and their parent genes difficult to achieve. An organism’s genome provides the raw data that contains the set of instructions for generating the complete set of proteins, providing the potential to obtain a complete list of proteins without having to rely exclusively on direct observations in the cell.
Computa- tional gene prediction systems, therefore, play an important role in compiling sets of putative proteins for each sequenced genome. This dissertation addresses the problem of computational gene prediction in eukaryotic genomes, presenting a framework for predicting precise single isoform protein coding genes in long contiguous stretches of DNA. The framework is extended to predict overlapping alterna- tively spliced exons in known protein coding regions. A main contribution of this work is to apply classifier stacking with sequential inference, for the first time, to the gene finding prob- lem and to develop a phylogenetic generalized hidden Markov model for the alternative splice site prediction problem.
First a linear weighting scheme is developed, which is extended to _ a statistical prediction model. The statistical model is then transformed to a new sequential inference model to predict alternatively spliced exons. il Prediction accuracy of the single isoform gene prediction methods are tested on three eukaryotic genomes: Arabidopsis thaliana, Oryza sativa and human. Applicatio n of the gene prediction methods are examined in other eukaryotic genomes.
The alternative ly spliced exon prediction model is tested in four Drosophila species under a variety of input conditions. Incorporating multiple sources of gene structure evidence is shown to substantially im- proveme single isoform gene prediction accuracy with performance beginning to rival the accuracy of expert human annotators. Results from the alternative exon prediction experi- ments demonstrate the potential to reliably predict new alternatively spliced forms of known genes. The use of cross-species sequence conservation information is shown to enhance the precision of alternatively spliced exon prediction.
Salzberg Readers: Steven L. Salzberg and Jason M. Eisner ili Acknowledgements I would like to thank my adviser Steven L. Salzberg, for his guidance, patience and support and for giving me the freedom to pursue challenging research problems.
Salzberg has made many helpful suggestions, which improved the quality of my work over the last several years. Thank you to J ason Eisner for informing me of important related work in machine learning and natural language processing. I would also like to thank other members of Dr. Salzberg’s group including Mihaela Pertea and William H.
Majoros with whom I had many enlightening discussions on gene finding work. I also benefited from the many useful discus- sions on bioinformatics topics with other members of the group including Pawel Gajer, Maria D. Delcher and Mihai Pop. Thanks to many people at.
The Institute for Genomic Research which were very helpful in providing useful data to work on including Brian Haas, Bernard Suh, Chunhui Yu, Sam Angioli, Ahwui Wang, Robin Buell, Malcolm Gardner, Jane Carlton, Elodie Ghedon and Brendon Loftus. | I would also like to thank Harold Gainer for advice and providing me with an interesting biological problem to work on and I thank S. Rao Kosaraju for his positive supervision. on this project.
Thank you to Marvin Cook for many productive study sessions, which helped me get more out of many of the courses we took together. Thank you to my wife Safia Ahmed Omar for her love and support and helping me to keep iv my life in proper perspective. Thanks to my family Leah Lewis, Karen Kramer, Wise D. Allen and especially my parents Wise and Joan Allen.
Without their support and encouragement my educational pursuits would not have been possible. Contents Abstract ii Acknowledgements iv List of Tables viii List of Figures | xi 1 Introduction 1.2 Computational Framework for Gene Prediction .1 Generalized Hidden Markov Models.2 Statistical Sequence Modeling .4 Integration of Extrinsic Evidence .5 Using Multiple Genomic Sequences. eee ee 2 Linear Combiner 2. ee 3 Statistical Combiner 3.1 Gene Structure Prediction withaGHMM .2 Representing Gene Structure Evidence.
Conditioned on Input Evidence. ee 4 Prediction of Alternatively Spliced Exons 4. Q Q Q Q Q Q Q ng 2g gà và và và 4.2 Biological Model of Splicing. Q Q LH Q HQ ng Q ng g A kg g vn v v g v va 4.1 Explicit Classification of Cassette Pxons.2 Implicit Prediction of Alternative Splicing .4 Computational Model for Alternative Splicing .2 A Generalized Hidden Markov Model.3 A Phylogenetic Generalized Hidden Markov Model .4 Recovering Exon Structure.và 5 Automated Gene Structure Annotation 5.1 TIGR Annotation Pipeline.2 Gene Structure Annotation Applications .3 Gene Structure Comparison .1 Testing on the ENCODE Regions .2 Evaluation of Evidence Tracks.0 00000 cece eee 6 Alternative Exon Prediction Performance 6.
ng gà vn và và 6.2 Sequence Conservation Patterns.0000 cee eee ee va 7 Conclusion Bibliography Vita vii List of Tables 2.1 Sequence intervals scored by the Linear Combiner. Both strands of a genomic sequence (“+” and “-”) are labeled simultaneously.2 Labels for the sequence intervals between the first and last signal in the se- quence (divided into two tables). Sigg and Sigg mark the index to the left of the start of the sequence (-1) and the right of the sequence respectively. Gene signals are listed along the top columns.1 The set of class labels that describe a local sequence interval used to construct gene models on the positive strand, denoted by the “+” symbol.
The non- coding label applies to both strands. Labels reflect partial and complete exons. Each entry asserts whether the condition in that column must be true (1) or false (0). 10 additional class labels are used to represent strand specific labels on the negative strand.
Q Q Q Q ng v g va va 5.1 Performance of the gene predictors on 1783 genes. SC = Statistical Combiner; SC-g = SC combining gene prediction programs only; LC2 = Linear Combiner using sequence alignments; LC1 = Linear Combiner using gene prediction pro- grams only; GA = GlimmerM; GM = GeneMark.hmm; GS = Genscan+. The columns are: number of whole genes correctly predicted (Correct Gene); num- ber of genes completely missed (Missed Gene); correctly predicted exons out of the 7510 total (Correct Exons); number of exons completely missed (ME); Pre- dicted exons overlapping a gene region but do not overlap a true exon (Inserted Exons); percentage of protein coding nucleotides correctly detected (Nucl Sn).2 Breakdown of combiner predictions when matching exactly 3, 2, 1 or 0 gene prediction programs. The first column (Combiner) refers to the four combiners.
The second column (# of GP) refers to the number matching gene prediction programs. The third column and fourth column count the number of times the combiner prediction is correct (CG) and not entirely correct (WG). The fifth column is the percentage of correct predictions.3 The number of gene models each gene finder exclusively predicts correctly in test set 2.4 Performance for gene predictors including Twinscan and the retrained Glim- merM in addition to the programs listed in Table 5. SC-5: SC using all 5 gene prediction programs; SC-3 = SC using three gene prediction programs; SC-ðg = SC using 5 gene prediction programs and no alignment data; LC2-3 = LC2 using three gene prediction programs; LC1-3 = LC1 using three gene prediction programs; TS = Twinscan; GM2 = newer GlimmerM output.
The three prediction programs used by SC-3, LC2-3 and LC1-3 are Twinscan, Gen- eMark.hmm and newer GlimmerM (GM2).9 Performance comparison of JIGSAW and SC-5 (from Table ð.6 JIGSAW performance in Oryza sativa. Sn (sensitivity) = percentage of test set correctly predicted. Sp (specificity) = percentage of predictions that are cor- rect. Performance measured on three criteria: Genes, Exons and Nucleotides (Nucl).
All results shown as percentages.7 Gene structure comparison. Each entry contains two values “A/M” with A being the average and M being the mean value. The Exon / Intron column is median exon length divided by median intron length, .8 JIGSAW using gene finders and non-Human EST data. Results show sensi- tivity (Sn) and specificity (Sp) measured on Genes, Exons and Nucleotides (Nucl).
All results shown as percentages.9 Results of applying JIGSAW with all available evidence. *KnownGene predicts multiple transcripts per gene locus with a transcript specificity of 47%.10 Comparison of EGASP prediction performance for exons and protein coding nucleotides among the different prediction methods. Sensitivity (Sn) and Speci- ficity (Sp) is given. The F-score is shown for the nucleotide predictions.11 EGASP prediction performance for Genes and gene transcripts (Gene Trans) measuring sensitivity (Sn) and specificity (Sp).
The F-score is given for the Gene predictions. Transcript to Gene Ratio shows the number of transcripts predicted per gene locus. ng ng và và 6. melanogaster annotated di-nucleotides conserved in D.
Di-nucleotides are separated according to splicing type: Acceptor (Acc) and Donor (donor) and splicing event type: alternative (Alt) or constitutive (Con). Pseudo splice sites are included for reference. melanogaster annotated exons missing at least one splice site in D. Percentages are organized by exon type: constitutive exons (CS) cassette exons (CE), exons with multiple splice sites (MS) and exons with intron retention (IR).
The second number associated with the MS and IR rows is the percentage of exons where the non-conserved splice site is constitutive (used in all isoforms).3 Results are shown for 8 versions of ExAlt using different combinations of infor- mant species plus Genscan. The informant species are D. ExAlt-ab initio uses no informant species.4 Prediction performance of ExAlt. Rows 1-2 show ExAlt performance using an input exon and default parameters (ExAlt-Exon) and no informant species (ExAlt-Exon-ab initio).
Rows 3-5 show ExAlt performance using an input coding frame with default parameters (ExAlt-Frame), no informant species (ExAlt-Frame-ab initio), and at most 1 exon predicted per test sequence (ExAlt-Frame-Single). Rows 6- 8 show ExAlt performance using no gene structure information with default parameters (ExAlt-Default), no informant species (ExAlt-Default-ab initio) and at most 1 exon prediction per test sequence (ExAlt-Default-Single).5 Exon prediction accuracy from Table 6.4 separated by exon splicing event.6 ExAlt results on the initial training and testing set in percentages. Included next to each measurement is the difference in percentage points compared to performance in the held out set in Table6. 161 List of Figures 1.1 Aschematic of double stranded DNA.
Each nucleotide is represented by an “L” shaped box and includes a 5 carbon sugar molecule (S), a phosphate residue (P) and one of four nitrogenous bases (a, c, t and g). Dashed lines indicate hydrogen bonds between two bases, c-g pairs form three hydrogen bonds and a-t pairs form two hydrogen bonds. The 3’ and 3’ labels denote the orientation ofeach strand.2 Example of protein coding gene structure. Gene contains two exons and one intron.
The initial exon includes an untranslated region (5’ UTR) and the translated region.