BOSTON UNIVERSITY COLLEGE OF ENGINEERING Dissertation ANALYSIS OF RIBOSOMAL PROTEIN BLOCK STRUCTURE: FUNCTIONAL CHARACTERIZATION, EVOLUTIONARY IMPLICATIONS AND DISTANT HOMOLOGY SEARCH USING DISCRETE STATE MODELS PAOLA FAVARETTO Laurea Degree, Universita’ degli Studi di Padova, Italy, 2001 Submitted in partial fulfillment of the requirements for the degree of Doctor of Philosophy 2007 UMI Number: 3246605 INFORMATION TO USERS The quality of this reproduction is dependent upon the quality of the copy submitted. Broken or indistinct print, colored or poor quality illustrations and photographs, print bleed-through, substandard margins, and improper alignment can adversely affect reproduction. In the unlikely event that the author did not send a complete manuscript and there are missing pages, these will be noted. Also, if unauthorized copyright material had to be removed, a note will indicate the deletion.
® UMI UMI Microform 3246605 Copyright 2007 by ProQuest Information and Learning Company. All rights reserved. This microform edition is protected against unauthorized copying under Title 17, United States Code. ProQuest Information and Learning Company 300 North Zeeb Road P.
Box 1346 Ann Arbor, MI 48106-1346 Approved by First Reader: CLG 7 GR 22 Temple F. Professor of Biomedical Engineering Second Reader: Ses C Dh. Scott Mohr, Ph. Professor of Chemistry Third Reader: Hyman Hartman, Ph.
Professor of Biology Fourth Reader: — Sandor Vajda, Ph. kẽ Professor of Biomedical Engineering Acknowledgments I would like to express deep gratitude to my advisor, Professor Temple F. Smith, for his guidance and support during the completion of this project. His unending drive and extraordinary enthusiasm for science have been of great inspiration and motivation in my work.
Iam grateful to Professor Scott C. Mohr for his thoughtful advice and encouraging help, especially in the early stage of the project. His meticulous work and rigorous attitude have shown me the importance of details and analytical thinking. Iam indebted to Professor Hyman Hartman for his important contribution to the project and his constant interest in my work.
His comments and insights were always very much appreciated. I would like to acknowledge Professor Sandor Vajda and Professor Lucia M. Vaina for serving on my Ph. committee, and Professor Jadwiga Bienkowska for serving on the prospectus committee.
My appreciation goes to all the members of the BioMolecular Engineering Research Cen- ter for their friendship and numerous stimulating conversations: Arjun Bhutkar, Prashant Vishwanath, Kavitha Venkatesan, Filiz Aslan, Hongxian He, Esther Epstein, Nancy Sands and Sean Quinlan. I would also like to thank my parents for always believing in me and for giving me the HH opportunity to pursue my goals. Finally, I would like to thank my husband Atul for his profound support and immense help throughout the completion of this degree. He has been a constant source of strength and encouragement, and I am honored to have him in my life.
iv ANALYSIS OF RIBOSOMAL PROTEIN BLOCK STRUCTURE: FUNCTIONAL CHARACTERIZATION, EVOLUTIONARY IMPLICATIONS AND DISTANT HOMOLOGY SEARCH USING DISCRETE STATE MODELS (Order No. ) PAOLA FAVARETTO Boston University College of Engineering, 2007 Major Professor: Temple F. Smith, Professor of Biomedical Engineering ABSTRACT The ribosome, a very complex molecular machine, plays a fundamental role in all liv- ing organisms and exhibits extraordinary engineering design concepts. This investigation examined the complexity of the translational apparatus, seeking to understand how its com- ponents evolved to their present configuration.
It includes detailed sequence and structural comparative analyses of the ribosome and its associated proteins with functional character- ization of these components across the three phylogenetic domains. Amino acid sequence alignments of ribosomal proteins revealed an unusual taxon-specific block structure, with some blocks universally conserved and others specific to one or two phylogenetic domains. Statistical and phylogenetic analyses of the universal blocks imply that modern Bacteria, Archaea and Eukarya clearly have a common ancestor, while the phylodomain-specific blocks suggest that these groups also share more recent, taxon-specific cenancestors. Major evolutionary implications of the observed block structure are: (7) the crenarchaeal, endosymbiotic origin of the modern eukaryotic translational apparatus; and (it) the occurrence of a prokaryotic bottleneck that drastically reduced the diversity of modern species progenitors about 2.2 billion years ago.
Surprisingly, the highly conserved blocks identified in most of the translation-related proteins do not associate consistently with any identifiable particular function or structural feature, or even with rRNA contacts. A comprehensive investigation of the rRNA-ribosomal protein interactions, however, demonstrated a major role of ribosomal proteins in constrain- ing the rRNA conformational space and stabilizing its correct, universally conserved core fold. In order to identify possible evolutionary relationships between the taxon-specific block structure and other proteins, a new stochastic tool for the identification of distant ho- mologous domains in single-, repeated- and multi-domain contexts was implemented. The approach uses sequence and structure information embedded in Discrete State Models, and a Markov threading technique to estimate the compatibility of any query sequence with the models under consideration.
The method was successfully applied to a variety of cases, in- cluding the ribosomal blocks, the WD40-repeat domain and the very diverse ubiquitin-like family. vi Contents Acknowledgements Hài Abstract List of Tables vii List of Figures XV 1 Introduction 1.1 Motivation and Significance.3 Protein Translation Apparatus. vii 2 Taxon-specific Block Structure in Ribosomal Proteins 10 2. c c c c Q Q Q n ng kg kg gi kg kia 10 2.
ng kg g kg kg cv kia 11 2. HQ kg ha 13 2. và kg ki à kia a 15 2. nà gà cv v k kg k k va 29 2.
LH gà kg va 48 2/71 Ribosomal Phylogeny Reconstruction .004 51 3 Functional Analysis of Ribosomal Proteins 54 3.1 Structural Role of Ribosomal Proteins .2 RNA-chaperone Activity of Ribosomal Proteins .3 Consensus Secondary Structure of rRNA.4 RNA Secondary-structure Prediction Software: mfold .5 Fold Prediction Scoring Scheme.1 Individual Protein Constraints .2 Binding Pathway Constraints .3 Artificial and Hypothetical Protein Constraints .4 Conclusions: Role of Ribosomal Proteins in rRNA Folding.7 rRNA Base-pair Substitutions.2 Conclusions: Sequence Composition Impact on rRNA Folding.9 Conclusions: Constraining rRNA Conformational Space. 87 Discrete State Models for Distant Homology Search of Ribosomal Blocks 89 4.2 Discrete State-space Models (DSMs) .1 Creation of Discrete State Models .2 Distant Homology Search: Problem Formulation .4 Method and Models Validation .1 Identification of Repeated-domain Proteins: WD40-repeat Family 44.2 WD40-repeat Model Validation.3 Identification of Small Domain Proteins and Concatenated Domains: Ubiquitin 2.5 Discrete State Models for Ribosomal Protein Blocks. ST aGLL----a áaa đa Discussion 5.1 Evolutionary Implications of the Ribosomal Protein Block Structure .2 The Archaeal Origin of the Eukaryotic Translational Apparatus.3 Ribosomal Block Structure and Interactions with rRNA .4 Protein Domain Identification Using DSMs. Ặ g ga aaAaAaaäa AT HH PDB Structures of Proteins in the Translational Apparatus 149 rRNA and Ribosomal Protein Interactions 152 Phylogenetic Reconstruction Parameters C.00000 pee ee D WD40-repeat Prediction in S.
cerevisiae 162 E PDB Structures of Ubiquitin Family 166 Bibliography 167 Vita 185 xi List of Tables 2.1 List of eukaryotic, archaeal and bacterial species used in the analysis .2 Amino acid similarity classes used in the refinement of the multi-sequence alignments.3 Ribosomal proteins in the 30S subunit and their phylogenetic assignment 16 2.4 Ribosomal proteins in the 50S subunit and their phylogenetic assignment 17 2.5 Structural folds of ribosomal proteins and their spatial location in the ribo- somal subunits .6 Interactions between ribosomal proteins and RNA helices in the 30S subunit 45 2.7 Interactions between ribosomal proteins and RNA helices in the 50S subunit 46 3.1 Examples of binding pathways and their impact on the conformational vari- ability of 16S and accuracy of predicted folds .2 Examples of binding pathways and their impact on the conformational vari- ability of 23S and accuracy of predicted folds .3 Domain variability in the 30S and 50Ssubunits.1 Generic DSMs used as competing models in the Bayesian estimation of pos- terior probabilltly.2 Comparison of predicted repeat boundaries for sequence PDB code 1GXRA 113 4.3 Ubiquitin-like protein subfamilies.4 Ubiquitin and ubiquitin-like model classes.5 Ubiquitin-like domain occurrences identified in a sample of representative complete eukaryotic gnomes. ee ee va 123 4.6 Ubiquitin-like domain occurrences identified in a sample of representative archaeal genome@s. HQ ng kg gi k k k kg sa 124 4.7 Ubiquitin-like domain occurrences identified in a sample of representative bacterial gnomes. 125 AI List of translational protein structures in the Protein Data Bank (PDB) .1 Distribution of rRNA-ribosomal protein interaction (RPI) patterns .2 Distribution of rRNA-ribosomal protein interaction across the amino acid types associated with residues in the universal blocks.3 Distribution of hydrogen-bond interactions between rRNA and ribosomal proteins (RPI)across block types .4 Distribution of all interactions within 3.7 A between rRNA and ribosomal proteins across block types.
ga kg ki kA 155 B.5 Distribution of rRNA-ribosomal protein interaction across the amino acid types associated with residues in non-universal blocks .6 Distribution of rRNA-ribosomal protein interaction across the amino acid types and the interaction types.7 Distribution of individual amino acids in the proteins of the small ribosomal subunit 2.00 ee xa 158 B8 Distribution of individual amino acids in the proteins of the large ribosomal subunit 2. ee 159 D1 Results of WD40-repeat search on the entire S. cerevisiae genome, identified by both the DSM method and the profile-based approach .2 Results of WD40-repeat search on the entire S. cerevisiae genome, identified only by the DSM method .1 List of PDB structures for ubiquitin-like proteins.
XỈV List of Figures 2.1 Ribosomal proteins distribution among the three phylogenetic domains.2 Schematic representation of the multiple alignments in the universal riboso- mai protein set (UN) of the 305 subunit.3 Schematic representation of the multiple alignments in the universal riboso- mal protein set (UN) of the 50S subunit.4 Schematic representation of the multiple alignments in the archaeal-eukaryotic ribosomal protein set (AE) of the 30S subunit.5 Schematic representation of the multiple alignments in the archaeal-eukaryotic ribosomal protein set (AP) of the 508 subunit.6 Length distribution of the taxon-specific blocks identified in the ribosomal proteins ©. ki kg kg va 23 2.7 Average number of blocks in ribosomal proteins.8 Block distribution among the ribosomal proteins .9 Average number of conserved positions in subsets of universal ribosomal pro- 2.10 Average number of informative positions in subsets of universal ribosomal 57500.11 Distribution of rRNA-ribosomal protein interactions (RPI) among block types 36 2.12 Distribution of rRNA-ribosomal protein interaction (RPI) types in both ri- bosomal subunits. c c cv cu 2 kg L k v va 37 2.13 Distributions of rRNA-ribosomal protein interactions among amino acids in ribosomal proteins .14 rRNA-ribosomal protein interactions mapped into the secondary structure of the 238 rRNA of H.15 rRNA-ribosomal protein interactions mapped into the secondary structure of the 238 rRNA of H.16 rRNA-ribosomal protein interactions mapped into the secondary structure of 16S rRNA of T. c c c c cv v g g g Q gà cv ga và va 44 2.17 Phylogenetic reconstructions using the positional variation among aligned subsets of ribosomal proteins .1 Consensus secondary-structure diagram for 16S rRNA .2 Consensus secondary-structure diagram for 238 rRNA (5’end) .3 Consensus secondary-structure diagram for 238 rRNA (3’ end) .4 Number of folds predicted when individual ribosomal protein constraints were applied to the FE.5 Number of folds predicted when individual ribosomal protein constraints were applied to the H.6 Assembly map for the small ribosomal subunit (308) .7 Assembly map for the large ribosomal subunit (50S) .8 Validation of results using artificial and hypothetical protein constraints .9 Correlation between the number of predicted folds and the number of applied constraints 2.10 Average percentage of native base pairs predicted correctly for 16S È.
cob sequences with canonical base-pair substitutions .11 Average total number of base pairs predicted for sequences with canonical base-pair substitutions .12 Average number of secondary-structure folds predicted within 5% of the min- imum free-energy fold for &. coli 16S sequences with canonical base-pair substitutions 2. cu cu ca v.v kg lv kg va 82 3.13 Average percentage of native base pairs predicted correctly for 23S H, maris- mortui sequences with canonical and non-canonical base-pair substitutions 83 Xvii 3.