Page v Contents Abbreviations ix Preface xii 1 From genomics to proteomics 1 1.2 The birth of large-scale biology 1 1.3 The genome, transcriptome and proteome 4 1.4 Functional genomics at the DNA and RNA levels 5 1.2 Large-scale mutagenesis 10 1.5 The need for proteomics 13 1.6 The scope of proteomics 17 1.1 Sequence and structural proteomics 17 1.7 The challenges of proteomics 21 2 Strategies for protein separation 23 2.2 Protein separation in proteomics—general principles 23 2.3 Principles of two-dimensional gel electrophoresis 24 2.1 General principles of protein separation by electrophoresis 24 2.2 Separation according to charge but not mass—isoelectric focusing 26 2.3 Separation according to mass but not charge—SDS-PAGE 28 2.4 Two-dimensional gel electrophoresis in proteomics 30 2.1 Limitations of 2DGE in proteomics 30 2.2 Improving the resolution of 2DGE 30 2.3 Improving the sensitivity of 2DGE 32 2.4 The representation of proteins on 2D-gels 34 2.5 The automation of 2DGE 34 2.5 Principles of liquid chromatography in proteomics 35 2.1 General principles of protein and peptide separation by chromatography 35 2.3 Ion exchange chromatography 38 2.4 Reverse-phase chromatography 39 2.5 Size exclusion chromatography 41 2.6 Multidimensional liquid chromatography 42 2.1 Comparison of multidimensional liquid chromatography and 2DGE 42 2.2 Strategies for multidimensional liquid chromatography in proteomics 42 Page ix Abbreviations 2DGE two-dimensional gel electrophoresis 3D-PSSM three dimensional position specific scoring matrix AC affinity chromatography ADP adenosine diphosphate AFM atomic force microscopy ATP adenosine triphosphate BCA bicinchoninic acid BCG Bacille Clamette-Guérin BLAST basic local alignment search tool CATH class, architecture, topology and homologous superfamily CBP calmodulin-binding protein CCD charge-coupled device CD circular dichroism cDNA complementary DNA CDS circular dichroism spectroscopy CE capillary electrophoresis CF chromatofocusing CGE capillary gel electrophoresis CHAPS 3-[(3-Cholamidopropyl)dimethylammonio]-1- propanesulfonate CHS chalcone synthase CID collision-induced dissociation COSY correlation spectroscopy cRNA complementary RNA DALPC direct analysis of large protein complexes DDBJ DNA database of Japan DHFR dihydrofolate reductase DIGE difference gel electrophoresis DNA deoxyribonucleic acid dsRNA double-stranded RNA DTT dithiothreitol EDC N, N' dimethylaminopropylethylcarbodiimide EGF epidermal growth factor EGFR epidermal growth factor receptor EGTA ethylene glycol-bis-(2-aminoethyl)-N, N, N', N' tetraacetic acid eIF eukaryotic initiation factor ELISA enzyme-linked immunosorbent assay EM electron microscopy ER endoplasmic reticulum ESI electrospray ionization EST expressed sequence tag FMN falvin mononucleotide fnII fibronectin type II domain FRET fluorescence resonance energy transfer FT-ICR Fourier transform ion-cyclotron resonance Page xii Preface Proteomics, a word in use for less than a decade, now describes a rapidly growing and maturing scientific discipline, and a burgeoning industry. Proteomics is the global analysis of proteins. It seeks to achieve what other large-scale enterprises in the life sciences cannot: a complete description of living cells in terms of all their functional components, brought about by the direct analysis of those components rather than the genes that encode them. The field of proteomics has grown rapidly in a short time, yet promises to provide more information about living systems than even the genomics revolution that started ten years before.
The reason for this is the richness of proteomics data. Genes have sequences, but proteins have sequences, structures, biochemical and physiological functions, and their activities are influenced by chemical modification, localization within or without the cell, and perhaps most importantly of all, their interactions with other molecules. If genes are the instruction carriers, proteins are the molecules that execute those instructions. Genes are the instruments of change over evolutionary timescales, but proteins are the molecules that define which changes are accepted and which are discarded.
It is from proteins that we shall learn how living cells and organisms are built and maintained, and how they fail when things go wrong. As is the case for any emerging scientific field, proteomics makes a lot of sense to those performing large-scale protein analysis on a day-to-day basis, and much less sense to those looking in from the outside. Proteomics abounds with jargon and acronyms. New technologies and variations appear on what can seem to be a daily basis.
It can be difficult to keep up, and even specialists in one area of proteomics sometimes have difficulties applying their knowledge in other specialized areas. It is my hope that this book will be useful to those who need a broad overview of proteomics and what it has to offer. It is not meant to provide expertise in any particular area: there are plenty of books on electrophoresis, mass spectrometry, bioinformatics etc. for the reader needing detailed treatment of particular technologies.
However, this book pulls together disparate information concerning the different proteomics technologies and their applications, and presents them in what I hope is a simple and user-friendly manner. After a brief introductory chapter, the various proteomics technologies are discussed in more detail: two-dimensional gel electrophoresis, multidimensional liquid chromatography, mass spectrometry, sequence analysis, structural analysis, methods for studying protein interactions, modifications, localization and function. Protein chips, an emerging and promising recent addition to the proteomics armory, are described in the penultimate chapter. The final chapter presents a few examples of how proteomics is being applied, particularly in the medical and pharmaceutical fields.
Again, this is not intended to be comprehensive coverage, but is provided so the reader has an overview of the scope of proteomics and its potential. At the end of each chapter is a short bibliography, containing some classic papers and useful reviews for those wanting to delve deeper into the subject. I have assumed that the reader has a working knowledge of molecular biology and biochemistry. This book would not have been possible without the help and support of many people, not least the team at Garland/BIOS for their patience, persistence and optimism in the face of tight deadlines.
I’d like to thank the many friends and colleagues who offered opinions on the individual chapters and pointed out potential errors or omissions, and in particular, I would like to thank all at the Fraunhofer Institute of Molecular Biology and Page 1 1 From genomics to proteomics 1.1 Introduction Proteomics is a rapidly growing area of molecular biology that is concerned with the systematic, large-scale analysis of proteins. It is based on the concept of the proteome as a complete set of proteins produced by a given cell or organism under a defined set of conditions. Proteins are involved in almost every biological function, so a comprehensive analysis of the proteins in the cell provides a unique global perspective on how these molecules interact and cooperate to create and maintain a working biological system. The cell responds to internal and external changes by regulating the level and activity of its proteins, so changes in the proteome, either qualitative or quantitative, provide a snapshot of the cell in action.
The proteome is a complex and dynamic entity that can be defined in terms of the sequence, structure, abundance, localization, modification, interaction and biochemical function of its components, providing a rich and varied source of data. The analysis of these various properties of the proteome requires an equally diverse range of technologies. This introductory chapter considers the importance of proteomics in the context of systems biology, discusses some of the major goals of proteomic analysis and introduces the major technology platforms. We begin by tracing the origins of proteomics in the genomics revolution of the 1990s and following its evolution from a concept to a mainstream technology with a current market value of over $1.2 The birth of large-scale biology The overall goal of molecular biology research is to determine the functions of genes and their products, allowing them to be linked into pathways and networks, and ultimately providing a detailed understanding of how biological systems work.
For most of the last 50 years, research in molecular biology has focused on the isolation and characterization of individual genes and proteins because there was neither the information nor the technology available for larger scale investigations. The only way to study biological systems was to break them down into their components, look at these individually, and attempt to reassemble each system from the bottom up. This approach is known as reductionism, and it dominated the molecular life sciences until the early 1990s. The face of biological research began to change in the 1990s as technological breakthroughs made it possible to carry out large-scale DNA sequencing.
Until this point, the sequences of individual genes and proteins had accumulated slowly and steadily as researchers cataloged their new discoveries. This can be seen from the steady growth in the Page 2 GenBank sequence database from 1980–1990 (Figure 1. The 1990s saw the advent of factory- style automated DNA sequencing, resulting in a massive explosion of sequence data (Figure 1. In the early 1990s, much of the new sequence data was represented by expressed sequence tags (ESTs), short fragments of DNA obtained by the random sequencing of cDNA libraries.
In 1995, the first complete cellular genome sequence was published, that of the bacterium Haemophilus influenzae. In the next few years, over 100 further genome sequences were completed, including our own human genome which was essentially finished in 2003. The large-scale sequencing projects ushered in the genomics era, which effectively removed the information bottleneck and brought about the realization that biological systems, while large and very complex, were ultimately finite. The idea began to emerge that it might be possible to study biological systems in a holistic manner simply by cataloging and enumerating the components if sufficient amounts of data could be collected and analyzed.
Unfortunately, while the technology for genome sequencing had advanced rapidly, the technology for studying the functions of the newly discovered genes lagged far behind. The sequence databases became clogged with anonymous sequences and gene fragments, and the problem was exacerbated by the Figure 1.1 Growth of the GenBank database in its first 20 years. Courtesy of GenBank. Page 3 unexpectedly large number of new genes found even in well-characterized organisms.
As an example, consider the bakers’ yeast Saccharomyces cerevisiae, which was thought to be one of the best-characterized model organisms prior to the completion of the genome-sequencing project in 1996. Over 2000 genes had been characterized in traditional experiments and it was thought that genome sequencing would identify at most a few hundred more. Scientists got a shock when they found the yeast genome contained over 6000 genes, nearly a third of which were unrelated to any previously identified sequence. Such genes were described as orphans because they could not be assigned to any classical gene family (Figure 1.
The availability of masses of anonymous sequence data for hundreds of different organisms has precipitated a number of fundamental changes in the way research is conducted in the molecular life sciences. Traditionally gene function had been studied by moving from phenotype to gene, an approach sometimes called forward genetics. An observed mutant phenotype (or purified protein) was used as the starting point to map and identify the corresponding gene, and this led to the functional analysis of that gene and its product. The opposite approach, sometimes termed reverse genetics, is to take an uncharacterized gene sequence and modify it to see the effect on phenotype.
As more uncharacterized sequences have accumulated in databases, the focus of research has shifted from forward to reverse genetics. Similarly, most research prior to 1995 was hypothesis- driven, in that the researcher put forward a hypothesis to explain a given observation, and then designed experiments to prove or disprove it. The genomics revolution instigated a progressive change towards discovery-driven research, in which the components of the system under investigation are collected irrespective of any hypothesis about how they might work. The final paradigm shift concerns the sheer volume of data generated in today’s experiments.
Whereas in the past researchers have focused on individual gene products and generated rather small amounts of data, Figure 1.2 Distribution of yeast genes by annotation status in the aftermath of the Saccharomyces cerevisiae genome project. (?? shows questionable open reading frames.