Mô Hình Dữ Liệu Nhị Phân - Phiên Bản Thứ Hai

Tài liệu nghiên cứu Modelling binary data second edition 1, tổng hợp lý thuyết và thực hành, cung cấp kiến thức chuyên sâu về ., phục vụ nghiên cứu và ứng dụng thực tiễn

Trường đại học

The University of Reading

Chuyên ngành

Applied Statistics

Người đăng

Ẩn danh

Thể loại

textbook

2003

397
5
0

Phí lưu trữ

75 Point

Mục lục chi tiết

1. Introduction

1.2. The scope of this book

1.3. Use of statistical software

1.4. Further reading

2. Statistical inference for binary data

2.1. The binomial distribution

2.2. Inference about the success probability

2.3. Comparison of two proportions

2.4. Comparison of two or more proportions

2.5. Further reading

3. Models for binary and binomial data

3.3. Methods of estimation

3.4. Fitting linear models to binomial data

3.5. Models for binomial response data

3.6. The linear logistic model

3.7. Fitting the linear logistic model to binomial data

3.8. Goodness of fit of a linear logistic model

3.9. Comparing linear logistic models

3.10. Linear trend in proportions

3.11. Comparing stimulus-response relationships

3.12. Non-convergence and overfitting

3.13. Some other goodness of fit statistics

3.14. Strategy for model selection

3.15. Predicting a binary response probability

3.16. Further reading

4. Bioassay and some other applications

4.1. The tolerance distribution

4.2. Estimating an effective dose

4.5. Non-linear logistic regression models

4.6. Applications of the complementary log-log model

4.7. Further reading

5. Model checking

5.1. Definition of residuals

5.2. Checking the form of the linear predictor

5.3. Checking the adequacy of the link function

5.4. Identification of outlying observations

5.5. Identification of influential observations

5.6. Checking the assumption of a binomial distribution

5.7. Model checking for binary data

5.8. Summary and recommendations

5.9. Further reading

6. Overdispersion

6.1. Potential causes of overdispersion

6.2. Modelling variability in response probabilities

6.3. Modelling correlation between binary responses

6.4. Modelling overdispersed data

6.5. A model with a constant scale parameter

6.6. The beta-binomial model

6.8. Further reading

7. Modelling data from epidemiological studies

7.1. Basic designs for aetiological studies

7.2. Measures of association between disease and exposure

7.3. Confounding and interaction

7.4. The linear logistic model for data from cohort studies

7.5. Interpreting the parameters in a linear logistic model

7.6. The linear logistic model for data from case-control studies

7.7. Matched case-control studies

7.8. Further reading

8. Mixed models for binary data

8.1. Fixed and random effects

8.2. Mixed models for binary data

8.4. Mixed models for longitudinal data analysis

8.5. Mixed models in meta-analysis

8.6. Modelling overdispersion using mixed models

8.7. Further reading

9. Exact Methods

9.1. Comparison of two proportions using an exact test

9.2. Exact logistic regression for a single parameter

9.3. Exact hypothesis tests

9.4. Exact confidence limits for βk

9.5. Exact logistic regression for a set of parameters

9.8. Further Reading

10. Some additional topics

10.1. Ordered categorical data

10.2. Analysis of proportions and percentages

10.3. Analysis of rates

10.4. Analysis of binary time series

10.5. Modelling errors in the measurement of explanatory variables

10.6. Multivariate binary data

10.7. Analysis of binary data from cross-over trials

10.8. Experimental design

11. Computer software for modelling binary data

11.1. Statistical packages for modelling binary data

11.2. Interpretation of computer output

11.3. Using packages to perform some non-standard analyses

11.4. Further reading

Preface to the second edition

Preface to the first edition

Appendix A Values of logit(p) and probit(p)

Appendix B Some derivations

Appendix C Additional data sets

References

Index of examples

Index

Tóm tắt

I. Tổng Quan Về Mô Hình Dữ Liệu Nhị Phân Phiên Bản Thứ Hai

Mô hình dữ liệu nhị phân là một công cụ quan trọng trong phân tích thống kê. Phiên bản thứ hai của tài liệu này cung cấp cái nhìn sâu sắc về các phương pháp và ứng dụng của mô hình này. Tài liệu không chỉ cập nhật các kỹ thuật mới mà còn nhấn mạnh tầm quan trọng của việc áp dụng chúng trong thực tiễn.

1.1. Định Nghĩa và Ứng Dụng Của Mô Hình Dữ Liệu Nhị Phân

Mô hình dữ liệu nhị phân được sử dụng để phân tích các biến có hai trạng thái. Các ứng dụng phổ biến bao gồm y tế, kinh tế và xã hội. Việc hiểu rõ mô hình này giúp cải thiện khả năng phân tích dữ liệu.

1.2. Lịch Sử Phát Triển Mô Hình Dữ Liệu Nhị Phân

Mô hình dữ liệu nhị phân đã trải qua nhiều giai đoạn phát triển. Từ những ngày đầu, các nhà nghiên cứu đã tìm ra các phương pháp mới để cải thiện độ chính xác và tính khả thi của mô hình.

II. Thách Thức Trong Phân Tích Dữ Liệu Nhị Phân

Phân tích dữ liệu nhị phân gặp phải nhiều thách thức. Những vấn đề này có thể ảnh hưởng đến độ chính xác của kết quả phân tích. Việc nhận diện và giải quyết các thách thức này là rất quan trọng.

2.1. Vấn Đề Về Độ Chính Xác Trong Mô Hình

Độ chính xác của mô hình có thể bị ảnh hưởng bởi nhiều yếu tố như kích thước mẫu và phương pháp thu thập dữ liệu. Cần có các phương pháp kiểm tra để đảm bảo độ chính xác.

2.2. Tác Động Của Overdispersion

Overdispersion xảy ra khi biến thiên trong dữ liệu lớn hơn so với dự đoán của mô hình. Điều này có thể dẫn đến sai lệch trong kết quả phân tích và cần được xử lý kịp thời.

III. Phương Pháp Phân Tích Dữ Liệu Nhị Phân Hiệu Quả

Có nhiều phương pháp để phân tích dữ liệu nhị phân. Mỗi phương pháp có ưu điểm và nhược điểm riêng. Việc lựa chọn phương pháp phù hợp là rất quan trọng để đạt được kết quả tốt nhất.

3.1. Mô Hình Logistic

Mô hình logistic là một trong những phương pháp phổ biến nhất để phân tích dữ liệu nhị phân. Nó giúp dự đoán xác suất của một sự kiện xảy ra dựa trên các biến độc lập.

3.2. Mô Hình Phân Tích Hồi Quy

Mô hình hồi quy cho phép phân tích mối quan hệ giữa các biến. Phương pháp này có thể được áp dụng để hiểu rõ hơn về các yếu tố ảnh hưởng đến kết quả.

IV. Ứng Dụng Thực Tiễn Của Mô Hình Dữ Liệu Nhị Phân

Mô hình dữ liệu nhị phân có nhiều ứng dụng thực tiễn trong các lĩnh vực khác nhau. Từ y tế đến kinh tế, mô hình này giúp đưa ra các quyết định dựa trên dữ liệu.

4.1. Ứng Dụng Trong Y Tế

Trong y tế, mô hình dữ liệu nhị phân được sử dụng để phân tích các yếu tố ảnh hưởng đến sức khỏe. Điều này giúp cải thiện chất lượng chăm sóc sức khỏe.

4.2. Ứng Dụng Trong Kinh Tế

Trong kinh tế, mô hình này giúp phân tích các quyết định đầu tư và rủi ro. Việc áp dụng mô hình giúp các nhà đầu tư đưa ra quyết định chính xác hơn.

V. Kết Luận Về Mô Hình Dữ Liệu Nhị Phân Phiên Bản Thứ Hai

Mô hình dữ liệu nhị phân là một công cụ mạnh mẽ trong phân tích thống kê. Phiên bản thứ hai của tài liệu này cung cấp cái nhìn sâu sắc và cập nhật về các phương pháp và ứng dụng của mô hình này.

5.1. Tương Lai Của Mô Hình Dữ Liệu Nhị Phân

Tương lai của mô hình dữ liệu nhị phân hứa hẹn sẽ có nhiều cải tiến. Các nghiên cứu mới sẽ tiếp tục mở rộng khả năng ứng dụng của mô hình này.

5.2. Tầm Quan Trọng Của Việc Đào Tạo

Đào tạo về mô hình dữ liệu nhị phân là rất cần thiết. Điều này giúp các nhà nghiên cứu và chuyên gia có thể áp dụng hiệu quả các phương pháp phân tích.

27/07/2025

Trích đoạn nội dung tài liệu

MODELLING BINARY DATA Second Edition CHAPMAN & HALL/CRC Texts in Statistical Science Series Series Editors C. Chatfield, University of Bath, UK Jim Lindsey, University of Liège, Belgium Martin Tanner, Northwestern University, USA J. Zidek, University of British Columbia, Canada Analysis of Failure and Survival Data Epidemiology — Study Design and Peter J. Smith Data Analysis The Analysis and Interpretation of M.

Woodward Multivariate Data for Social Scientists Essential Statistics, Fourth Edition David J. Bartholomew, Fiona Steele, D. Rees Irini Moustaki, and Jane Galbraith A First Course in Linear Model Theory The Analysis of Time Series — Nalini Ravishanker and Dipak K. Dey An Introduction, Fifth Edition Interpreting Data — A First Course C.

Chatfield in Statistics Applied Bayesian Forecasting and Time A. Anderson Series Analysis An Introduction to Generalized A. Harrison Linear Models, Second Edition Applied Nonparametric Statistical A. Dobson Methods, Third Edition Introduction to Multivariate Analysis P.

Collins Applied Statistics — Principles and Introduction to Optimization Methods Examples and their Applications in Statistics D. Everitt Bayesian Data Analysis Large Sample Methods in Statistics A. da Motta Singer Beyond ANOVA — Basics of Applied Markov Chain Monte Carlo — Stochastic Statistics Simulation for Bayesian Inference R. Gamerman Computer-Aided Multivariate Analysis, Mathematical Statistics Third Edition K.

Clark Modeling and Analysis of Stochastic A Course in Categorical Data Analysis Systems T. Kulkarni A Course in Large Sample Theory Modelling Binary Data, Second Edition T. Collett Data Driven Statistical Methods Modelling Survival Data in Medical P. Collett Decision Analysis — A Bayesian Approach J.

Smith Multivariate Analysis of Variance and Repeated Measures — A Practical Elementary Applications of Probability Approach for Behavioural Scientists Theory, Second Edition D. Tuckwell Multivariate Statistics — Elements of Simulation A Practical Approach B. Riedwyl Practical Data Analysis for Designed Statistical Methods for SPC and TQM Experiments D. Yandell Statistical Methods in Agriculture and Practical Longitudinal Data Analysis Experimental Biology, Second Edition D.

Hasted Practical Statistics for Medical Research Statistical Process Control — Theory and D. Altman Practice, Third Edition Probability — Methods and Measurement G. O’Hagan Statistical Theory, Fourth Edition Problem Solving — A Statistician’s B. Lindgren Guide, Second Edition Statistics for Accountants, Fourth Edition C.

Letchford Randomization, Bootstrap and Statistics for Technology — Monte Carlo Methods in Biology, A Course in Applied Statistics, Second Edition Third Edition B. Chatfield Readings in Decision Analysis Statistics in Engineering — S. French A Practical Approach Sampling Methodologies with A. Metcalfe Applications Statistics in Research and Development, Poduri S.

Rao Second Edition Statistical Analysis of Reliability Data R. Kimber, The Theory of Linear Models T. Jørgensen MODELLING BINARY DATA Second Edition David Collett School of Applied Statistics The University of Reading, UK CHAPMAN & HALL/CRC A CRC Press Company Boca Raton London New York Washington, D. CRC Press Taylor & Francis Group 6000 Broken Sound Parkway NW, Suite 300 Boca Raton, FL 33487-2742 © 2003 by Taylor & Francis Group, LLC CRC Press is an imprint of Taylor & Francis Group, an Informa business No claim to original U.

Government works Version Date: 20140219 International Standard Book Number-13: 978-1-4200-5738-6 (eBook - PDF) This book contains information obtained from authentic and highly regarded sources. Reasonable efforts have been made to publish reliable data and information, but the author and publisher cannot assume responsibility for the validity of all materials or the consequences of their use. The authors and publishers have attempted to trace the copyright holders of all material reproduced in this publication and apologize to copyright holders if permission to publish in this form has not been obtained. If any copyright material has not been acknowledged please write and let us know so we may rectify in any future reprint.

Except as permitted under U. Copyright Law, no part of this book may be reprinted, reproduced, transmitted, or utilized in any form by any electronic, mechanical, or other means, now known or hereafter invented, including photocopying, microfilming, and recording, or in any information stor- age or retrieval system, without written permission from the publishers. For permission to photocopy or use material electronically from this work, please access www.com (http://www.com/) or contact the Copyright Clearance Center, Inc. (CCC), 222 Rosewood Drive, Danvers, MA 01923, 978-750-8400.

CCC is a not-for-profit organization that pro- vides licenses and registration for a variety of users. For organizations that have been granted a pho- tocopy license by the CCC, a separate system of payment has been arranged. Trademark Notice: Product or corporate names may be trademarks or registered trademarks, and are used only for identification and explanation without intent to infringe. Visit the Taylor & Francis Web site at http://www.com and the CRC Press Web site at http://www.com To Janet Contents 1 Introduction 1 1.2 The scope of this book 14 1.3 Use of statistical software 15 1.4 Further reading 16 2 Statistical inference for binary data 19 2.1 The binomial distribution 19 2.2 Inference about the success probability 23 2.3 Comparison of two proportions 31 2.4 Comparison of two or more proportions 38 2.5 Further reading 42 3 Models for binary and binomial data 45 3.3 Methods of estimation 50 3.4 Fitting linear models to binomial data 53 3.5 Models for binomial response data 56 3.6 The linear logistic model 58 3.7 Fitting the linear logistic model to binomial data 59 3.8 Goodness of fit of a linear logistic model 65 3.9 Comparing linear logistic models 71 3.10 Linear trend in proportions 78 3.11 Comparing stimulus-response relationships 81 3.12 Non-convergence and overfitting 85 3.13 Some other goodness of fit statistics 87 3.14 Strategy for model selection 91 3.15 Predicting a binary response probability 98 3.16 Further reading 101 4 Bioassay and some other applications 103 4.1 The tolerance distribution 103 4.2 Estimating an effective dose 106 4.5 Non-linear logistic regression models 118 CONTENTS 4.6 Applications of the complementary log-log model 122 4.7 Further reading 128 5 Model checking 129 5.1 Definition of residuals 130 5.2 Checking the form of the linear predictor 135 5.3 Checking the adequacy of the link function 146 5.4 Identification of outlying observations 150 5.5 Identification of influential observations 154 5.6 Checking the assumption of a binomial distribution 168 5.7 Model checking for binary data 169 5.8 Summary and recommendations 185 5.9 Further reading 193 6 Overdispersion 195 6.1 Potential causes of overdispersion 195 6.2 Modelling variability in response probabilities 199 6.3 Modelling correlation between binary responses 201 6.4 Modelling overdispersed data 202 6.5 A model with a constant scale parameter 206 6.6 The beta-binomial model 211 6.8 Further reading 213 7 Modelling data from epidemiological studies 215 7.1 Basic designs for aetiological studies 216 7.2 Measures of association between disease and exposure 219 7.3 Confounding and interaction 223 7.4 The linear logistic model for data from cohort studies 226 7.5 Interpreting the parameters in a linear logistic model 230 7.6 The linear logistic model for data from case-control studies 242 7.7 Matched case-control studies 250 7.8 Further reading 264 8 Mixed models for binary data 269 8.1 Fixed and random effects 269 8.2 Mixed models for binary data 270 8.4 Mixed models for longitudinal data analysis 284 8.5 Mixed models in meta-analysis 291 8.6 Modelling overdispersion using mixed models 293 8.7 Further reading 300 9 Exact Methods 303 9.1 Comparison of two proportions using an exact test 303 9.2 Exact logistic regression for a single parameter 307 CONTENTS 9.3 Exact hypothesis tests 312 9.4 Exact confidence limits for βk 317 9.5 Exact logistic regression for a set of parameters 318 9.8 Further Reading 323 10 Some additional topics 325 10.1 Ordered categorical data 325 10.2 Analysis of proportions and percentages 329 10.3 Analysis of rates 330 10.4 Analysis of binary time series 331 10.5 Modelling errors in the measurement of explanatory variables 331 10.6 Multivariate binary data 332 10.7 Analysis of binary data from cross-over trials 333 10.8 Experimental design 333 11 Computer software for modelling binary data 335 11.1 Statistical packages for modelling binary data 335 11.2 Interpretation of computer output 339 11.3 Using packages to perform some non-standard analyses 341 11.4 Further reading 349 Appendix A Values of logit(p) and probit(p) 351 Appendix B Some derivations 353 B.1 An algorithm for fitting a GLM to binomial data 353 B.2 The likelihood function for a matched case-control study 357 Appendix C Additional data sets 361 C.2 Toxicity of rotenone 361 C.4 Analgesic potency of four compounds 362 C.5 Vasoconstriction of the fingers 363 C.6 Treatment of neuralgia 363 C.9 Cancer of the cervix 367 C.10 Endometrial cancer 367 References 369 Index of examples 379 Index 381 Preface to the second edition The aim of the first edition of this book was to describe the modelling approach to binary data analysis, with an emphasis on practical applications.

That edi- tion was prepared in 1989–90, but in the intervening period there have been a number of important methodological and computational advances. These include the development of techniques for analysing data with more than one level of variation through the use of mixed models, and procedures that lead to an exact version of logistic regression. These methodological advances have been accompanied by developments in statistical computing, with the result that modern computer-based methods for modelling binary data can now be implemented using a wide range of computer packages. This new edition has been prepared so that the text may continue to realise the aims of the first edition, by providing a comprehensive practical guide to statistical methods for use in modelling binary data.

I hope that this book will also continue to meet the needs of statisticians in the pharmaceutical industry; those engaged in agricultural, biological, epidemiological, industrial and medical research; numerate scientists in universities and research institutes, and students fol- lowing undergraduate or postgraduate programmes that feature statistical modelling. The first seven chapters correspond to those in the first edition. Specifically, Chapter 1 introduces a number of example data sets, and statistical proce- dures based on the binomial distribution are described in Chapter 2. Chap- ter 3 introduces the modelling approach, with emphasis being given to the linear logistic model, and Chapter 4 covers bioassay, non-linear logistic mod- els and some other applications.

Model checking diagnostics are described and illustrated in Chapter 5, and the phenomenon of overdispersion is discussed in Chapter 6. Chapter 7 describes the use of linear logistic models in the analysis of data from epidemiological studies, and shows how the estimated parameters in such models can be interpreted in terms of odds ratios. The opportunity has been taken to revise and update the material in each of these chapters. In particular, the emphasis of the first edition on the GLIM software has been eliminated, so that the illustrative examples are now independent of any specific package.

Two chapters have been added. Chapter 8 presents an introduction to mixed models for binary data analysis. The use of these models in multilevel mod- elling, longitudinal data analysis and meta-analysis is considered in detail. Material on the use of mixed models for overdispersion, originally in Chap- ter 6, is also included in this chapter.

Exact methods for modelling binary PREFACE TO THE SECOND EDITION data, which include Fisher’s exact test as a special case, are introduced in Chapter 9. This chapter shows how exact estimates of parameters in a logistic regression model can be found, and how exact hypothesis tests about such parameters can be conducted. By their very nature, the topics considered in Chapters 8 and 9 are a little more sophisticated, and so the mathematical level of these chapters is slightly higher than that of the earlier chapters. Additional topics are covered in Chapter 10, which includes a substantial sec- tion devoted to modelling ordered categorical data.

Chapter 11, on computer software, has been re-written to reflect developments in this area over the last ten years. There are now so many different packages that can be used in modelling binary data that it is no longer practicable to give a compre- hensive guide to the output from them.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ