Phân Tích Dữ Liệu Kinh Doanh và Khoa Học Dữ Liệu cho Các Vấn Đề Kinh Doanh

Khám phá cách phân tích dữ liệu và khoa học dữ liệu giải quyết các vấn đề kinh doanh hiệu quả, nâng cao quyết định và tối ưu hóa quy trình.

Trường đại học

Data Analytics Corp.

Chuyên ngành

Business Analytics

Người đăng

Ẩn danh

Thể loại

book

2021

416
2
0

Phí lưu trữ

75 Point

Mục lục chi tiết

Preface

Acknowledgments

1. Part I Beginning Analytics

1.1. Introduction to Business Data Analytics: Setting the Stage

1.2. Types of Business Problems

1.3. The Role of Information in Business Decision Making

1.4. The Data-Information Nexus

1.4.1. Data and Information Confusion

1.4.2. The Data Component

1.4.3. The Extractor Component

1.4.4. The Information Component

1.5. Data Sources, Organization, and Structures

1.5.1. Data Dimensions: A Taxonomy for Defining Data

1.5.1.1. Taxonomy Component #1: Source
1.5.1.2. Taxonomy Component #2: Domain
1.5.1.3. Taxonomy Component #3: Levels
1.5.1.4. Taxonomy Component #4: Continuity
1.5.1.5. Taxonomy Component #5: Measurement Scale

1.5.2. External Database Structures

1.5.3. Internal Database Structures

1.6. Basic Data Handling

1.6.1. Case Study 1: Customer Transactions Data

1.6.2. Case Study 2: Measures of Order Fulfillment

1.6.3. Importing Your Data

1.6.3.1. Importing a CSV Text File into Pandas
1.6.3.2. Importing Large Files in Chunks
1.6.3.3. Checking Your Imported Data

1.6.4. Merging or Joining DataFrames

1.6.4.1. Boolean Operators and Indicator Functions
1.6.4.2. Pandas Query Method

1.7. Data Visualization: The Basics

1.7.1. Background for Data Visualization

1.7.2. Gestalt Principles of Visual Design

1.7.3. Issues Complicating Data Visualization

1.7.3.1. Human Visual Limitations
1.7.3.2. Data Visualization Tools
1.7.3.3. Types of Visuals
1.7.3.4. What to Look for in a Graph

1.7.4. Visualizing Spatial Data

1.7.4.1. Visualizing Continuous Spatial Data
1.7.4.2. Visualizing Categorical Spatial Data
1.7.4.3. Visualizing Continuous and Categorical Spatial Data

1.7.5. Visualizing Temporal (Time Series) Data

1.7.5.1. Properties of Temporal (Time Series) Data
1.7.5.2. Visualizing Time Series Data
1.7.5.3. Times Series Complications

1.7.6. Taylor Series Expansion for Growth Rates

1.8. Advanced Data Handling: Preprocessing Methods

1.8.1. A Family of Transformations

1.8.2. Dummy or One-Hot Encoding

1.8.3. Handling Missing Data

1.8.4. Mean and Variance of Standardized Variable

1.8.5. Mean and Variance of Adjusted Standardized Variable

1.8.6. Unbiased Estimators of μ and σ 2

2. Part II Intermediate Analytics

2.1. OLS Regression: The Basics

2.1.1. Basic OLS Concept

2.1.1.1. The Disturbance Term and the Residual
2.1.1.2. The Gauss-Markov Theorem

2.1.2. Analysis of Variance

2.1.2.1. Basic OLS Regression
2.1.2.2. The Log-Log Model
2.1.2.3. Model Set-up
2.1.2.4. ANOVA for Basic Regression

2.1.3. Basic Multiple Regression

2.1.3.1. ANOVA for Multiple Regression
2.1.3.2. Alternative Measures of Fit: AIC and BIC

2.1.4. Case Study: Expanded Analysis

2.1.5. Predictive Analysis: Introduction

2.1.5.1. Simulation Tool for Prediction Application

2.2. Time Series Analysis

2.2.1. Time Series Basics

2.2.1.1. Time Series Definition
2.2.1.2. Time Series Concepts

2.2.2. Importing a Date/Time Variable

2.2.3. The Data Cube and Time Series Data

2.2.4. Handling Dates and Times in Python and Pandas

2.2.5. Aggregating Datetime Measures

2.2.6. Converting Time Periods in Pandas

2.2.7. Date-Time Mini-Language

2.2.8. Some Calendrical Calculations

2.2.9. Time Series Generation Process: AR(1) Model

2.2.10. Visualization for AR(1) Detection

2.2.11. Durbin-Watson Test Statistic

2.2.12. Lagged Dependent and Independent Variables

2.2.12.1. Lagged Independent Variable: ARDL(0, 1)
2.2.12.2. Lagged Dependent Variable: ARDL(1, 0)
2.2.12.3. Lagged Dependent and Independent Variables: ARDL(1, 1)

2.2.13. Further Exploration of Time Series Analysis

2.2.13.1. Step 1: Identification of a Model
2.2.13.2. Step 2: Estimation of the Model
2.2.13.3. Step 3: Validation of the Model
2.2.13.4. Step 4: Forecasting with the Model

2.2.14. Useful Algebra Results

2.2.15. Mean and Variance of Yt

2.2.16. Time Trend Addition

2.3. Extending the Cross-tab

2.3.1. Creating a Frequency Table

2.3.2. Hypothesis Testing: A First Step

2.3.3. Cross-tabs and Hypothesis Tests

2.3.4. Plotting a Frequency Table

2.3.5. Pearson Chi-Square Statistic

3. Part III Advanced Analytics

3.1. Advanced Data Handling for Business Data Analytics

3.1.1. Supervised and Unsupervised Learning

3.1.2. Working with the Data Cube

3.1.3. The Data Cube and DataFrame Indexing

3.1.4. Sampling From a DataFrame

3.1.4.1. Simple Random Sampling (SRS)
3.1.4.2. Stratified Random Sampling
3.1.4.3. Cluster Random Sampling

3.1.5. Index Sorting of a DataFrame

3.1.6. Splitting a DataFrame: The Train-Test Splits

3.1.6.1. Model Tuning of Hyperparameters
3.1.6.2. Incorrect Use of Testing Data
3.1.6.3. Creating the Training/Testing Data Sets
3.1.6.4. Recombining the Data Sets

3.1.7. Primer on Random Numbers

3.2. Advanced OLS for Business Data Analytics

3.2.1. Link Functions: An Introduction

3.2.2. Data Standardization for Regression Analysis

3.2.3. One-Hot and Effects (or Sum) Encoding

3.2.4. Case Study Application

3.2.5. Heteroskedasticity Issues and Tests

3.2.5.1. Digression on Multicollinearity
3.2.5.2. Detection with VIF and the Condition Index
3.2.5.3. Principal Component Regression and High-Dimensional Data

3.2.6. Predictions and Scenario Analysis

3.2.6.1. Prediction Error Analysis (PEA)

3.2.7. Panel Data Models

3.3. Classification with Supervised Learning Methods

3.3.1. Case Study: Background

3.3.2. Properties of this Problem

3.3.3. A Model for the Binary Problem

3.3.4. Case Study: Train-Test Data Split

3.3.5. Case Study: Logit Model Training

3.3.6. Making and Assessing Predictions

3.3.7. Classification with a Logit Model

3.3.7.1. Case Study: Predicting

3.3.8. Background: Bayes Theorem

3.3.9. The Naive Adjective: A Simplifying Assumption

3.3.10. Case Study: Naive Bayes Training

3.3.11. Decision Trees for Classification

3.3.11.1. Partitioning by Constants
3.3.11.2. Gini Index and Entropy
3.3.11.3. Case Study: Growing a Tree
3.3.11.4. Case Study: Predicting with a Tree

3.3.12. Support Vector Machines

3.3.12.1. Case Study: SVC Application
3.3.12.2. Case Study: Prediction

3.3.13. Classifier Accuracy Comparison

3.4. Grouping with Unsupervised Learning Methods

3.4.1. Training and Testing Data Sets

3.4.2. Forms of Hierarchical Clustering

3.4.3. Agglomerative Algorithm Description

3.4.4. Metrics and Linkages

3.4.5. Case Study Application

3.4.6. Examining More than One Solution

3.4.7. Case Study Application

3.4.8. Mixture Model Clustering

List of Figures

Tóm tắt

I. Giới thiệu về Phân Tích Dữ Liệu Kinh Doanh Khám Phá Khoa Học Dữ Liệu

Phân tích dữ liệu kinh doanh là một lĩnh vực quan trọng trong việc ra quyết định. Nó sử dụng các phương pháp khoa học dữ liệu để biến dữ liệu thành thông tin có giá trị. Việc hiểu rõ về phân tích dữ liệu giúp các doanh nghiệp tối ưu hóa quy trình và nâng cao hiệu quả hoạt động. Trong bối cảnh hiện đại, khoa học dữ liệu trở thành một công cụ không thể thiếu cho các nhà quản lý và nhà phân tích.

1.1. Tầm Quan Trọng của Phân Tích Dữ Liệu trong Kinh Doanh

Phân tích dữ liệu giúp doanh nghiệp hiểu rõ hơn về thị trường và khách hàng. Nó cung cấp cái nhìn sâu sắc về xu hướng và hành vi tiêu dùng, từ đó hỗ trợ việc ra quyết định chiến lược.

1.2. Các Loại Dữ Liệu Kinh Doanh Thường Gặp

Dữ liệu kinh doanh có thể được phân loại thành nhiều loại, bao gồm dữ liệu định lượng và định tính. Việc phân loại này giúp xác định phương pháp phân tích phù hợp.

II. Thách Thức Trong Phân Tích Dữ Liệu Kinh Doanh Những Vấn Đề Cần Giải Quyết

Mặc dù phân tích dữ liệu mang lại nhiều lợi ích, nhưng cũng tồn tại nhiều thách thức. Các vấn đề như dữ liệu không đầy đủ, dữ liệu không chính xác và khó khăn trong việc tích hợp dữ liệu từ nhiều nguồn khác nhau là những trở ngại lớn. Những thách thức này cần được nhận diện và giải quyết để tối ưu hóa quy trình phân tích.

2.1. Dữ Liệu Không Đầy Đủ và Không Chính Xác

Dữ liệu không đầy đủ có thể dẫn đến những quyết định sai lầm. Việc đảm bảo chất lượng dữ liệu là rất quan trọng trong quá trình phân tích.

2.2. Khó Khăn Trong Tích Hợp Dữ Liệu

Tích hợp dữ liệu từ nhiều nguồn khác nhau có thể gây khó khăn. Các doanh nghiệp cần có chiến lược rõ ràng để xử lý vấn đề này.

III. Phương Pháp Phân Tích Dữ Liệu Kinh Doanh Các Kỹ Thuật Hiệu Quả

Có nhiều phương pháp để thực hiện phân tích dữ liệu kinh doanh. Các kỹ thuật như hồi quy, phân tích chuỗi thời gian và phân tích thống kê là những công cụ quan trọng. Việc lựa chọn phương pháp phù hợp sẽ giúp tối ưu hóa kết quả phân tích.

3.1. Hồi Quy Phân Tích Mối Quan Hệ Giữa Các Biến

Hồi quy là một trong những phương pháp phổ biến nhất trong phân tích dữ liệu. Nó giúp xác định mối quan hệ giữa các biến và dự đoán kết quả.

3.2. Phân Tích Chuỗi Thời Gian Dự Đoán Xu Hướng Tương Lai

Phân tích chuỗi thời gian giúp doanh nghiệp dự đoán xu hướng trong tương lai dựa trên dữ liệu lịch sử. Đây là một công cụ mạnh mẽ trong việc lập kế hoạch và ra quyết định.

IV. Ứng Dụng Thực Tiễn của Phân Tích Dữ Liệu Kinh Doanh Kết Quả Nghiên Cứu

Nhiều doanh nghiệp đã áp dụng phân tích dữ liệu để cải thiện hiệu suất và tăng trưởng. Các nghiên cứu cho thấy rằng việc sử dụng khoa học dữ liệu có thể dẫn đến tăng trưởng doanh thu và cải thiện sự hài lòng của khách hàng. Những ứng dụng này chứng minh giá trị của khoa học dữ liệu trong môi trường kinh doanh hiện đại.

4.1. Cải Thiện Quy Trình Kinh Doanh Thông Qua Phân Tích

Nhiều doanh nghiệp đã cải thiện quy trình sản xuất và phân phối nhờ vào việc phân tích dữ liệu. Điều này giúp tiết kiệm chi phí và thời gian.

4.2. Tăng Cường Sự Hài Lòng Của Khách Hàng

Phân tích dữ liệu giúp doanh nghiệp hiểu rõ hơn về nhu cầu của khách hàng, từ đó cải thiện sản phẩm và dịch vụ, tăng cường sự hài lòng của khách hàng.

V. Kết Luận Tương Lai Của Phân Tích Dữ Liệu Kinh Doanh

Tương lai của phân tích dữ liệu kinh doanh rất hứa hẹn. Với sự phát triển của công nghệ và khoa học dữ liệu, các doanh nghiệp sẽ có nhiều cơ hội hơn để tối ưu hóa quy trình và nâng cao hiệu quả. Việc đầu tư vào phân tích dữ liệu sẽ mang lại lợi ích lâu dài cho doanh nghiệp.

5.1. Xu Hướng Mới Trong Khoa Học Dữ Liệu

Các xu hướng như trí tuệ nhân tạo và học máy đang ngày càng trở nên phổ biến trong phân tích dữ liệu. Những công nghệ này sẽ mở ra nhiều cơ hội mới cho doanh nghiệp.

5.2. Tầm Quan Trọng Của Đào Tạo Nhân Lực

Đào tạo nhân lực trong lĩnh vực khoa học dữ liệu là rất cần thiết. Doanh nghiệp cần chuẩn bị đội ngũ nhân viên có kỹ năng để tận dụng tối đa các công cụ phân tích.

10/07/2025
Business analytics data science for business problems

Trích đoạn nội dung tài liệu

Paczkowski Business Analytics Data Science for Business Problems Business Analytics Walter R. Paczkowski Business Analytics Data Science for Business Problems Walter R. Paczkowski Data Analytics Corp. Plainsboro, NJ, USA ISBN 978-3-030-87022-5 ISBN 978-3-030-87023-2 (eBook) https://doi.1007/978-3-030-87023-2 © Springer Nature Switzerland AG 2021 This work is subject to copyright.

All rights are reserved by the Publisher, whether the whole or part of the material is concerned, specifically the rights of translation, reprinting, reuse of illustrations, recitation, broadcasting, reproduction on microfilms or in any other physical way, and transmission or information storage and retrieval, electronic adaptation, computer software, or by similar or dissimilar methodology now known or hereafter developed. The use of general descriptive names, registered names, trademarks, service marks, etc. in this publication does not imply, even in the absence of a specific statement, that such names are exempt from the relevant protective laws and regulations and therefore free for general use. The publisher, the authors, and the editors are safe to assume that the advice and information in this book are believed to be true and accurate at the date of publication.

Neither the publisher nor the authors or the editors give a warranty, expressed or implied, with respect to the material contained herein or for any errors or omissions that may have been made. The publisher remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. This Springer imprint is published by the registered company Springer Nature Switzerland AG The registered company address is: Gewerbestrasse 11, 6330 Cham, Switzerland Preface I analyze business data—and I have been doing this for a long time. I was an analyst and department head, a consultant and trainer, worked on countless problems, written many books and reports, and delivered numerous presentations to all levels of management.

This book reflects insights I gained from this experience about Business Data Analytics that I want to share. There are three questions you should quickly ask about this sharing. The first is obvious: “Share what?” The second logically follows: “Share with whom?” The third is more subtle: “How does this book differ from other data analytic books?” The first is about focus, the second is about target, and the third is about competitive comparison. So, let me address each question.

The Book’s Focus My experience has been with practical business problems. When I finished my academic training with a Ph. in economics and a heavy statistics exposure, I immediately started my professional career with an AT&T internal consulting group, The Analytical Support Center (ASC). I quickly learned that I needed both a theoretical, technical understanding of quantitative work—how to estimate a regression model, for example—and an understanding of how to deal with messy data beyond the nice, clean data sets I used as a graduate student.

My time at the ASC was a great learning experience that I carried throughout my professional career at AT&T, including Bell Labs, and into my own consulting business. The lessons I learned were that good, solid data analysis for practical business problems requires: 1. A theoretical understanding of statistical, econometric, and (in the current era) machine learning methods 2. Data handling capabilities encompassing data organizing, preprocessing, and wrangling 3.

Programming knowledge in at least one software language v vi Preface These three components form a synergistic whole, a unifying approach if you wish, for doing business data analytics, and, in fact, any type of data analysis. This synergy implies that one part does not dominate any of the other two. They work together, feeding each other with the goal of solving only one overarching problem: how to provide decision makers with rich information extracted from data. Recognizing this problem was the most valuable lesson of all.

All the analytical tools and know how must have a purpose and solving this problem is that purpose—there is no other. I show this problem and the synergy of the three components for solving it as a triangle in Fig. This triangle represents the almost philosophical approach I take for any form of business data analysis and is the one I advocate for all data analyses. Theoretical Framework Problem: Provide Rich Information Data Raw Data Programming Handling Processed Data Literacy Empirical Stage of Analysis Preface vii CEO) to provide actionable, insightful, and useful rich information relevant for their problem.

If the limitations of a methodology prevent you from accomplishing your charge, then your life as an analyst will be short-lived, to say the least. This will hold if you either do not know these limitations or simply choose to ignore them. Another methodological approach might be better, one that has fewer problems, or is just more applicable. There is a dichotomy in methodology training.

Most graduate-level statistics and econometric programs, and the newer Data Science programs, do an excellent job instructing students in the theory behind the methodologies. The focus of these academic programs is largely to train the next generation of academic professionals, not the next generation of business analytical professionals. Data Science programs, of which there are now many available online and “in person,” often skim the surface of the theoretical underpinnings since their focus is to prepare the next generation of business analysts, those who will tackle the business decision makers’ tough problems, and not the academic researchers. Something in between the academic and data science training is needed for successful business data analysts.

Data handling is not as obvious since it is infrequently taught and talked about in academic programs. In those programs, beginner students work with clean data with few problems and that are in nice, neat, and tidy data sets. They are frequently just given the data. More advanced students may be required to collect data, most often at the last phase of training for their thesis or dissertation, but these are small efforts, especially when compared to what they will have to deal with post training.

The post-training work involves: • Identifying the required data from diverse, disparate, and frequently disconnected data sources with possibly multiple definitions of the same quantitative concept • Dealing with data dictionaries • Dealing with samples of a very large database—how to draw the sample and determine the sample size • Merging data from disparate sources • Organizing data into a coherent framework appropriate for the statisti- cal/econometric/machine learning methodology chosen • Visualizing complex multivariate data to understand relationships, trends, pat- terns, and anomalies inside the data sets This is all beyond what is provided by most training programs. Finally, there is the programming. First, let me say that there is programming and then there is programming. The difference is scale and focus.

Most people, when they hear about programming and programming languages, immediately think about large systems, especially ones needing a considerable amount of time (years?) to fully specify, develop, test, and deploy. They would be correct regarding large-scale, complex systems that handle a multitude of interconnected operations. Online ordering systems easily come to mind. Customer interfaces, inventory management, production coordination, supply chain management, price maintenance and dynamic pricing platforms, shipping and tracking, billing, and viii Preface collections are just a few components of these systems.

The programming for these is complex to say the least. As a business data analyst, you would not be involved in this type of program- ming although you might have to know about and access the subsystems of one or more of these larger systems. And major businesses are composed of many larger systems! You might have to write code to access the data, manipulate the retrieved data, and so forth, basically write programming code to do all the data handling I described above. And for this you need to know programming and languages.

There are many programming languages available. Only a few are needed for most business data analysis problems. In my experience, these are: • SQL • Python • R Julia should be included because it is growing in popularity due to its performance and ease of use. For this book, I will use Python because its ecosystem is strongly oriented toward machine learning with strong modeling, statistics, data visualization, and programming functionalities.

In fact, its programming paradigm is clear to use, which is a definite advantage over other languages. The Target Audience The target audience for this book consists of business data analysts, data scientists, and market research professionals, or those aspiring to be any of these, in the private sector. You would be involved in or responsible for a myriad of quantitative analyses for business problems such as, but not limited to: • Demand measurement and forecasting • Predictive modeling • Pricing analytics including elasticity estimation • Customer satisfaction assessment • Market and advertisement research • New product development and research To meet these tasks, you will have a need to know basic data analytical methods and some advanced methods, including data handling and management. This book will provide you with this needed background by: • Explaining the intuition underlying analytic concepts • Developing the mathematical and statistical analytic concepts • Demonstrating analytical concepts using Python • Illustrating analytical concepts with case studies Preface ix This book is also suitable for use in colleges and universities offering courses and certifications in business data analytics, data sciences, and market research.

It could be used as a major or supplemental textbook. Since the target audience consists of either current or aspiring business data analysts, it is assumed that you have or are developing a basic understanding of fun- damental statistics at the “Stat 101” level: descriptive statistics, hypothesis testing, and regression analysis. Knowledge of econometric and market research principles, while not required, would be beneficial. In addition, a level of comfort with calculus and some matrix algebra is recommended, but not required.

Appendices will provide you with some background as needed. The Book’s Competitive Comparison There are many books on the market that discuss the three themes of this book: analytic methods, data handling, and programming languages. But they do them separately as opposed to a synergistic, analytic whole. They are given separate treatment so that you must cover a wide literature just to find what is needed for a specific business problem.

Also, once found, you must translate the material into business terms. This book will present the three themes so you can more easily master what is needed for your work. The Book’s Structure I divided this book into three parts. In Part I, I cover the basics of business data analytics including data handling, preprocessing, and visualization.

In some instances, the basic analytic toolset is all you need to address problems raised by business executives. Part II is devoted to a richer set of analytic tools you should know at a minimum. These include regression modeling, time series analysis, and statistical table analysis. Part III extends the tools from Part II with more advanced methods: advanced regression modeling, classification methods, and grouping methods (a.

The three parts lead naturally from basic principles and methods to complex methods. I illustrate this logical order in Fig. Embedded in the three parts are case study examples of business problems using (albeit, fictitious, fake, or simulated) business transactions data designed to be indicative of what business data analysts use every day. Using simulated data x Preface Part III Analytics Progression Advanced Analytics: Going Further Business Data Part II Intermediate Analytics: Gaining Insight Part I Beginning Analytics: Getting Started Fig.

2 This is a flow chart of the three parts of this book. The parts move progressively from basics to advanced. At the end of Part I, you should be able to do basic analyses of business data. At the end of Part II, you should be able to do regression and times series analysis.

At the end of Part III, you should be able to do advanced machine learning work for instructional purposes is certainly not without precedence. See, for example, Gelman et al. Data handling, visualization, and modeling are all illustrated using Python. All examples are in Jupyter notebooks available on Github.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ