Data Locality in Distributed Computing Applications - Shubhada Nayak

Khám phá luận án Nayak Shubhada Nithyananda 2021 về nghiên cứu khoa học tiên tiến. Phân tích chi tiết các đóng góp và ứng dụng trong lĩnh vực chuyên môn.

Trường đại học

California State University, Northridge

Chuyên ngành

Khoa học máy tính

Người đăng

Ẩn danh

Thể loại

Luận văn thạc sĩ

2021

58
0
0

Phí lưu trữ

30 Point

Tóm tắt

I. Tổng quan về luận án Shubhada Nayak về Data Locality 2021

Luận án của Shubhada Nayak năm 2021 tại Đại học California State Northridge tập trung vào vấn đề Data Locality trong các ứng dụng điện toán phân tán trên nền tảng đám mây. Nghiên cứu phân tích mối quan hệ giữa vị trí dữ liệu và hiệu suất xử lý trong môi trường phân tán, đặc biệt nhấn mạnh vào các thách thức khi triển khai ứng dụng MapReduce trên hệ thống Hadoop. Luận án đề xuất các phương pháp tối ưu hóa vị trí dữ liệu nhằm giảm thiểu chi phí truyền dữ liệu qua mạng, vốn là nguồn tài nguyên đắt đỏ so với khả năng tính toán của các nút. Kết quả nghiên cứu cung cấp cơ sở lý thuyết và thực nghiệm cho việc cải thiện hiệu suất hệ thống phân tán thông qua chiến lược phân bổ dữ liệu thông minh.

1.1. Mục tiêu nghiên cứu của luận án

Mục tiêu chính của luận án nhằm giải quyết vấn đề di chuyển dữ liệu dư thừa trong các hệ thống phân tán bằng cách tối ưu vị trí dữ liệu. Shubhada Nayak đề xuất hai phương pháp chính: DCR (Data Computation Ratio) và NDCR (Normalized Data Computation Ratio) để đánh giá khả năng tính toán của từng nút trong cụm. Các phương pháp này giúp phân bổ dữ liệu gần hơn với các nút có khả năng xử lý mạnh hơn, từ đó giảm thiểu thời gian truyền dữ liệu và cải thiện hiệu suất tổng thể. Nghiên cứu cũng xem xét tác động của yếu tố đồng nhất trong cụm tới hiệu quả phân bổ dữ liệu.

1.2. Đối tượng và phạm vi nghiên cứu

Nghiên cứu tập trung vào môi trường điện toán đám mây sử dụng framework Hadoop và MapReduce. Phạm vi khảo sát bao gồm kiến trúc HDFS, YARN và các thành phần liên quan như rack awareness. Luận án tiến hành thực nghiệm trên các cụm máy ảo AWS với các kích thước dữ liệu khác nhau (từ 1MB đến 455MB) để đánh giá hiệu quả của các phương pháp đề xuất. Kết quả thực nghiệm được so sánh với phương pháp phân bổ dữ liệu mặc định của Hadoop nhằm minh chứng tính ưu việt của các giải pháp mới.

II. Phân tích vấn đề Data Locality trong hệ thống phân tán

Data Locality trở thành thách thức lớn trong các hệ thống phân tán khi khối lượng dữ liệu tăng trưởng nhanh chóng. Vấn đề cốt lõi nằm ở sự không đồng nhất giữa khả năng tính toán của các nút và vị trí lưu trữ dữ liệu. Hệ thống Hadoop mặc định phân bổ dữ liệu ngẫu nhiên mà không xem xét đặc điểm phần cứng của từng nút, dẫn đến tình trạng nút yếu xử lý dữ liệu xa nguồn, gây lãng phí băng thông mạng. Ngoài ra, cơ chế nhân bản dữ liệu mặc định (replication factor = 3) tiêu tốn đáng kể tài nguyên lưu trữ và băng thông. Luận án chỉ ra rằng hiệu suất hệ thống giảm sút đáng kể khi dữ liệu được phân bổ không phù hợp với khả năng xử lý của các nút.

2.1. Nguyên nhân gây lãng phí tài nguyên mạng

Sự không đồng nhất trong cụm máy tính là nguyên nhân chính gây lãn phí tài nguyên mạng. Các nút có cấu hình phần cứng khác nhau (CPU, RAM, tốc độ đĩa) dẫn đến khả năng xử lý khác biệt. Khi dữ liệu được phân bổ ngẫu nhiên, các tác vụ thường được giao cho những nút yếu hơn, buộc phải kéo dữ liệu từ xa. Kết quả là thời gian xử lý tăng lên đáng kể do chi phí truyền dữ liệu qua mạng. Nghiên cứu cũng chỉ ra rằng cơ chế nhân bản dữ liệu mặc định không xem xét đặc điểm phần cứng, dẫn đến tình trạng dư thừa không cần thiết.

2.2. Hạn chế của cơ chế phân bổ dữ liệu mặc định

Cơ chế phân bổ dữ liệu mặc định của Hadoop dựa trên thuật toán round-robin, không xem xét khả năng tính toán của từng nút. Điều này dẫn đến tình trạng mất cân bằng tải nghiêm trọng trong cụm. Các nút mạnh thường không được tận dụng tối đa trong khi các nút yếu phải xử lý khối lượng công việc lớn. Ngoài ra, cơ chế này không thích ứng với sự thay đổi động của cụm (thêm/bớt nút, thay đổi cấu hình). Luận án nhấn mạnh sự cần thiết của các thuật toán phân bổ dữ liệu thích ứng, có khả năng tự điều chỉnh dựa trên tình trạng hoạt động thực tế của cụm.

III. Giải pháp tối ưu Data Locality bằng phương pháp DCR và NDCR

Shubhada Nayak đề xuất hai phương pháp phân bổ dữ liệu mới: DCR (Data Computation Ratio) và NDCR (Normalized Data Computation Ratio) nhằm tối ưu vị trí dữ liệu trong hệ thống phân tán. Phương pháp DCR tính toán tỷ lệ giữa khối lượng công việc xử lý và dung lượng lưu trữ của từng nút, từ đó phân bổ dữ liệu ưu tiên cho các nút có khả năng xử lý mạnh hơn. NDCR cải tiến DCR bằng cách chuẩn hóa tỷ lệ này theo dung lượng lưu trữ, giúp so sánh công bằng giữa các nút có dung lượng khác nhau. Các phương pháp này được triển khai thông qua API REST, cho phép tích hợp linh hoạt với hệ thống Hadoop hiện có.

3.1. Cách thức hoạt động của phương pháp DCR

Phương pháp DCR hoạt động bằng cách thu thập dữ liệu hiệu suất từ các nút thông qua các tác vụ thử nghiệm nhỏ. Hệ thống tính toán tỷ lệ giữa khối lượng công việc xử lý (tính bằng số lượng tác vụ hoàn thành) và dung lượng lưu trữ của từng nút. Dữ liệu sau đó được phân bổ ưu tiên cho các nút có tỷ lệ DCR cao nhất. Quá trình này được thực hiện định kỳ để thích ứng với sự thay đổi của cụm. Kết quả thực nghiệm cho thấy phương pháp DCR giảm tới 35% thời gian xử lý so với cơ chế mặc định khi áp dụng trên cụm AWS.

3.2. Ưu điểm của phương pháp NDCR so với DCR

NDCR cải tiến DCR bằng cách chuẩn hóa tỷ lệ xử lý/dung lượng lưu trữ theo dung lượng lưu trữ tối đa của cụm. Điều này giúp so sánh công bằng giữa các nút có dung lượng lưu trữ khác nhau, đặc biệt hữu ích trong môi trường đám mây nơi dung lượng lưu trữ có thể thay đổi động. NDCR cũng tích hợp cơ chế phát hiện các nút chậm (slow nodes) và điều chỉnh phân bổ dữ liệu tương ứng. Kết quả thực nghiệm cho thấy NDCR cải thiện hiệu suất tới 42% so với DCR trong các tình huống có sự chênh lệch lớn về dung lượng lưu trữ giữa các nút.

IV. Kết luận và ứng dụng thực tiễn của luận án 2021

Nghiên cứu của Shubhada Nayak năm 2021 đã chứng minh hiệu quả vượt trội của các phương pháp DCR và NDCR trong việc tối ưu Data Locality cho các ứng dụng điện toán phân tán. Kết quả thực nghiệm cho thấy cả hai phương pháp đều cải thiện đáng kể hiệu suất so với cơ chế mặc định của Hadoop, đặc biệt trong môi trường có sự không đồng nhất về phần cứng. Luận án cũng đề xuất hướng nghiên cứu tiếp theo về tích hợp cơ chế nhân bản dữ liệu thông minh và xử lý các nút chậm. Các kết quả này có giá trị tham khảo cho các nhà phát triển hệ thống phân tán, đặc biệt trong lĩnh vực xử lý dữ liệu lớn (Big Data).

4.1. Đánh giá hiệu quả thực nghiệm

Kết quả thực nghiệm trên các kích thước dữ liệu khác nhau (từ 1MB đến 455MB) cho thấy phương pháp NDCR đạt hiệu suất cao nhất, giảm tới 42% thời gian xử lý so với cơ chế mặc định. Phương pháp DCR cũng đạt được cải thiện 35% trong cùng điều kiện. Cả hai phương pháp đều thể hiện ưu việt trong việc xử lý các file dung lượng lớn. Nghiên cứu cũng chỉ ra rằng hiệu quả cải thiện tỷ lệ thuận với mức độ không đồng nhất của cụm.

4.2. Hướng phát triển trong tương lai

Luận án đề xuất một số hướng nghiên cứu tiếp theo bao gồm tích hợp cơ chế nhân bản dữ liệu thông minh (smart replication) dựa trên đặc điểm phần cứng, phát triển thuật toán phát hiện và cô lập các nút chậm tự động, và mở rộng nghiên cứu sang các nền tảng điện toán phân tán khác ngoài Hadoop. Ngoài ra, việc áp dụng các kỹ thuật học máy để dự đoán hiệu suất nút cũng là một hướng nghiên cứu hứa hẹn. Các giải pháp này hứa hẹn sẽ mang lại những cải tiến vượt trội cho hiệu suất hệ thống phân tán trong tương lai.

Tóm tắt và mô tả trên trang này được tạo với sự hỗ trợ của AI. Nếu bạn thấy nội dung không chính xác hoặc có vấn đề, vui lòng Báo lỗi nội dung.

30/05/2026
Nayak shubhada nithyananda thesis 2021

Trích đoạn nội dung tài liệu

CALIFORNIA STATE UNIVERSITY, NORTHRIDGE Data Locality in Distributed Computing Applications in the Cloud A thesis submitted in partial fulfillment of the requirements For the degree of Master of Science in Computer Science By Shubhada Nithyananda Nayak May 2021 Signature The thesis of Shubhada Nithyananda Nayak is approved: Dr. Robert D McIlhenny Date Dr. John Noga Date Dr. Mahdi Ebrahimi, Chair Date California State University, Northridge ii Acknowledgment I would like to express my profound gratitude to my advisor and chair, Dr.

Mahdi Ebrahimi, for his continued motivation and guidance. Thank you for your support, feedback, and valuable input throughout this thesis project. I would also like to extend warm thanks to my thesis committee members Dr. Robert Mcllhenny and Dr.

John Noga, for their support during this thesis project. I am indebted to my family for having given me this opportunity to pursue higher education in the United States. I credit this thesis to my mother, Mrs. Nayana Nayak, for her unconditional love and encouragement.

I am forever grateful to my sister and her family for placing their faith in my abilities and for their constant support in my difficult times. This would not have been possible without the support of my fiancé; thank you for empowering me to achieve the best. iii Table of Contents Signature. Hadoop and Hadoop Ecosystem.

Hadoop YARN Architecture. Application Workflow in Hadoop YARN. Word Count Example in MapReduce. Rack Awareness in Hadoop.

Replica Placement using Rack Awareness. Categories of Data Locality. Data Locality in Hadoop. Data Distribution Based on Node Capabilities.

Adaptive Data Placement. Building the Distributed Computing Framework on Cloud. Laying the Bricks. Building the REST APIs.

Conclusions and Future Work. 49 v List of Tables Table 1:Key-Value Pairs from Map Phase. 12 Table 2:Sort & Shuffle Phase. 13 Table 3: Reduce Phase.

13 Table 4:Text file size and number of words. 23 Table 5:Compute DCR. 35 Table 6: File Splits using DCR. 36 Table 7:NDCR Compute Rules.

37 Table 8:AWS cluster. 39 vi List of Figures Figure 1: HDFS Architecture .7 Figure 3: Hadoop YARN Architecture .8 Figure 4: Application Workflow in YARN .9 Figure 5: Map Reduce Architecture. 11 Figure 6: Rack Awareness in Hadoop. 16 Figure 7: Default Word Count 1.

42 Figure 8: DCR Word Count 1. 43 Figure 9: NDCR Word Count 1. 43 Figure 10: Default Word Count 32KB file. 44 Figure 11: DCR Word Count 32KB file.

44 Figure 12: NDCR Word Count 32KB file. 45 Figure 13: Default Word Count 455MB text file. 45 Figure 14: DCR Word Count 455MB file. 46 Figure 15: NDCR Word Count 455MB file.

46 Figure 16: Performance comparisons of Data Distribution Schemes. 47 vii Abstract Data Locality in Distributed Computing Applications in the Cloud By Shubhada Nithyananda Nayak Master of Science in Computer Science The data processing in data-intensive applications has become increasingly complex. Data placement in a distributed computing environment is critical as it is the primary factor for determining tasks' performance and scheduling. A job is divided into multiple smaller tasks in a distributed computing system and run in various nodes in a large-scale cluster.

The data placement strategies from a homogeneous cluster can hinder the performance in a heterogenous cluster by increasing the overhead of data transfer of unprocessed data from slow nodes to fast nodes. In a distributed computing environment, the emphasis is on the programming model and the distributed file system. Data movement is expensive than compute movement; the goal is to move the task closer to the data. This thesis explores data locality based on the capabilities of computation nodes and how this can be dynamically improved based on the load status and studies behavior on pushing data locality to stages as far as reduce.

The distributed computing framework is simulated on AWS cloud with Leader and worker nodes performing word count operations on files of varying sizes based on their data computation ratios computed off their computing capabilities calibrated using matrix multiplication and dynamically improving them ratios based on CPU and memory usage statistics. The results show enhanced computation times and thus a novel way to achieve data locality in heterogeneous computing systems. Introduction With the evolution of technology, distributed systems are becoming very ubiquitous. In its simplest definition, a distributed computer system is a group of computing machines working together to achieve a single task and appear as a single machine to the end-user.

It is becoming very complicated in data to measure the total volume of data stored electronically. There are various streaming data sources such as social networking websites, e-commerce user traffic, data archive stores, etc. Most importantly, these digital streams are growing rapidly as data becomes a crucial tool to add value to the business. The trend is for this data footprint to grow as more machines in the computing space will generate even more essential data.

System logs, sensor data, retail transactions, user behavior, contribute to the growing mountain of data. As the volume of data available to access for the public is increasing, the goal is not merely to manage private data but also to develop the tools and technology to extract and analyze value from publicly available data resources. No matter how great the computing algorithms are, they can be easily defeated by the volume of data it has to process. The elephant in the room is although the storage capacities have spiraled, the disk access speed, which is the speed at which data can be read from the disk, has not increased.

This means long wait times to read all the data from the drive and even slower writes to the disk; all this is when we consider a single disk for the operation. What if there were multiple disks? This would reduce the read-write times. If numerous disks were storing a portion of the data, working in parallel, we could read and write in under a few minutes. When we think of parallel operations, of the many problems, the first to arise is fault tolerance.

As many hardware pieces begin to operate together, the chances of one or a few failing steadily increase. A solution to avoid this is replication; redundant copies of data are maintained by the 1 system. In the event of a hardware failure, there is a copy readily available. The second problem demanding attention is how to combine the data; a data read from one disk may need to be combined with data from several hundred disks.

This involves data transfer between the computing nodes, which may prove to be a bottleneck, thereby bringing down the performance and increasing the computing times. MapReduce provides a programming model that simplifies and abstracts the problem of read-write and reforms it into computation over a set of keys and values. There are two parts to the MapReduce paradigm: the Map phase and the Reduce phase, and the sort, shuffle, and spill occur here in between. Hadoop is a distributed computing framework which, unlike the traditional systems, enables multiple workloads to run on the same data parallelly at a massive scale on industry commodity hardware.

Hadoop comes with reliable shared storage: HDFS (Hadoop Distributed File System) and data analysis with MapReduce programming model. Brief History The creator of Hadoop is Doug Cutting. Hadoop has its origins in Apache Nutch, an open-source web search engine and the Apache Lucene Project. Apache Nutch was started in 2002 as a crawler and a search system engine, but soon it was realized the architecture would not scale for billions of pages on the web.

Based on the Google File System used in production at Google, a way was sought to solve storage needs for the huge files generated as a part of the web crawl and indexing process. Google introduced MapReduce to the world, Nutch had a working implementation of MapReduce and later replaced all major algorithms with running using MapReduce. Yahoo further developed the Nutch project into Hadoop, a system that ran a web scale. Since then, Hadoop has seen a rapid rate of adoption at an enterprise scale.

The industry has recognized Hadoop’s pivotal 2 role as a distributed storage system and analysis platform for big data. Hadoop is the most widely used framework for batch processing jobs and smaller tasks where the time isn’t the primary consideration. Hadoop and Hadoop Ecosystem Hadoop Ecosystem is a suite that provides various services for big data problems. It encompasses the Apache project and various commercial tools and solutions.

The main elements of Hadoop are Hadoop Distributed File System, MapReduce, Yarn, and Hadoop Common. These tools work collectively to provide services such as analysis, storage, data maintenance, etc. Following are the components that collectively form a Hadoop ecosystem: HDFS: Hadoop Distributed File System YARN: Yet Another Resource Negotiator MapReduce: Programming paradigm for data processing Spark: In-Memory data processing PIG, HIVE: Query-based processing of data services HBase: NoSQL database Spark MLlib: Machine learning algorithm libraries Zookeeper: Cluster manager Oozie: job scheduler All the components revolve around data; that is the essence of Hadoop, which makes it easier for data handling and analysis. 3 • HDFS: Hadoop Distributed File System is the major component of the Hadoop Ecosystem; it handles storage of large amounts of data that is structured and unstructured across a cluster of nodes and thereby maintains metadata in the form of log files.

HDFS consists of two core components which are the Name Node and Data Node. Name Node is the primary node that contains metadata, file system logs which is data about data. It requires fewer resources than the data nodes, plays the role of a leader node. The Data Nodes are the industry commodity hardware.

The HDFS is the heart of the system, plays the role of maintaining coordination between cluster and the hardware. Figure 1: HDFS Architecture • YARN: Yet Another Resource Negotiator manages the resources in the cluster. It handles scheduling and resource allocation for the Hadoop ecosystem. It mainly consists of Resource Manager, Nodes Manager, Application Manager.

Resource allocation for the system's applications is taken care of by the resource manager; the allocation of node resources like CPU, memory is taken care of by the node manager. The Application Manager is the interface between the node manager and resource manager and handles the negotiations between them. 4 • MapReduce: By leveraging parallel and distributed algorithms, MapReduce carries the processing logic making it possible to write applications that can handle big data sets. The two most essential functions in this programming paradigm are Map() and Reduce().

Sorting, filtering, and organizing of the data are taken care of in the map phase. Map phase generates key-value pairs, which are processed in the reduce phase. Reduce as the name goes, aggregates the map data. It feeds on the output generated by the map phase and combines them to give the result.

• Pig: It is a query-based language that works on Pig Latin language. It provides a platform for structuring the data flow, processing, and analyzing big data sets. Pig abstracts the activities of MapReduce via commands and stores the results after processing in HDFS. It is a significant part of the Hadoop ecosystem and plays a pivotal role in optimizing and easing programming for Hadoop users.

• HIVE: Hive leverages SQL interface for reading and writing of big data sets. It is a query language called Hive Query Language. It enables real-time, and batches are processing while also supporting all the SQL datatypes making the processing quick. In short, it is a data warehouse software that facilitates the reading, writing, and management of large datasets residing in distributed storage systems.

• Apache Spark: It is a platform that handles all processing in-memory for an operation like real- time streaming and batch processing of enormous data sets.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ