Luận án tiến sĩ: Thiết kế kết nối hybrid cho các bộ gia tốc phần cứng đa dạng

Luận án tiến sĩ phân tích hybrid interconnect design for heterogeneous hardware accelerators, xây dựng cơ sở lý luận, kiểm chứng thực nghiệm, đóng góp tri thức mới cho ngành.

Trường đại học

Technische Universiteit Delft

Chuyên ngành

Computer Science

Người đăng

Ẩn danh

Thể loại

thesis

2015

165
1
0

Phí lưu trữ

45 Point

Tóm tắt

I. Thiết kế kết nối hybrid

Thiết kế kết nối hybrid là trọng tâm của luận án, tập trung vào việc tối ưu hóa giao tiếp dữ liệu giữa các bộ gia tốc phần cứng đa dạng. Luận án đề xuất một phương pháp heuristic dựa trên phân tích giao tiếp dữ liệu để thiết kế hệ thống với kết nối tùy chỉnh. Các giải pháp bao gồm bộ nhớ cục bộ chia sẻ dựa trên crossbar, DMA hỗ trợ xử lý song song, bộ đệm cục bộ và nhân bản phần cứng. Phương pháp này đặc biệt hữu ích cho các hệ thống nhúng nơi tài nguyên phần cứng bị hạn chế.

1.1. Phương pháp heuristic

Phương pháp heuristic được sử dụng để thiết kế kết nối tùy chỉnh, dựa trên phân tích giao tiếp dữ liệu của ứng dụng. Các giải pháp như crossbar, DMA và bộ đệm cục bộ được xem xét để tối ưu hóa hiệu suất hệ thống.

1.2. Tối ưu hóa tài nguyên

Luận án nhấn mạnh việc tối ưu hóa tài nguyên phần cứng bằng cách sử dụng các giải pháp kết nối hybrid, giúp giảm thiểu việc sử dụng tài nguyên trong khi vẫn đảm bảo hiệu suất cao.

II. Bộ gia tốc phần cứng đa dạng

Luận án tập trung vào việc thiết kế kết nối cho các bộ gia tốc phần cứng đa dạng, bao gồm các lõi xử lý chuyên dụng và các IP core. Các bộ gia tốc này được thiết kế để xử lý các tác vụ cụ thể, mang lại hiệu suất cao và tiết kiệm năng lượng so với các bộ xử lý đa năng.

2.1. Kiến trúc đa dạng

Các bộ gia tốc phần cứng được thiết kế với kiến trúc đa dạng, bao gồm các lõi xử lý chuyên dụng và các IP core, giúp tối ưu hóa hiệu suất cho các tác vụ cụ thể.

2.2. Hiệu suất và năng lượng

Các bộ gia tốc phần cứng đa dạng mang lại hiệu suất cao và tiết kiệm năng lượng, đặc biệt trong các ứng dụng yêu cầu xử lý dữ liệu lớn.

III. Kết nối phần cứng đa dạng

Luận án đề xuất một kết nối phần cứng đa dạng dựa trên phân tích giao tiếp dữ liệu, bao gồm NoC và bộ nhớ cục bộ chia sẻ. Phương pháp này giúp tối ưu hóa kết nối giữa các lõi tính toán và bộ nhớ cục bộ, giảm thiểu việc sử dụng tài nguyên phần cứng.

3.1. NoC và bộ nhớ chia sẻ

Kết nối hybrid bao gồm NoC và bộ nhớ cục bộ chia sẻ, giúp tối ưu hóa giao tiếp dữ liệu giữa các lõi tính toán và bộ nhớ.

3.2. Thuật toán ánh xạ thích ứng

Thuật toán ánh xạ thích ứng được đề xuất để kết nối các lõi tính toán và bộ nhớ cục bộ với kết nối hybrid, giúp tối ưu hóa hiệu suất hệ thống.

IV. Ứng dụng thực tế

Luận án đã triển khai các phương pháp trên các nền tảng phần cứng thực tế, chứng minh hiệu quả trong việc cải thiện hiệu suất và giảm tiêu thụ năng lượng. Các kết quả thực nghiệm cho thấy các phương pháp đề xuất không chỉ cải thiện hiệu suất hệ thống mà còn giảm tiêu thụ năng lượng so với các hệ thống cơ bản.

4.1. Kết quả thực nghiệm

Các kết quả thực nghiệm trên các nền tảng phần cứng thực tế cho thấy sự cải thiện đáng kể về hiệu suất và tiêu thụ năng lượng.

4.2. Ứng dụng trong xử lý ảnh

Luận án cũng đề xuất một kiến trúc phần cứng hỗ trợ xử lý ảnh theo luồng, chứng minh hiệu quả trong các ứng dụng thực tế.

21/02/2025

Trích đoạn nội dung tài liệu

Phạm Quốc Cường H YBRID I NTERCONNECT D ESIGN FOR H ETEROGENEOUS H ARDWARE A CCELERATORS H YBRID I NTERCONNECT D ESIGN FOR H ETEROGENEOUS H ARDWARE A CCELERATORS Proefschrift ter verkrijging van de graad van doctor aan de Technische Universiteit Delft, op gezag van de Rector Magnificus prof. Luyben, voorzitter van het College voor Promoties, in het openbaar te verdedigen op dinsdag 14 April 2015 om 12:30 uur door Cuong PHAM-QUOC Master of Engineering in Computer Science Ho Chi Minh City University of Technology - HCMUT, Vietnam geboren te Tien Giang, Vietnam. This dissertation has been approved by the Promotor: Prof.M Bertels Copromotor: Dr. Al-Ars Composition of the doctoral committee: Rector Magnificus voorzitter Prof.M Bertels Technische Universiteit Delft, promotor Dr.

Al-Ars Technische Universiteit Delft, copromotor Independent members: Prof. Charbon Technische Universiteit Delft Prof. Becker Karlsruhe Institute of Technology Prof. Dinh-Duc Vietnam National University - Ho Chi Minh City Prof.

Luigi Carro Universidade Federal do Rio Grande do Sul Dr. Silla Universitat Politècnica de València Prof.-J van der Veen Technische Universiteit Delft, reservelid Keywords: Hybrid interconnect, hardware accelerators, data communication, quan- titative data usage, automated design. Copyright © 2015 by Cuong Pham-Quoc All rights reserved. No part of this publication may be reproduced, stored in a re- trieval system, or transmitted, in any form or by any means, electronic, mechan- ical, photocopying, recording, or otherwise, without permission of the author.

ISBN 978-94-6186-448-2 Cover design: Cuong Pham-Quoc Printed in The Netherlands To my wife and my son A BSTRACT Heterogeneous multicore systems are becoming increasingly important as the need for computation power grows, especially when we are entering into the big data era. As one of the main trends in heterogeneous multicore, hardware accelerator systems provide application specific hardware circuits and are thus more energy efficient and have higher performance than general purpose pro- cessors, while still providing a large degree of flexibility. However, system perfor- mance dose not scale when increasing the number of processing cores due to the communication overhead which increases greatly with the increasing number of cores. Although data communication is a primary anticipated bottleneck for sys- tem performance, the interconnect design for data communication among the accelerator kernels has not been well addressed in hardware accelerator systems.

A simple bus or shared memory is usually used for data communication between the accelerator kernels. In this dissertation, we address the issue of interconnect design for heterogeneous hardware accelerator systems. Evidently, there are dependencies among computations, since data produced by one kernel may be needed by another kernel. Data communication patterns can be specific for each application and could lead to different types of intercon- nect.

In this dissertation, we use detailed data communication profiling to de- sign an optimized hybrid interconnect that provides the most appropriate sup- port for the communication pattern inside an application while keeping the hard- ware resource usage for the interconnect minimal. Firstly, we propose a heuristic- based approach that takes application data communication profiling into ac- count to design a hardware accelerator system with a custom interconnect. A number of solutions are considered including crossbar-based shared local mem- ory, direct memory access (DMA) supporting parallel processing, local buffers, and hardware duplication. This approach is mainly useful for embedded sys- tem where the hardware resources are limited.

Secondly, we propose an auto- mated hybrid interconnect design using data communication profiling to define an optimized interconnect for accelerator kernels of a generic hardware accel- erator system. The hybrid interconnect consists of a network-on-chip (NoC), vii viii A BSTRACT shared local memory, or both. To minimize hardware resource usage for the hybrid interconnect, we also propose an adaptive mapping algorithm to con- nect the computing kernels and their local memories to the proposed hybrid in- terconnect. Thirdly, we propose a hardware accelerator architecture to support streaming image processing.

In all presented approaches, we implement the ap- proach using a number of benchmarks on relevant reconfigurable platforms to show their effectiveness. The experimental results show that our approaches not only improve system performance but also reduce overall energy consumption compared to the baseline systems. A CKNOWLEDGMENTS It is not easy to write this last part of the dissertation, but this is an exciting period because it lets me take a careful look at the whole last four years, starting from 2011. First, I would like to thank the Vietnam International Education Develop- ment (VIED) for their funding.

Without this funding, I would not have been in the Netherlands. I would like to express special appreciation and thanks to my promoter, Prof. Koen Bertels, who had a difficult decision, but a successful one, when ac- cepting me as his Ph. At that time, my spoken English was not very good but he tried very hard to understand our Skype-based discussion.

During my time at the Computer Engineering Lab, he has introduced me to so many great ideas and has given me freedom to do my research. Koen, without you, I would have had no chance to write this dissertation. Another significant appreciation and thanks are given to my daily supervisor, but he always says that I am his friend, Dr. Zaid Al-Ars, who has guided me a lot not only in doing re- search but also in writing a paper.

Zaid, I can never forget the many hours you have spent correcting my papers. Without you, I would have no publication and, of course, no dissertation. Besides these two great persons, I would like to say thank you to Veronique from Valorisation Center - TUDelft, Lidwina - CE sec- retary, and Eef and Erik - CE system administrators, for their support. I would like to thank my colleagues, Razvan, for your DWARV compiler and, Vlad, for the Molen platform upon which I have conducted the experiments.

Thank you, Ernst, for your time translating my abstract and my proposition into Dutch. I need to say thank you to Prof. Anh-Vu Dinh-Duc. This is the third time I have written his name in my thesis.

The first and the second times were as my supervisor while this time is as a committee member. He has been there at many steps of my learning journey. I also appreciate all the committee members’ time and the remarks they gave me. Life is not only doing research.

Without relaxing time and parties, we have no energy and no ideas. So, thank you to the ANCB group, a group of Vietnamese students, for the very enjoyable parties. Those parties and relaxing time helped ix x A CKNOWLEDGMENTS me refresh my mind after the tiring working days. I am sure that I cannot say thank you to everybody who has supported me during the last four years because it would take a hundred pages, but I am also sure that I will never forget.

Let me keep your kindness in my mind. I am extremely grateful for my family and my wife’s family, especially my fa- ther in law and my mother in law who have helped me to take care of my son when I could not be at home. Without you, I would not have had the peace of mind to do my work. Last but most importantly, I would like to say thank you so much my wife and my son.

You raise me up, and you make me stronger. Without your love and your support, I cannot do anything. Our family is going to reunite in the next couple of months after a long period of connecting together through a “hybrid inter- connect” - a combination of video-calls, telephone calls, emails, social networks, and traveling. Phạm Quốc Cường Delft, April 2015 C ONTENTS Abstract vii Acknowledgments ix List of Figures xv List of Tables xix 1 Introduction 1 1.

8 2 Background and Related Work 11 2.1 On-chip Interconnect .2 System-level Hybrid Interconnect .1 Mixed topologies hybrid interconnect .2 Mixed architectures hybrid interconnect .3 Interconnect in Hardware Accelerator Systems .4 Data Communication Optimization Technique .1 Software level optimization .2 Hardware level optimization. 30 3 Communication Driven Hybrid Interconnect Design 33 3.1 Overview Hybrid Interconnect Design .2 Data Communication Driven Quantitative Execution Model .1 Baseline execution model .2 Ideal execution model .3 Parallelizing kernel processing. 41 xi xii C ONTENTS 3. 42 4 Bus-based Interconnect with Extensions 45 4.2 Bus-based hardware accelerator systems .3 Different Interconnect Solutions .1 Assumptions and definitions .2 Bus-based interconnect .3 Bus-based with a consolidation of a DMA .4 Bus-based with a consolidation of a crossbar .5 Bus-based with both a DMA and a crossbar .6 NoC-based interconnect.

62 5 Heuristic Communication-aware Hardware Optimization 63 5.2 Custom Interconnect and System Design .3 Heuristic-based algorithm. 78 6 Automated Hybrid Interconnect Design 81 6.2 Automated Hybrid Interconnect Design .1 Modeling system components.2 Custom interconnect design .3 Adaptive mapping function .1 Embedded system results .2 High performance computing results. 103 7 Accelerator Architecture for Stream Processing 105 7.2 Background and Related Work .1 Streaming image processing with hardware acceleration .2 Canny edge detection algorithm .1 Hardware-software streaming model .3 Multiple clock domains .4 Case Study: Canny Edge Detection. 117 8 Conclusions and Future Work 119 8.

122 Bibliography 125 List of Publications 143 Curriculum Vitæ 145 L IST OF F IGURES 1.1 The evolution of the on-chip interconnects .2 (a) Directly shared local memory; (b) Bus; (c) Crossbar; (d) Network- on-Chip .4 Examples of NoC topologies: (a) 2D-mesh; (b) ring; (c) hypercube; (d) tree; and (e) star.5 A generic hardware accelerator architecture .1 (a) The generic FPGA-based accelerator architecture; (b) The generic FPGA-based accelerator system with our hybrid interconnect.2 Hybrid interconnect design steps .3 Example of a QDU graph .4 The sequential diagrams for the baseline (left) and ideal execution model (right) .5 An example of data parallelism processing compared to serial pro- cessing .6 An example of instruction parallelism processing compared to se- rial processing .1 The bus is used as interconnect .2 The DMA is used as a consolidation to the bus .3 The crossbar is used as a consolidation to the bus .4 The DMA and the crossbar are used as consolidations to the bus .5 The NoC is used as interconnect of the hardware accelerators .6 The communication profiling graph generated by QUAD tool for the jpeg application. 57 xv xvi L IST OF F IGURES 4.7 Comparison between computation (Comp.), hardware accelerator execution (HW Acc.), and theoretical com- munication (Theoretical Comm.) times normalized to software time 59 4.8 Speed-up of hardware accelerators with respect to software and bus-based model .9 Comparison of resource utilization and energy consumption nor- malized to bus-based model .1 (a) HW1 and HW2 share their memories using a crossbar; (b) Struc- ture of the crossbar for the Molen architecture .2 Local buffer at HW2 .3 QUAD graph for the Canny edge detection application .4 Final system for Canny based on the Molen architecture and pro- posed solutions .t software) of hardware accelerators using Molen plat- form with and without using custom interconnect .6 The contribution of each solution to the speed-up .1 Shared local memory with and without crossbar in a hardware ac- celerator system.2 The NoC is used as interconnect of the kernels in a hardware accel- erator system.3 Illustrated NoC-based interconnect data communication for a hard- ware accelerator system.4 The speed-up of the baseline system compared to the software.5 The overall application and the kernels speed-up of the proposed system compared to the software and baseline system.6 Interconnect resource usage normalized to the resource usage for the kernels .7 Energy consumption comparison between the baseline system and the system using custom interconnect with NoC normalized to the baseline system.8 The speed-up of the baseline high performance computing system w.9 The overall application and the kernels speed-up of the proposed system compared to the software and baseline system. 97 L IST OF F IGURES xvii 6.10 Interconnect resource usage normalized to the resource usage for the kernels.

Nội dung được bảo vệ bản quyền — Tải xuống đầy đủ

Luận án tiến sĩ với tiêu đề "Thiết kế kết nối hybrid cho các bộ gia tốc phần cứng đa dạng" mang đến cái nhìn sâu sắc về việc tối ưu hóa thiết kế kết nối cho các bộ gia tốc phần cứng, giúp nâng cao hiệu suất và khả năng tương thích của hệ thống. Tài liệu này không chỉ trình bày các phương pháp thiết kế hiện đại mà còn phân tích các ứng dụng thực tiễn trong lĩnh vực công nghệ thông tin và điện tử. Độc giả sẽ tìm thấy những lợi ích rõ ràng từ việc áp dụng các kỹ thuật này, từ việc cải thiện tốc độ xử lý đến việc giảm thiểu chi phí sản xuất.

Để mở rộng thêm kiến thức về các chủ đề liên quan, bạn có thể tham khảo Luận văn thạc sĩ hcmute nghiên cứu và thiết kế phần cứng cho bộ biến đổi wavelet thuận fdwt hỗ trợ roi trong chuẩn nén ảnh jpeg2000, nơi bạn sẽ tìm thấy những nghiên cứu về thiết kế phần cứng trong lĩnh vực nén ảnh. Ngoài ra, Luận văn thạc sĩ kỹ thuật điện tử polar code decoder hardware design for 5g implemented on fpga sẽ cung cấp cái nhìn về thiết kế phần cứng cho các ứng dụng 5G, một lĩnh vực đang phát triển mạnh mẽ. Cuối cùng, Luận án tiến sĩ khoa học máy tính một số phương pháp tiếp cận cho bài toán lập lịch cá nhân sẽ giúp bạn khám phá các phương pháp lập lịch trong nghiên cứu khoa học máy tính, mở rộng thêm kiến thức về tối ưu hóa quy trình làm việc. Những tài liệu này sẽ là cơ hội tuyệt vời để bạn đào sâu hơn vào các khía cạnh khác nhau của thiết kế và ứng dụng công nghệ.