Phạm Quốc Cường H YBRID I NTERCONNECT D ESIGN FOR H ETEROGENEOUS H ARDWARE A CCELERATORS H YBRID I NTERCONNECT D ESIGN FOR H ETEROGENEOUS H ARDWARE A CCELERATORS Proefschrift ter verkrijging van de graad van doctor aan de Technische Universiteit Delft, op gezag van de Rector Magnificus prof. Luyben, voorzitter van het College voor Promoties, in het openbaar te verdedigen op dinsdag 14 April 2015 om 12:30 uur door Cuong PHAM-QUOC Master of Engineering in Computer Science Ho Chi Minh City University of Technology - HCMUT, Vietnam geboren te Tien Giang, Vietnam. This dissertation has been approved by the Promotor: Prof.M Bertels Copromotor: Dr. Al-Ars Composition of the doctoral committee: Rector Magnificus voorzitter Prof.M Bertels Technische Universiteit Delft, promotor Dr.
Al-Ars Technische Universiteit Delft, copromotor Independent members: Prof. Charbon Technische Universiteit Delft Prof. Becker Karlsruhe Institute of Technology Prof. Dinh-Duc Vietnam National University - Ho Chi Minh City Prof.
Luigi Carro Universidade Federal do Rio Grande do Sul Dr. Silla Universitat Politècnica de València Prof.-J van der Veen Technische Universiteit Delft, reservelid Keywords: Hybrid interconnect, hardware accelerators, data communication, quan- titative data usage, automated design. Copyright © 2015 by Cuong Pham-Quoc All rights reserved. No part of this publication may be reproduced, stored in a re- trieval system, or transmitted, in any form or by any means, electronic, mechan- ical, photocopying, recording, or otherwise, without permission of the author.
ISBN 978-94-6186-448-2 Cover design: Cuong Pham-Quoc Printed in The Netherlands To my wife and my son A BSTRACT Heterogeneous multicore systems are becoming increasingly important as the need for computation power grows, especially when we are entering into the big data era. As one of the main trends in heterogeneous multicore, hardware accelerator systems provide application specific hardware circuits and are thus more energy efficient and have higher performance than general purpose pro- cessors, while still providing a large degree of flexibility. However, system perfor- mance dose not scale when increasing the number of processing cores due to the communication overhead which increases greatly with the increasing number of cores. Although data communication is a primary anticipated bottleneck for sys- tem performance, the interconnect design for data communication among the accelerator kernels has not been well addressed in hardware accelerator systems.
A simple bus or shared memory is usually used for data communication between the accelerator kernels. In this dissertation, we address the issue of interconnect design for heterogeneous hardware accelerator systems. Evidently, there are dependencies among computations, since data produced by one kernel may be needed by another kernel. Data communication patterns can be specific for each application and could lead to different types of intercon- nect.
In this dissertation, we use detailed data communication profiling to de- sign an optimized hybrid interconnect that provides the most appropriate sup- port for the communication pattern inside an application while keeping the hard- ware resource usage for the interconnect minimal. Firstly, we propose a heuristic- based approach that takes application data communication profiling into ac- count to design a hardware accelerator system with a custom interconnect. A number of solutions are considered including crossbar-based shared local mem- ory, direct memory access (DMA) supporting parallel processing, local buffers, and hardware duplication. This approach is mainly useful for embedded sys- tem where the hardware resources are limited.
Secondly, we propose an auto- mated hybrid interconnect design using data communication profiling to define an optimized interconnect for accelerator kernels of a generic hardware accel- erator system. The hybrid interconnect consists of a network-on-chip (NoC), vii viii A BSTRACT shared local memory, or both. To minimize hardware resource usage for the hybrid interconnect, we also propose an adaptive mapping algorithm to con- nect the computing kernels and their local memories to the proposed hybrid in- terconnect. Thirdly, we propose a hardware accelerator architecture to support streaming image processing.
In all presented approaches, we implement the ap- proach using a number of benchmarks on relevant reconfigurable platforms to show their effectiveness. The experimental results show that our approaches not only improve system performance but also reduce overall energy consumption compared to the baseline systems. A CKNOWLEDGMENTS It is not easy to write this last part of the dissertation, but this is an exciting period because it lets me take a careful look at the whole last four years, starting from 2011. First, I would like to thank the Vietnam International Education Develop- ment (VIED) for their funding.
Without this funding, I would not have been in the Netherlands. I would like to express special appreciation and thanks to my promoter, Prof. Koen Bertels, who had a difficult decision, but a successful one, when ac- cepting me as his Ph. At that time, my spoken English was not very good but he tried very hard to understand our Skype-based discussion.
During my time at the Computer Engineering Lab, he has introduced me to so many great ideas and has given me freedom to do my research. Koen, without you, I would have had no chance to write this dissertation. Another significant appreciation and thanks are given to my daily supervisor, but he always says that I am his friend, Dr. Zaid Al-Ars, who has guided me a lot not only in doing re- search but also in writing a paper.
Zaid, I can never forget the many hours you have spent correcting my papers. Without you, I would have no publication and, of course, no dissertation. Besides these two great persons, I would like to say thank you to Veronique from Valorisation Center - TUDelft, Lidwina - CE sec- retary, and Eef and Erik - CE system administrators, for their support. I would like to thank my colleagues, Razvan, for your DWARV compiler and, Vlad, for the Molen platform upon which I have conducted the experiments.
Thank you, Ernst, for your time translating my abstract and my proposition into Dutch. I need to say thank you to Prof. Anh-Vu Dinh-Duc. This is the third time I have written his name in my thesis.
The first and the second times were as my supervisor while this time is as a committee member. He has been there at many steps of my learning journey. I also appreciate all the committee members’ time and the remarks they gave me. Life is not only doing research.
Without relaxing time and parties, we have no energy and no ideas. So, thank you to the ANCB group, a group of Vietnamese students, for the very enjoyable parties. Those parties and relaxing time helped ix x A CKNOWLEDGMENTS me refresh my mind after the tiring working days. I am sure that I cannot say thank you to everybody who has supported me during the last four years because it would take a hundred pages, but I am also sure that I will never forget.
Let me keep your kindness in my mind. I am extremely grateful for my family and my wife’s family, especially my fa- ther in law and my mother in law who have helped me to take care of my son when I could not be at home. Without you, I would not have had the peace of mind to do my work. Last but most importantly, I would like to say thank you so much my wife and my son.
You raise me up, and you make me stronger. Without your love and your support, I cannot do anything. Our family is going to reunite in the next couple of months after a long period of connecting together through a “hybrid inter- connect” - a combination of video-calls, telephone calls, emails, social networks, and traveling. Phạm Quốc Cường Delft, April 2015 C ONTENTS Abstract vii Acknowledgments ix List of Figures xv List of Tables xix 1 Introduction 1 1.
8 2 Background and Related Work 11 2.1 On-chip Interconnect .2 System-level Hybrid Interconnect .1 Mixed topologies hybrid interconnect .2 Mixed architectures hybrid interconnect .3 Interconnect in Hardware Accelerator Systems .4 Data Communication Optimization Technique .1 Software level optimization .2 Hardware level optimization. 30 3 Communication Driven Hybrid Interconnect Design 33 3.1 Overview Hybrid Interconnect Design .2 Data Communication Driven Quantitative Execution Model .1 Baseline execution model .2 Ideal execution model .3 Parallelizing kernel processing. 41 xi xii C ONTENTS 3. 42 4 Bus-based Interconnect with Extensions 45 4.2 Bus-based hardware accelerator systems .3 Different Interconnect Solutions .1 Assumptions and definitions .2 Bus-based interconnect .3 Bus-based with a consolidation of a DMA .4 Bus-based with a consolidation of a crossbar .5 Bus-based with both a DMA and a crossbar .6 NoC-based interconnect.
62 5 Heuristic Communication-aware Hardware Optimization 63 5.2 Custom Interconnect and System Design .3 Heuristic-based algorithm. 78 6 Automated Hybrid Interconnect Design 81 6.2 Automated Hybrid Interconnect Design .1 Modeling system components.2 Custom interconnect design .3 Adaptive mapping function .1 Embedded system results .2 High performance computing results. 103 7 Accelerator Architecture for Stream Processing 105 7.2 Background and Related Work .1 Streaming image processing with hardware acceleration .2 Canny edge detection algorithm .1 Hardware-software streaming model .3 Multiple clock domains .4 Case Study: Canny Edge Detection. 117 8 Conclusions and Future Work 119 8.
122 Bibliography 125 List of Publications 143 Curriculum Vitæ 145 L IST OF F IGURES 1.1 The evolution of the on-chip interconnects .2 (a) Directly shared local memory; (b) Bus; (c) Crossbar; (d) Network- on-Chip .4 Examples of NoC topologies: (a) 2D-mesh; (b) ring; (c) hypercube; (d) tree; and (e) star.5 A generic hardware accelerator architecture .1 (a) The generic FPGA-based accelerator architecture; (b) The generic FPGA-based accelerator system with our hybrid interconnect.2 Hybrid interconnect design steps .3 Example of a QDU graph .4 The sequential diagrams for the baseline (left) and ideal execution model (right) .5 An example of data parallelism processing compared to serial pro- cessing .6 An example of instruction parallelism processing compared to se- rial processing .1 The bus is used as interconnect .2 The DMA is used as a consolidation to the bus .3 The crossbar is used as a consolidation to the bus .4 The DMA and the crossbar are used as consolidations to the bus .5 The NoC is used as interconnect of the hardware accelerators .6 The communication profiling graph generated by QUAD tool for the jpeg application. 57 xv xvi L IST OF F IGURES 4.7 Comparison between computation (Comp.), hardware accelerator execution (HW Acc.), and theoretical com- munication (Theoretical Comm.) times normalized to software time 59 4.8 Speed-up of hardware accelerators with respect to software and bus-based model .9 Comparison of resource utilization and energy consumption nor- malized to bus-based model .1 (a) HW1 and HW2 share their memories using a crossbar; (b) Struc- ture of the crossbar for the Molen architecture .2 Local buffer at HW2 .3 QUAD graph for the Canny edge detection application .4 Final system for Canny based on the Molen architecture and pro- posed solutions .t software) of hardware accelerators using Molen plat- form with and without using custom interconnect .6 The contribution of each solution to the speed-up .1 Shared local memory with and without crossbar in a hardware ac- celerator system.2 The NoC is used as interconnect of the kernels in a hardware accel- erator system.3 Illustrated NoC-based interconnect data communication for a hard- ware accelerator system.4 The speed-up of the baseline system compared to the software.5 The overall application and the kernels speed-up of the proposed system compared to the software and baseline system.6 Interconnect resource usage normalized to the resource usage for the kernels .7 Energy consumption comparison between the baseline system and the system using custom interconnect with NoC normalized to the baseline system.8 The speed-up of the baseline high performance computing system w.9 The overall application and the kernels speed-up of the proposed system compared to the software and baseline system. 97 L IST OF F IGURES xvii 6.10 Interconnect resource usage normalized to the resource usage for the kernels.