Systematic Design Methodologies for Multi-Engine Deep Learning Accelerators
Doctoral thesis, 2026
Multi-engine DL accelerators generally fall into two categories: model-specific and flexible. Model-specific accelerators are co-designed to efficiently execute one or a few closely related models. Flexible accelerators, by contrast, are designed to support a broad range of DL workloads. Designing and implementing accelerators in either category that fully exploit specialization and parallelism, and thus optimize performance and efficiency, requires systematic exploration based on quantitative evaluation of design alternatives. Existing multi-engine DL accelerator design approaches range from intuition-driven to exploration-based methodologies. However, even the latter typically leave key architectural parameters unexplored by fixing them a priori based on expert knowledge and intuition. In many cases, the fixed parameters are more consequential for accelerator specialization and parallelism than the explored parameters. Consequently, the full potential of the multi-engine paradigm is often left unexploited.
To fully exploit the potential of multi-engine DL accelerators, this thesis presents design methodologies that increase the extent to which key design choices are based on exploration and quantitative evaluation of design alternatives. The first contribution of this thesis is the Fixed Budget Hybrid CNN Accelerator (FiBHA). FiBHA proposes a hybrid, model-specific, multi-engine architecture and an accompanying design methodology. FiBHA targets a specific class of DL models and relies primarily on empirical analysis to exploit opportunities for specialization and parallelism. To expand the scope and co-design accelerators for a broader class of models, the work moved to a more analytical approach. The second contribution of this thesis comprises MCCM, a fast analytical cost model for evaluating model-specific multi-engine accelerators, and MCExplorer, a design space exploration framework built upon it. Together, they enable orders-of-magnitude faster evaluation and systematic exploration of model-specific multi-engine accelerator designs. Unlike existing approaches that rely on predefined design choices, MCExplorer quantitatively evaluates alternative architectural configurations across a broader design space. To further expand the scope, the work extends to flexible, in addition to model-specific, multi-engine accelerators. The third contribution of the thesis is a design methodology for flexible multi-engine DL accelerators, termed MEDEM. To support a wide range of diverse DL workloads, a flexible multi-engine accelerator must have an engine combination with complementary capabilities to ensure that different layers across these diverse workloads are processed efficiently. Existing work builds flexible multi-engine accelerators by combining expert-selected, independently optimized engines. However, independently optimized engines may perform best on largely overlapping subsets of workloads, and thus their combination does not necessarily improve overall workload coverage. MEDEM presents an alternative design methodology where the engines are co-designed, then curated to find a combination that maximizes the coverage of diverse workloads.
Using a systematic approach based on modeling and quantitative evaluation of a wider space of design alternatives, the proposed methodologies identify accelerator architectures that better exploit the specialization and parallelism inherent in DL workloads. This applies to both model-specific accelerators, as FiBHA, MCCM, and MCExplorer demonstrate, and to flexible ones, as MEDEM shows. As a result, these methodologies identify architectures that consistently outperform the state-of-the-art, achieving considerable improvements in latency, throughput, energy, and energy-delay product (EDP).
Accelerators
Co-design
Deep Neural Networks (DNNs)
FPGA
Multi-engine Accelerators
Design Methodology
Deep Learning (DL)
Author
Fareed Mohammad Qararyah
Chalmers, Computer Science and Engineering (Chalmers), Computer and Network Systems
MCExplorer: Exploring the Design Space of Multiple Compute-Engine Deep Learning Accelerators
Transactions on Architecture and Code Optimization,;Vol. 22(2025)
Journal article
An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators
IEEE International Symposium on Performance Analysis of Systems and Software Ispass,;(2025)p. 239-250
Journal article
An Efficient Hybrid Deep Learning Accelerator for Compact and Heterogeneous CNNs
Transactions on Architecture and Code Optimization,;Vol. 21(2024)
Journal article
FiBHA: Fixed Budget Hybrid CNN Accelerator
Proceedings - Symposium on Computer Architecture and High Performance Computing,;(2022)p. 180-190
Paper in proceeding
Fareed Qararyah, Mohammad Ali Maleki, and Pedro Trancoso, MEDEM: Multi-Engine DL Accelerator Design Methodology
One promising approach is to build AI hardware from several specialized processing units, each suited to different types of computations within AI models. The challenge is deciding how the units should be designed, how hardware resources should be divided among them, which computations should run on each unit, and how data should move between them. Today, many key decisions are still fixed in advance based on expert intuition rather than systematic exploration and comparison, meaning that better designs may be missed.
This thesis develops methods for making such decisions systematically by exploring and quantitatively comparing many design alternatives. These methods can be used both for hardware tailored to particular AI models and for more flexible hardware that supports different models. By exploring a wider range of possibilities, these methods can identify designs that use hardware more effectively, allowing faster and more hardware- and energy-efficient AI.
Principer för beräknande minnesenheter (PRIDE)
Swedish Foundation for Strategic Research (SSF) (DnrCHI19-0048), 2021-01-01 -- 2025-12-31.
Quantumstack: Programmering av kvantdatorn
Swedish Foundation for Strategic Research (SSF) (FU21-0063), 2022-06-01 -- 2027-05-31.
Very Efficient Deep Learning in IOT (VEDLIoT)
European Commission (EC) (EC/H2020/957197), 2020-11-01 -- 2023-10-31.
Areas of Advance
Information and Communication Technology
Subject Categories (SSIF 2025)
Computer Systems
DOI
10.63959/chalmers.dt/5928
ISBN
978-91-8103-471-4
Doktorsavhandlingar vid Chalmers tekniska högskola. Ny serie: 5928
Publisher
Chalmers
EDIT room EC
Opponent: Prof. Fabrizio Ferrandi, Politecnico di Milano, Italy