Systematic Design Methodologies for Multi-Engine Deep Learning Accelerators
Doctoral thesis, 2026

Domain-Specific Accelerators (DSAs) have become a key driver of performance and efficiency improvements in the post-Moore's Law era. These improvements stem from specializing the hardware for domain workloads and exploiting the workloads' inherent parallelism. In the domain of Deep Learning (DL), workloads (models) comprise multiple operations, known as layers, that exhibit parallelism opportunities and diverse computational characteristics. Consequently, DSAs with multiple computational units (engines) provide a natural architectural paradigm to fully exploit the specialization and parallelism potential inherent in such multi-layered models.


Multi-engine DL accelerators generally fall into two categories: model-specific and flexible. Model-specific accelerators are co-designed to efficiently execute one or a few closely related models. Flexible accelerators, by contrast, are designed to support a broad range of DL workloads. Designing and implementing accelerators in either category that fully exploit specialization and parallelism, and thus optimize performance and efficiency, requires systematic exploration based on quantitative evaluation of design alternatives. Existing multi-engine DL accelerator design approaches range from intuition-driven to exploration-based methodologies. However, even the latter typically leave key architectural parameters unexplored by fixing them a priori based on expert knowledge and intuition. In many cases, the fixed parameters are more consequential for accelerator specialization and parallelism than the explored parameters. Consequently, the full potential of the multi-engine paradigm is often left unexploited.


To fully exploit the potential of multi-engine DL accelerators, this thesis presents design methodologies that increase the extent to which key design choices are based on exploration and quantitative evaluation of design alternatives. The first contribution of this thesis is the Fixed Budget Hybrid CNN Accelerator (FiBHA). FiBHA proposes a hybrid, model-specific, multi-engine architecture and an accompanying design methodology. FiBHA targets a specific class of DL models and relies primarily on empirical analysis to exploit opportunities for specialization and parallelism. To expand the scope and co-design accelerators for a broader class of models, the work moved to a more analytical approach. The second contribution of this thesis comprises MCCM, a fast analytical cost model for evaluating model-specific multi-engine accelerators, and MCExplorer, a design space exploration framework built upon it. Together, they enable orders-of-magnitude faster evaluation and systematic exploration of model-specific multi-engine accelerator designs. Unlike existing approaches that rely on predefined design choices, MCExplorer quantitatively evaluates alternative architectural configurations across a broader design space. To further expand the scope, the work extends to flexible, in addition to model-specific, multi-engine accelerators. The third contribution of the thesis is a design methodology for flexible multi-engine DL accelerators, termed MEDEM. To support a wide range of diverse DL workloads, a flexible multi-engine accelerator must have an engine combination with complementary capabilities to ensure that different layers across these diverse workloads are processed efficiently. Existing work builds flexible multi-engine accelerators by combining expert-selected, independently optimized engines. However, independently optimized engines may perform best on largely overlapping subsets of workloads, and thus their combination does not necessarily improve overall workload coverage. MEDEM presents an alternative design methodology where the engines are co-designed, then curated to find a combination that maximizes the coverage of diverse workloads.


Using a systematic approach based on modeling and quantitative evaluation of a wider space of design alternatives, the proposed methodologies identify accelerator architectures that better exploit the specialization and parallelism inherent in DL workloads. This applies to both model-specific accelerators, as FiBHA, MCCM, and MCExplorer demonstrate, and to flexible ones, as MEDEM shows. As a result, these methodologies identify architectures that consistently outperform the state-of-the-art, achieving considerable improvements in latency, throughput, energy, and energy-delay product (EDP).

Accelerators

Co-design

Deep Neural Networks (DNNs)

FPGA

Multi-engine Accelerators

Design Methodology

Deep Learning (DL)

EDIT room EC
Opponent: Prof. Fabrizio Ferrandi, Politecnico di Milano, Italy

Author

Fareed Mohammad Qararyah

Chalmers, Computer Science and Engineering (Chalmers), Computer and Network Systems

MCExplorer: Exploring the Design Space of Multiple Compute-Engine Deep Learning Accelerators

Transactions on Architecture and Code Optimization,;Vol. 22(2025)

Journal article

An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators

IEEE International Symposium on Performance Analysis of Systems and Software Ispass,;(2025)p. 239-250

Journal article

An Efficient Hybrid Deep Learning Accelerator for Compact and Heterogeneous CNNs

Transactions on Architecture and Code Optimization,;Vol. 21(2024)

Journal article

FiBHA: Fixed Budget Hybrid CNN Accelerator

Proceedings - Symposium on Computer Architecture and High Performance Computing,;(2022)p. 180-190

Paper in proceeding

Fareed Qararyah, Mohammad Ali Maleki, and Pedro Trancoso, MEDEM: Multi-Engine DL Accelerator Design Methodology

Artificial intelligence (AI) is increasingly used in applications that shape our everyday lives. As its use grows, AI systems require substantial computing power, energy, and computer hardware. Some applications must also produce results within strict time limits. AI systems must therefore be fast while using energy and hardware resources efficiently.


One promising approach is to build AI hardware from several specialized processing units, each suited to different types of computations within AI models. The challenge is deciding how the units should be designed, how hardware resources should be divided among them, which computations should run on each unit, and how data should move between them. Today, many key decisions are still fixed in advance based on expert intuition rather than systematic exploration and comparison, meaning that better designs may be missed.


This thesis develops methods for making such decisions systematically by exploring and quantitatively comparing many design alternatives. These methods can be used both for hardware tailored to particular AI models and for more flexible hardware that supports different models. By exploring a wider range of possibilities, these methods can identify designs that use hardware more effectively, allowing faster and more hardware- and energy-efficient AI.

Principer för beräknande minnesenheter (PRIDE)

Swedish Foundation for Strategic Research (SSF) (DnrCHI19-0048), 2021-01-01 -- 2025-12-31.

Quantumstack: Programmering av kvantdatorn

Swedish Foundation for Strategic Research (SSF) (FU21-0063), 2022-06-01 -- 2027-05-31.

Very Efficient Deep Learning in IOT (VEDLIoT)

European Commission (EC) (EC/H2020/957197), 2020-11-01 -- 2023-10-31.

Areas of Advance

Information and Communication Technology

Subject Categories (SSIF 2025)

Computer Systems

DOI

10.63959/chalmers.dt/5928

ISBN

978-91-8103-471-4

Doktorsavhandlingar vid Chalmers tekniska högskola. Ny serie: 5928

Publisher

Chalmers

EDIT room EC

Opponent: Prof. Fabrizio Ferrandi, Politecnico di Milano, Italy

More information

Latest update

8/31/2026