Compiler and Hardware Co-Design for Accelerator Architectures
Abstract: Heterogeneous accelerator architectures offer an efficient path to performance for compute-intensive workloads. However, full-stack integration remains difficult. We present EAAC (Extensible Accelerator Architecture), a flexible and extensible compiler and hardware architecture designed to lower the overhead of hardware-compiler co-development for rapid prototyping of hardware accelerators. EAAC targets static data-flow workloads with predictable memory access patterns. By using MLIR, we enable possible integration with a range of different frontends that emit MLIR. And by providing a minimal set of compiler functionality that enable necessary data-orchestration we lower the effort needed to get a simple implementation of a hardware acceleration unit up and running, while providing a relatively blank and un-opinionated starting point for further work. We validate this approach with a prototype GEMM accelerator combining a systolic array unit with an embedded RISC-V core, compiled end-to-end through the EAAC MLIR pipeline. On a synthetic fully-connected-layer workload, the accelerator achieves a 28x speedup in execution time over a RISC-V-only baseline, with the GEMM operation itself accounting for only 1.64\% of total execution time. We further characterize the compiler's hardware-semaphore allocation, showing that the number of semaphores required scales linearly with instruction count in the worst case, and identify this as a concrete target for future optimization.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper presents EAAC, which stands for Extensible Accelerator Architecture. EAAC is a system that helps engineers build special computer hardware called accelerators and connect that hardware to software more easily.
An accelerator is like a special-purpose machine inside a computer. Instead of asking a general-purpose processor to do every task, an accelerator can be designed to do one type of job very quickly. For example, a graphics processor is built to perform many calculations at the same time, which is useful for games and artificial intelligence.
The main idea of the paper is:
Hardware and software should be designed together so that new accelerators can be tested and used more quickly.
The authors create both:
- a hardware design that can work with different accelerator units, and
- a compiler that translates programs into instructions for that hardware.
2. What questions are the researchers asking?
The researchers are mainly trying to answer these questions:
- Can a flexible hardware system make it easier to add new accelerator units?
- Can a compiler automatically manage where data is stored and when it should be moved?
- Can different hardware units safely share memory without interfering with one another?
- How much faster is the accelerator system than using only a small general-purpose processor?
- Does the compiler use a reasonable number of synchronization tools as programs become larger?
The authors are not trying to create a completely automatic system that designs all hardware. Instead, they want to provide a useful starting point that hardware designers can extend.
3. How did they do the research?
The overall design
EAAC uses several important ideas.
MLIR is used as the basis of the compiler. MLIR is a flexible way of describing programs before they are turned into machine instructions. It is similar to writing a recipe in a form that can later be translated into instructions for different kinds of machines.
This is useful because different programs, such as machine-learning programs or image-processing programs, can potentially be converted into MLIR and then compiled for EAAC.
Managing memory
EAAC uses several levels, or tiers, of memory:
- Tier 0 is small and fast memory used directly by the accelerator units.
- Tiers 1 and 2 are larger but slower memories.
This is similar to keeping frequently used school supplies on your desk while storing less-used supplies in a cupboard. If the desk becomes full, some items are moved to the cupboard and brought back when needed.
The compiler decides:
- where each piece of data should be stored,
- when data should be moved between memory tiers, and
- when memory can be reused.
It uses DMA, or direct memory access, to move data without making the main processor handle every individual transfer.
Keeping hardware units synchronized
Several hardware units may need to use the same data. They must not read data before it is ready or overwrite data that another unit is still using.
EAAC uses hardware semaphores to solve this problem. A semaphore works like a traffic light or a permission token:
- a hardware unit waits if the data is not ready,
- it receives permission when the data becomes available,
- and it can then safely read or write the memory.
The compiler studies the program and inserts these synchronization instructions where necessary. It also removes some of them when it can prove that they are not needed.
The prototype accelerator
To test EAAC, the authors built a prototype containing three main parts:
- A GEMM accelerator, which quickly performs matrix multiplication.
- A small RISC-V processor, which handles other calculations and control tasks.
- Load and store units, which move data between the accelerator and the host computer.
GEMM means general matrix-matrix multiplication. A matrix is a rectangular grid of numbers. Matrix multiplication is used heavily in neural networks, image processing, and scientific calculations.
The GEMM unit contains a systolic array. This is a group of small calculators arranged like a grid. Numbers flow through the grid, allowing many multiplications and additions to happen at the same time—similar to an assembly line where each worker performs one step of a larger task.
The researchers tested a workload based on a small neural-network layer. It performed 24 separate matrix multiplications, followed by additional operations such as adding biases and applying ReLU, a common neural-network function.
4. What did they find?
The accelerator was much faster
The version using only the RISC-V processor took:
- 3,591,489 clock cycles
The version using the EAAC accelerator took:
- 127,348 clock cycles
This means the accelerator was about 28 times faster for the tested workload.
| Implementation | Execution time |
|---|---|
| RISC-V processor only | 3,591,489 cycles |
| EAAC accelerator | 127,348 cycles |
This is important because it shows that a specialized hardware unit can perform a calculation much more quickly than a small general-purpose processor.
Most of the accelerator's time was spent outside the main calculation
The actual GEMM calculation took only 2,092 cycles, which was about 1.64% of the accelerator's total execution time.
Most of the time was spent by the RISC-V processor doing extra integer calculations, such as preparing data and finishing the workload. This tells the researchers that the GEMM unit itself is very fast, but the surrounding software and data-management tasks still need improvement.
In other words, building a fast engine is not enough if loading, preparing, and organizing the materials takes much longer.
The compiler's semaphore use can become a problem
The researchers also tested larger and more difficult programs. In the worst case, the number of semaphores needed increased roughly in a straight-line relationship with the number of instructions.
That means a program with twice as many instructions might need about twice as many synchronization resources.
This could become a problem because hardware can provide only a limited number of semaphores. In the prototype, there were 16 hardware semaphores.
However, the researchers found that they could reduce semaphore use by examining more details about the hardware. For example, if two operations are already guaranteed to happen in order, the compiler does not need to add another semaphore between them.
The prototype used many hardware resources
The GEMM unit used most of the FPGA resources, especially the digital signal-processing units used for fast arithmetic. The prototype was designed mainly to prove that the EAAC idea works, so it was not yet highly optimized.
The maximum measured clock speed was about 104.6 MHz. The slowest path through the system was inside the RISC-V processor.
5. Why are these results important?
The results show that EAAC can support the complete process from a high-level program to working hardware:
- A program is represented using MLIR.
- The compiler assigns memory locations.
- The compiler plans data movement.
- The compiler adds synchronization instructions.
- Hardware units execute the resulting instructions.
This is valuable because creating a custom accelerator normally requires engineers to build many separate pieces of software and hardware. That can take a long time and make experimentation difficult.
EAAC provides a shared structure so that different kinds of accelerators—whether written by hand or produced by other design tools—can communicate in a consistent way.
6. What could this mean in the future?
EAAC could make it easier for researchers and engineers to design new accelerators for tasks such as:
- artificial intelligence,
- image and video processing,
- scientific calculations,
- digital signal processing, and
- other jobs with predictable data movement.
The system is especially suitable for programs where the compiler can predict what data will be needed and when. It may be less suitable for programs with many unpredictable branches, loops, or irregular memory accesses.
The paper also shows several areas for future improvement:
- reducing the number of semaphores required,
- improving the compiler's handling of the RISC-V processor,
- making data movement faster,
- reducing hardware resource use, and
- testing larger and more realistic workloads.
Simple conclusion
The paper introduces EAAC, a framework that helps custom hardware and software work together. Its compiler organizes memory and synchronization, while its hardware provides special-purpose units for fast calculations.
In the experiment, an EAAC system with a matrix-multiplication accelerator was 28 times faster than using only a small RISC-V processor. The work suggests that combining compiler design with hardware design can make accelerator development easier and faster, although the system still needs improvements before it can handle much larger and more complicated programs.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Limited workload evaluation: The architecture is evaluated primarily on one synthetic dense-layer workload consisting of 24
16×16int8 GEMM tiles; its effectiveness on representative CNNs, transformers, DSP pipelines, image-processing workloads, and non-machine-learning applications remains untested. - No comparison with relevant alternatives: The reported 28× speedup is compared only with the embedded, educational RISC-V baseline. Comparisons against a stronger RISC-V processor, vector extensions, CPU implementations, GPUs, FPGAs, SNAX, CGRAs, or HLS-generated accelerators are absent.
- Unclear contribution of the accelerator versus the baseline choice: Because the baseline uses a simple single-issue Sodor-derived core, the measured speedup may substantially reflect the weakness of the baseline rather than the general benefit of EAAC. The paper does not quantify this effect using more capable processor configurations.
- No energy or power evaluation: Although energy efficiency motivates accelerator architectures in the introduction, the paper reports neither dynamic power, static power, energy per operation, nor energy comparisons with the baseline or alternative architectures.
- No ASIC-area or technology-portability analysis: Resource results are reported for one Kintex-7 FPGA, with no ASIC synthesis, technology scaling analysis, or evaluation on other FPGA families. The area, timing, and power implications of the architecture outside this target are unresolved.
- Unvalidated claims about extensibility: The prototype contains only a GEMM unit, a load/store unit, and one RISC-V core. The effort required to integrate heterogeneous units such as convolution, pooling, DSP, or irregular control-oriented accelerators is not demonstrated.
- No quantitative integration-effort study: The paper claims that EAAC lowers hardware/software co-development cost, but does not measure development time, code size, number of required compiler changes, or engineering effort against existing frameworks.
- Insufficient evidence for frontend interoperability: Although MLIR is presented as enabling multiple frontends, the evaluation does not demonstrate compilation from TensorFlow, PyTorch, ONNX, or other independent MLIR-producing toolchains.
- Restricted support for dynamic programs: EAAC assumes statically allocated memrefs, predictable lifetimes, and statically analyzable dataflow. The paper does not characterize which dynamic shapes, indirect accesses, data-dependent control flow, loops, recursion, or runtime allocation patterns are unsupported or how they could be handled.
- Unclear correctness guarantees for synchronization: The paper describes semaphore chaining, generation tags, alias analysis, and buffer reuse, but provides no formal proof, model checking, litmus tests, or systematic validation that the compiler prevents all data hazards and deadlocks.
- Deadlock and liveness behavior is not analyzed: The effects of semaphore recycling, spin-waiting, chained predicates, and resource exhaustion on deadlock, starvation, or bounded completion are not established.
- Semaphore scalability remains unresolved: The worst-case number of required semaphores grows linearly with instruction count, and the evaluated workload was deliberately reduced because larger workloads exceeded practical semaphore resources. The maximum scalable program size and hardware requirements are therefore unknown.
- Semaphore optimization is evaluated only synthetically: The compiler evaluation uses generated worst-case buffer lifetimes and does not report results for a broad suite of real application graphs. It is unclear how often the proposed pruning strategies reduce synchronization overhead in practice.
- No systematic exploration of hardware parameters: The effects of the number of semaphores, generation-bit width, memory-tier sizes, bus widths, DMA bandwidth, array dimensions, and hardware-unit count are not evaluated.
- No synchronization-overhead breakdown: The paper does not separately measure semaphore programming, acquisition, chaining, spin-waiting, and synchronization contention costs, making it difficult to identify the dominant scalability bottlenecks.
- Memory-planning quality is not benchmarked: The modified linear-scan spilling strategy is described, but there is no comparison with optimal allocation, ILP/CP-based allocation, alternative heuristics, or other scratchpad-management techniques in terms of transfer volume and execution time.
- DMA and memory-hierarchy performance is underexplored: The evaluation does not report DMA latency, bandwidth utilization, tier-by-tier traffic, overlap efficiency, or sensitivity to memory-access patterns. Consequently, the benefit of the multi-tier hierarchy is not isolated.
- Compiler compilation time is not reported: EAAC is motivated partly by rapid iteration, but compilation latency, memory consumption, scalability with program size, and the cost of semaphore allocation are not quantified.
- The RISC-V code-generation path is a known bottleneck: Instruction-memory transfer and LLVM lowering account for significant overhead, yet no optimization or alternative code-generation path is evaluated. It remains unclear whether the reported performance would persist after realistic compiler/runtime improvements.
- The role of host initialization is ambiguous: Both measurements include binary-transfer time, but the paper does not provide results excluding initialization or distinguish one-time setup costs from steady-state execution, which limits interpretation for repeated inference or streaming workloads.
- No repeated-execution or throughput study: The evaluation appears to measure a single invocation. It does not assess amortized performance across batches, continuous streams, repeated kernel launches, or workloads where initialization and configuration can be reused.
- Limited treatment of contention and concurrency: The prototype contains a single GEMM unit, so it cannot demonstrate contention among multiple instances of the same accelerator or among several concurrent hardware units. The claimed partial out-of-order execution and distributed scheduling model therefore remain only partially validated.
- No robustness analysis under resource exhaustion: The compiler’s behavior when semaphore capacity, scratchpad capacity, instruction memory, data memory, or DMA bandwidth is insufficient is not fully specified. Feasibility diagnostics, fallback strategies, and compile-time failure modes remain open questions.
- Hardware interface assumptions are not standardized: The paper states that custom units need only a memory interface and an instruction-receive method, but does not provide a formal interface specification, timing protocol, verification methodology, or compatibility tests for independently developed units.
- Security and isolation are unexplored: Shared scratchpad memory, memory-mapped semaphores, DMA, and embedded processors are not analyzed for protection, privilege separation, faulty accelerator behavior, or isolation between mutually untrusted workloads.
- Fault handling is absent: The architecture does not discuss recovery from malformed instructions, DMA errors, hardware-unit faults, semaphore corruption, or RISC-V exceptions.
- The claimed flexibility–efficiency position is not empirically established: EAAC is argued to occupy a middle ground between fixed HLS accelerators and general-purpose CGRAs, but no quantitative comparison of programmability, performance, area, energy, or supported workload diversity is provided.
- Data-type and numerical-format support is unclear: The prototype uses int8 inputs and int32 accumulation, but the compiler and hardware framework’s support for floating point, mixed precision, sparsity, quantization variants, and wider or irregular element types is not demonstrated.
- No end-to-end application-quality validation: The synthetic workload does not establish accuracy, numerical correctness, or model-level behavior for a complete deployed application, particularly after bias addition, ReLU, and requantization.
- Reproducibility is incomplete: Although repositories are linked, the paper does not provide sufficient details about compiler versions, synthesis constraints, generated binaries, exact benchmark inputs, simulation versus FPGA execution methodology, or scripts needed to reproduce all reported results.
- The effect of hardware frequency limitations is unresolved: The prototype’s critical path lies in the RISC-V core, but the paper does not evaluate how replacing or optimizing that core changes system-level performance, resource use, or the balance between computation and orchestration overhead.
Practical Applications
The paper’s main practical contribution is a reusable hardware–compiler integration framework for statically analyzable, data-intensive workloads. Its demonstrated 28× speedup is promising, but it was measured on a synthetic dense-layer workload and a prototype FPGA implementation; therefore, applications should be understood as opportunities rather than validated production results.
Immediate Applications
- Rapid prototyping of FPGA accelerators for machine-learning inference
- Sector: AI hardware, embedded systems, edge computing.
- Teams can use EAAC’s MLIR-based pipeline to integrate custom units for GEMM, convolution-like kernels, ReLU, pooling, quantization, and other tensor operations without developing a complete memory-management and synchronization stack from scratch.
- A practical workflow would be:
- 1. Export a model or kernel into MLIR.
- 2. Bufferize tensor operations into statically allocated memrefs.
- 3. Map supported operations to EAAC hardware units.
- 4. Compile DMA transfers, scratchpad placement, semaphore dependencies, and embedded RISC-V control code.
- 5. Package the resulting FlatBuffer and load it through the runtime.
- Feasibility dependencies: The workload must have predictable memory access and mostly static data dependencies. Dynamic control flow, irregular sparse operations, and data-dependent memory access may require additional compiler and hardware support.
- Edge-AI devices for low-latency classification and anomaly detection
- Sector: Industrial monitoring, predictive maintenance, IoT, robotics, automotive.
- The prototype’s combination of an int8 systolic GEMM unit and an embedded RISC-V core could support compact neural-network inference close to sensors, reducing transfers to a CPU, GPU, or cloud service.
- Potential products include FPGA-based sensor hubs, vibration-anomaly detectors, small vision systems, and embedded controllers for factory equipment.
- The RISC-V core can execute control, activation, and requantization operations while the accelerator performs matrix multiplication.
- Feasibility dependencies: Production deployment would require validation across representative models, power measurements, real sensor workloads, fault handling, and comparison against commercial microcontrollers, NPUs, and FPGA IP.
- Reusable accelerator IP blocks for FPGA-based system-on-chip designs
- Sector: Semiconductor design, embedded computing, industrial electronics.
- EAAC’s memory-mapped data, instruction, and semaphore interfaces allow independently developed RTL, Chisel, or HLS-generated units to coexist in one accelerator fabric.
- Hardware teams could create a library of compatible units—for example, GEMM, FFT, image filters, compression, encryption, or signal-processing blocks—and connect them through the shared scratchpad and DMA hierarchy.
- This could reduce integration work when replacing one accelerator implementation with another.
- Feasibility dependencies: Each unit must implement the EAAC instruction and memory-interface conventions, and the compiler must contain suitable code-generation support. Area, timing, and bus-contention limits remain design-specific.
- Compiler-assisted DMA and scratchpad-memory management
- Sector: Embedded software, FPGA programming, high-performance computing.
- EAAC’s static allocation and lifetime-based spilling can be used immediately for applications whose working sets exceed the fastest on-chip memory.
- The compiler can determine when buffers should move between Tier 0, Tier 1, and Tier 2 memory, then emit explicit DMA operations. This provides a repeatable alternative to manually inserting transfers and buffer reuse logic.
- Likely early adopters include developers of image pipelines, DSP kernels, and fixed-size neural-network inference engines.
- Feasibility dependencies: The memory sizes, transfer bandwidths, and buffer lifetimes must be accurately represented. Poor allocation or excessive spilling can eliminate the benefit of acceleration.
- Education and research infrastructure for hardware–software co-design
- Sector: Academia and engineering training.
- The open-source compiler, accelerator, and minimal RISC-V core can serve as a laboratory platform for courses and research on MLIR, FPGA architecture, DMA scheduling, dataflow execution, hardware semaphores, and RISC-V integration.
- Students can add a functional unit and study the effect of memory hierarchy, synchronization, bus width, or systolic-array dimensions without implementing an entire toolchain.
- Feasibility dependencies: The existing prototype is FPGA-oriented and includes educationally simple components. Additional documentation, tests, simulators, and reproducible build flows would improve classroom usability.
- Benchmarking and experimentation with synchronization strategies
- Sector: Computer architecture and compiler research.
- EAAC provides a concrete testbed for evaluating hardware semaphores, dependency pruning, semaphore reuse, generation tags, and data-driven scheduling.
- Researchers can compare compiler policies by measuring semaphore count, execution time, memory traffic, FPGA resource use, and synchronization stalls.
- This is especially useful because the paper identifies semaphore allocation as a current scalability bottleneck.
- Feasibility dependencies: Results depend strongly on the target hardware’s number of semaphore entries, execution overlap, memory bandwidth, and instruction-granularity assumptions.
- Specialized DSP and image-processing pipelines
- Sector: Telecommunications, medical imaging, industrial vision, audio processing.
- Regular operations such as FIR filtering, transforms, matrix operations, pixel-wise functions, and fixed pipelines can be implemented as EAAC functional units connected by statically scheduled buffers.
- The semaphore and DMA system can support producer–consumer streaming between stages, including buffer broadcasting to multiple consumers.
- Feasibility dependencies: The current evaluation focuses on GEMM rather than DSP or image-processing kernels. New functional units, numerical formats, and domain-specific code-generation passes would be required.
Long-Term Applications
- A portable accelerator ecosystem spanning multiple hardware implementations
- Sector: Semiconductor platforms, cloud acceleration, embedded systems.
- With a sufficiently standardized EAAC interface, the same MLIR-level program could target different FPGA or ASIC configurations containing different combinations of functional units.
- Possible outcomes include accelerator modules for linear algebra, cryptography, signal processing, and machine learning that share compiler and runtime infrastructure while differing in implementation technology.
- This would resemble an accelerator-oriented software ecosystem in which hardware vendors expose compatible memory, instruction, and synchronization abstractions.
- Dependencies and risks: Hardware units currently require application-specific code generation. Portability would require a stable ISA or operation specification, capability discovery, versioning, formal interface definitions, and validation across devices.
- ASIC accelerators for energy-efficient edge inference
- Sector: Healthcare devices, robotics, autonomous systems, industrial IoT.
- EAAC’s scratchpad-centric architecture, explicit data reuse, and avoidance of general-purpose instruction execution for large kernels could be adapted into ASICs optimized for low power and predictable latency.
- An ASIC implementation could combine fixed-function systolic arrays with small programmable RISC-V controllers, potentially supporting multiple related model families rather than only one neural network.
- Dependencies and risks: The paper validates only an FPGA prototype. ASIC feasibility requires physical-design studies, power and thermal analysis, manufacturing cost justification, security review, and evidence that the added flexibility does not compromise efficiency.
- Scalable multi-accelerator dataflow systems
- Sector: Robotics, autonomous vehicles, high-performance embedded computing.
- Multiple accelerator units could execute partially out of order, communicating through shared buffers and chained hardware semaphores. A larger system might support pipelines such as sensor preprocessing → feature extraction → matrix computation → decision logic.
- This could provide predictable latency while allowing independent stages to overlap.
- Dependencies and risks: The current prototype has limited semaphore resources, and worst-case semaphore requirements grow linearly with instruction count. Larger systems will need better dependency analysis, hierarchical synchronization, semaphore virtualization, or local synchronization within functional units.
- Automatic hardware-aware scheduling and cost modeling
- Sector: Compiler technology and electronic design automation.
- Future EAAC versions could incorporate detailed models of memory bandwidth, DMA latency, functional-unit occupancy, semaphore pressure, and energy consumption.
- The compiler could then select buffer placements, operation ordering, semaphore elimination, tile sizes, and hardware-unit assignments based on an explicit performance or energy objective.
- Such a tool could support design-space exploration before committing to an FPGA or ASIC implementation.
- Dependencies and risks: Accurate cost models require hardware measurements or simulation. Optimization must avoid trading away correctness when operations overlap or reuse memory addresses.
- Automatic integration of neural-network frontends
- Sector: AI software infrastructure.
- Since MLIR can receive representations from frameworks such as TensorFlow and PyTorch, a mature EAAC toolchain could compile quantized model subgraphs directly into EAAC programs.
- A possible product would be an edge-inference compiler that partitions a model between EAAC units and the embedded RISC-V core, inserts DMA transfers, and reports unsupported operations.
- Dependencies and risks: The present work demonstrates a synthetic fully connected layer, not complete models. Real deployment requires support for convolutions, attention, normalization, dynamic shapes, model partitioning, numerical calibration, and robust fallback behavior.
- Support for real-time robotics and control pipelines
- Sector: Robotics, drones, autonomous vehicles, industrial automation.
- Deterministic data-driven scheduling and explicit synchronization could be useful for fixed-rate pipelines involving image or lidar preprocessing, sensor fusion, control-law evaluation, and actuator command preparation.
- Hardware units could be specialized for perception or signal processing while the RISC-V core handles supervisory control.
- Dependencies and risks: Real-time use requires bounded worst-case execution time, interrupt handling, fault recovery, priority management, and guarantees against deadlock or semaphore starvation. These properties are not established by the current prototype.
- Acceleration of privacy-sensitive or safety-critical local computation
- Sector: Healthcare, finance, defense, and public infrastructure.
- Local FPGA or ASIC accelerators could process medical signals, biometric data, fraud-detection features, or industrial telemetry without sending raw data to cloud systems.
- Static memory allocation and constrained data movement may also simplify certain forms of security auditing and information-flow analysis.
- Dependencies and risks: The paper does not evaluate confidentiality, side-channel resistance, secure boot, isolation, or formal verification. These would be required before deployment in regulated or safety-critical environments.
- Compiler and architecture standards for open hardware research
- Sector: Academia, open-source EDA, public research policy.
- EAAC could contribute to a broader open ecosystem in which researchers publish accelerator units together with MLIR dialects, interface descriptions, runtime packages, and reproducible FPGA configurations.
- Funding agencies and universities could use such platforms to make accelerator results more comparable and easier to reproduce than designs tied to proprietary toolchains.
- Dependencies and risks: This requires sustained maintenance, standardized benchmarks, compatibility testing, licensing clarity, and independent replication. The current evaluation is limited in workload size and hardware diversity, so broader claims about general-purpose accelerator performance would be premature.
Glossary
- Accelerator architecture: A hardware design specialized to speed up particular computational workloads. “Heterogeneous accelerator architectures offer an efficient path to performance for compute-intensive workloads.”
- Address aliasing: A condition in which different names or references resolve to the same memory address. “If the compiler finds an instance of address aliasing within a sliding window, it locates the semaphores guarding the given buffers and chains them in order of use.”
- Atomic operation: An indivisible operation that completes without interference from concurrent operations. “To accelerate atomic access to the semaphore pairs, EAAC utilizes hardware semaphores.”
- AXI4-Stream: A widely used streaming interconnect protocol for transferring data between hardware components. “A load and store unit that provides fully bidirectional data transfers to and from the host using AXI4-Stream interfaces.”
- Bufferization: A compiler transformation that converts abstract tensor values into explicit memory buffers. “To convert tensors into a memory representation that is closer to hardware, bufferization on tensors is used.”
- Calyx: A hardware compiler infrastructure and accelerator design framework. “Existing frameworks address this primarily by raising the abstraction level of accelerator implementation through high-level synthesis (HLS), accelerator design languages, or hardware compiler infrastructures such as Allo, Calyx, and CIRCT.”
- Chisel: A hardware-construction language embedded in Scala that supports parameterized hardware generators. “This allows accelerators produced through different methods, including hand-written RTL, Chisel generators, and HLS-generated designs, to coexist within a shared framework.”
- Coarse-grained reconfigurable array (CGRA): A reconfigurable architecture composed of relatively large programmable processing elements. “Coarse-Grained Reconfigurable Arrays (CGRAs) occupy a different point in the design space.”
- Compiler backend: The compiler component that translates an intermediate representation into target-specific code or instructions. “Domain-specific hardware acceleration has historically complicated compiler backend development.”
- Compiler–hardware co-design: Joint development of compiler software and hardware so that each is optimized for the capabilities of the other. “We propose EAAC (Extensible Accelerator ArChitecture), a framework for hardware-compiler co-development focused on reducing the cost of hardware/software integration.”
- Control and status register (CSR): A processor register used to configure hardware or report its state. “SNAX utilizes RISC-V processor cores for synchronization, configuration, and control of the hardware accelerator units, specifically through control and status registers in the RISC-V core.”
- Data dependency: A relationship in which one operation requires data produced or modified by another operation. “EAAC is specifically designed for workloads that are optimal for acceleration, often characterized by regular and predictable memory access patterns, minimal control flow, and statically analyzable data dependencies.”
- Data-flow graph: A graph representing operations as nodes and data dependencies as edges. “To ensure correctness, the EAAC compiler will allocate a virtual semaphore for every edge between a producer and consumer node in the dataflow graph.”
- Data-driven scheduling: Scheduling in which an operation executes when its required input data becomes available. “EAAC operates exclusively using data-driven scheduling; unlike a typical von Neumann architecture, it does not have a program counter, branching, or looping.”
- Direct memory access (DMA): Hardware-supported transfer of data between memory regions without requiring continuous processor involvement. “The EAAC compiler is capable of static memory allocation and manages a multi-tiered memory hierarchy using explicit direct memory access (DMA) control between layers.”
- Decoupled access-execute: An architecture that separates data movement or access operations from computation operations so they can proceed independently. “A flexible hardware architecture based on decoupled access-execute, with phased execution allowing for data transfers to overlap with execution.”
- Dialect: In MLIR, an extensible collection of operations, types, and attributes for a particular abstraction or domain. “Developers can define custom IR operations and transformation passes to support compiling for targets other than classical assembly languages.”
- Digital signal processing: Computational manipulation of sampled signals such as audio, video, or sensor data. “Examples include GEMM, digital signal processing, and image processing.”
- Embedded processor: A processor integrated into a larger hardware system to perform local control or computation. “We demonstrate this design through a prototype GEMM accelerator combining a systolic array unit with an embedded RISC-V core.”
- Fan-out: The number of destination components receiving a signal or data value from one source. “This mechanism also extends to support broadcasting behavior for instruction outputs with a fan-out of more than one.”
- FlatBuffer: A serialized binary data format designed for efficient access without requiring unpacking or parsing the entire object. “The final output from the compiler is in the form of a FlatBuffer, which is imported into an assembler and runtime that configures and programs the accelerator.”
- Formal description: A precise, machine-processable specification of a system’s behavior or interface. “Act dissolves tensor operations into a set of IR-like primitive tensor operations; ISA instructions are then defined in terms of these primitives.”
- General matrix multiplication (GEMM): Matrix multiplication, usually expressed as or , and a core workload in machine learning. “We validate this approach with a prototype GEMM accelerator combining a systolic array unit with an embedded RISC-V core.”
- Generation tag: Metadata distinguishing different uses or allocations of the same physical synchronization or storage resource. “To prevent aliasing, the compiler appends a semaphore generation tag to each new allocation.”
- High-level synthesis (HLS): The automated generation of hardware descriptions from high-level algorithmic code. “Existing frameworks address this primarily by raising the abstraction level of accelerator implementation through high-level synthesis (HLS), accelerator design languages, or hardware compiler infrastructures such as Allo, Calyx, and CIRCT.”
- Intermediate representation (IR): A structured compiler representation between source code and target machine code. “MLIR is intended to function as an intermediate representation supported by multiple frontends.”
- Instruction set architecture (ISA): The programmer-visible specification of a processor’s instructions, registers, data types, and behavior. “Act dissolves tensor operations into a set of IR-like primitive tensor operations; ISA instructions are then defined in terms of these primitives.”
- Linear scan: A compiler algorithm commonly used for register allocation and interval-based resource assignment. “Data movement scheduling is integrated into the memory allocation algorithm using a modified approach inspired by linear-scan.”
- Lowering: A compiler transformation that translates a higher-level representation into a lower-level representation. “In this pipeline, operations are first lowered into the LLVM dialect, which is later compiled into RISV-V assembly.”
- Memory-mapped: Organized so that hardware devices or control registers are accessed through ordinary memory addresses. “The architecture additionally allows for a degree of out-of-order processing between different hardware units.”
- MemRef: An MLIR type representing a reference to a memory buffer together with metadata such as shape and memory space. “Memrefs are MLIR's core dialect for memory references, and act as pointers to buffers in memory with additional information about element types, shapes, and memory space.”
- Memory hierarchy: A system of memory levels with differing capacities, latencies, and access costs. “EAAC organizes its memories into a configurably tiered hierarchy interleaved with programmable DMA units.”
- MLIR: A compiler infrastructure and extensible intermediate-representation framework for heterogeneous and domain-specific systems. “MLIR is a compiler framework designed with heterogeneous and domain-specific computing in mind.”
- Out-of-order execution: Execution in which operations may run in an order different from their original program order when their dependencies permit it. “The architecture additionally allows for a degree of out-of-order processing between different hardware units.”
- Predicate: A Boolean condition used to determine whether an action should occur. “These events are logged and can be used to create predicates, such as "semaphore x generated event y".”
- Read-after-write (RAW) hazard: A dependency hazard in which a read must not occur before an earlier write to the same location completes. “In the subsequent pass, the compiler analyzes the address space of the accelerator (SPM Tier 0) to identify cases where read-after-write (RAW) errors due to address-space reuse are likely to occur.”
- ReLU: The rectified linear unit activation function, defined as . “24 independent int8 matmul tiles are summed in int32, then bias-added, ReLU'd, and requantized back to int8.”
- Requantization: Conversion of numerical values from one quantized precision or scale to another. “24 independent int8 matmul tiles are summed in int32, then bias-added, ReLU'd, and requantized back to int8.”
- Register renaming: A processor technique that maps architectural registers to distinct physical registers to eliminate false dependencies. “The use of generation tags can be compared to register renaming in Tomasulo's algorithm, where multiple versions of the same architectural register can coexist and are distinguished by different tags.”
- RISC-V: An open instruction set architecture designed around a modular and extensible specification. “We validate this approach with a prototype GEMM accelerator combining a systolic array unit with an embedded RISC-V core.”
- Scratchpad memory (SPM): Programmer- or compiler-managed on-chip memory that provides predictable, low-latency access. “Because individual instructions represent large units of work, we chose to construct the architecture as a memory-to-memory system---with the Tier 0 SPM being the working memory for the hardware units.”
- Semaphore: A synchronization mechanism based on a counter or state that regulates access to shared resources. “The synchronization primitives chosen are counting or binary semaphore pairs.”
- Systolic array: A regular grid of processing elements that passes data rhythmically between neighboring elements for highly parallel computation. “We validate this approach with a prototype GEMM accelerator combining a systolic array unit with an embedded RISC-V core.”
- Tensor: A multidimensional numerical data structure commonly used to represent machine-learning data. “EAAC and Act both treat tensors as first-class datatypes.”
- TileLink: A hardware interconnect protocol used to connect processor cores and memory-mapped devices. “Sodor has been modified with a TileLink data-memory interface that allows it to interact with the hardware semaphores and the main Tier 0 scratchpad memory.”
- Triggered instruction: An instruction whose execution is enabled by satisfaction of a specified event or condition. “Semaphore initializations are issued when a given set of predicates is satisfied, in a fashion inspired by triggered instructions.”
- Virtual semaphore: A compiler-level synchronization object that has not yet been assigned to a physical hardware semaphore. “In practice, the compiler manages the state of buffers using virtual semaphores, which are semaphores not assigned to physical hardware semaphores yet.”
- Von Neumann architecture: A computer architecture in which instructions and data are stored in memory and execution is generally controlled by a program counter. “EAAC operates exclusively using data-driven scheduling; unlike a typical von Neumann architecture, it does not have a program counter, branching, or looping.”



