Papers
Topics
Authors
Recent
Search
2000 character limit reached

Compiler and Hardware Co-Design for Accelerator Architectures

Published 24 Sep 2026 in cs.AR | (2609.30099v1)

Abstract: Heterogeneous accelerator architectures offer an efficient path to performance for compute-intensive workloads. However, full-stack integration remains difficult. We present EAAC (Extensible Accelerator Architecture), a flexible and extensible compiler and hardware architecture designed to lower the overhead of hardware-compiler co-development for rapid prototyping of hardware accelerators. EAAC targets static data-flow workloads with predictable memory access patterns. By using MLIR, we enable possible integration with a range of different frontends that emit MLIR. And by providing a minimal set of compiler functionality that enable necessary data-orchestration we lower the effort needed to get a simple implementation of a hardware acceleration unit up and running, while providing a relatively blank and un-opinionated starting point for further work. We validate this approach with a prototype GEMM accelerator combining a systolic array unit with an embedded RISC-V core, compiled end-to-end through the EAAC MLIR pipeline. On a synthetic fully-connected-layer workload, the accelerator achieves a 28x speedup in execution time over a RISC-V-only baseline, with the GEMM operation itself accounting for only 1.64\% of total execution time. We further characterize the compiler's hardware-semaphore allocation, showing that the number of semaphores required scales linearly with instruction count in the worst case, and identify this as a concrete target for future optimization.

Summary

  • The paper presents EAAC, a compiler-hardware framework designed to reduce the integration cost of custom accelerators by addressing memory orchestration, synchronization, and code generation challenges in accelerator development.
  • EAAC uses an extensible hardware interface paired with an MLIR-based compiler to support memory, synchronization, and control mechanisms for custom hardware units. Specifically, the prototypes aim to validate this approach through an eight-bit systolic GEMM unit implementation reaching a 28x speedup over a RISC-V-only model.
  • EAAC's memory-to-memory accelerator model decouples access and execution using a tiered scratchpad hierarchy and specialized DMA engines, which improves data flow and enhances specific tasks, such as matrix operations, activation functions, and pooling.

Research problem and design objective

“Compiler and Hardware Co-Design for Accelerator Architectures” presents EAAC, an extensible compiler–hardware framework intended to reduce the integration cost of custom accelerators into a complete execution stack (2609.30099). The paper addresses a specific gap in accelerator-development infrastructure: HLS and accelerator design languages reduce the cost of implementing hardware, but they generally produce application-specific hardware and do not, by themselves, solve memory orchestration, synchronization, code generation, runtime configuration, or integration with heterogeneous functional units. Conversely, conventional compiler infrastructures provide mature software abstractions but offer limited support for novel spatial architectures with explicit scratchpad management and distributed execution.

EAAC deliberately does not attempt to synthesize hardware automatically. Its objective is narrower and more pragmatic: provide a minimal, extensible hardware interface and a corresponding MLIR-based compiler pipeline through which independently developed accelerator units can share memory, synchronization, and control mechanisms. The target workload class consists of statically analyzable data-flow applications with regular memory accesses, limited control flow, and predictable buffer lifetimes. GEMM, DSP kernels, and image-processing pipelines are representative examples.

The design occupies a distinct point in the accelerator design space. It is more flexible than an HLS-generated accelerator specialized for one application, but less general than a CGRA or a conventional processor. The cost of this flexibility is that hardware developers remain responsible for implementing domain-specific functional units and their code-generation support. EAAC therefore provides infrastructure for co-development rather than an automated hardware-generation methodology.

EAAC architecture

EAAC implements a memory-to-memory accelerator model based on decoupled access and execute. Instead of exposing scalar registers as the primary instruction operands, instructions manipulate statically allocated memory buffers. Functional units interpret those buffers according to their operation semantics. This makes coarse-grained operations such as GEMM, activation functions, and pooling natural instruction units.

The architecture contains an arbitrary number of functional units, a tiered scratchpad hierarchy, programmable DMA engines, a load/store interface to the host, and a hardware semaphore array. Tier 0 is directly visible to accelerator units and acts as the shared working memory. Lower tiers provide additional capacity, with DMA operations explicitly moving buffers between tiers. The system has no program counter, branches, or loops in the conventional sense. Instead, execution is driven by an instruction stream and by data availability.

Figure 1

Figure 1: EAAC’s abstract architecture, including functional units, tiered memory, DMA interfaces, and hardware semaphores.

The central control mechanism is data-driven synchronization through binary or counting semaphore pairs. A semaphore represents the availability state of a buffer or buffer segment. Producers and consumers acquire and release the corresponding semaphore state, allowing functional units to operate asynchronously while preventing data hazards. Counting semaphores additionally support streaming access, in which a buffer is produced and consumed incrementally rather than as a single indivisible object.

EAAC extends this mechanism with semaphore chaining, event predicates, broadcasting, and generation tags. Chaining allows one semaphore transition to trigger another, enabling fences and producer-to-multiple-consumer communication. Generation tags distinguish successive logical allocations that reuse the same physical semaphore address. This is analogous to register renaming: a finite physical resource can represent multiple logically distinct synchronization instances without permitting stale operations to alias a newer allocation.

The use of dedicated hardware semaphore units, rather than memory-based atomic operations, is intended to avoid synchronization traffic competing with ordinary data traffic. The resulting design assumes that synchronization latency and access behavior should remain predictable. The paper nevertheless acknowledges that explicit synchronization introduces overhead, which motivates compiler-side semaphore pruning.

MLIR compilation pipeline

EAAC uses MLIR as the compiler substrate, enabling frontends that already lower into MLIR dialects to target the framework. The pipeline begins after tensor bufferization, when operations manipulate statically allocated memrefs with explicit lifetimes. This choice provides the compiler with the information required for memory placement, lifetime analysis, DMA insertion, and dependency construction. MLIR’s extensibility is particularly important because EAAC must support both generic data movement and target-specific code generation.

The compiler consists of three principal phases: static memory allocation and data planning, dependency and semaphore management, and code generation. MiniMalloc supplies the underlying static allocation mechanism. EAAC overlays a tier-aware spilling algorithm on top of MiniMalloc rather than modifying the allocator itself. Buffers are processed according to lifetime start points; when a tier lacks sufficient capacity, the compiler spills a buffer whose next use is farthest away, inserts a transfer to a lower tier, and reloads the buffer before its next use.

Figure 2

Figure 2: Tier-aware buffer spilling based on buffer lifetimes and next-use information.

This strategy is a form of linear-scan allocation extended to a multi-tier scratchpad hierarchy. Buffers are normally single-assignment and exist in only one tier at a time, except for constants that are statically placed in lower memory. Once the allocator determines offsets and spill points, subsequent passes materialize those decisions as explicit DMA operations.

The dependency phase initially constructs virtual semaphores for producer–consumer edges in the data-flow graph. This conservative strategy guarantees correctness under arbitrary partial out-of-order execution, but it can generate many synchronization objects. The compiler subsequently eliminates semaphores when it can establish that operations are already serialized by hardware behavior, by sufficient graph distance, or by known execution ordering between functional units. It also identifies potential read-after-write hazards caused by scratchpad address reuse and chains the relevant semaphore instances to preserve ordering.

Finally, virtual semaphores are mapped to a finite set of physical semaphore addresses. When physical addresses are reused, generation tags prevent aliasing between logically distinct instances. The allocator selects addresses based on an estimated release time, attempting to minimize serialization introduced by resource recycling. This design makes the compiler’s quality dependent on hardware-specific knowledge: the more accurately the compiler models unit ordering and run-ahead behavior, the more aggressively it can remove synchronization.

Code generation is split between accelerator-specific lowering and RISC-V lowering. Users provide the code generation needed for application-specific hardware units, while EAAC supplies common passes for memory orchestration and synchronization. Unsupported operations can be offloaded to an embedded RISC-V processor. The final accelerator program is serialized as a FlatBuffer, and the runtime configures the hardware and loads any generated RISC-V code.

Prototype accelerator

The prototype validates the framework with an 8-bit, 16×1616 \times 16 GEMM unit implemented as a systolic array. It also integrates a small five-stage, single-issue RISC-V core, load/store units connected through AXI4-Stream interfaces, a tiered memory system, and a bank of 16 hardware semaphores.

Figure 3

Figure 3: Prototype EAAC accelerator integrating a systolic GEMM unit, RISC-V control core, memory system, and semaphore subsystem.

The GEMM unit performs the matrix-intensive portion of a synthetic dense-layer workload. The complete workload comprises 24 independent 16×1616 \times 16 int8 matrix multiplications, accumulation into int32, bias addition, ReLU, and requantization to int8. The GEMM unit and the RISC-V core therefore divide the workload: the systolic array handles matrix multiplication, while the RISC-V core performs integer operations and residual control-oriented computation.

The prototype uses a 512-bit memory-transfer bus and a separate 16-bit semaphore bus. It contains 64 kB of Tier 0 memory and 128 kB each of Tier 1 and Tier 2 memory. The RISC-V core has 64 kB of instruction memory and 4 kB of data memory. The implementation reaches a maximum reported clock frequency of 104.602 MHz, with the critical path located in the RISC-V core rather than in the GEMM unit or semaphore subsystem.

Resource utilization is dominated by the systolic array. The complete design consumes 64,414 LUTs, 71,891 flip-flops, 16 block-RAM units of the reported B36 category, and 266 DSPs. The GEMM unit alone accounts for 40,384 LUTs and 258 DSPs, or 97% of the DSP resources. This distribution is consistent with the paper’s positioning of the prototype as a proof-of-concept integration vehicle rather than an optimized accelerator implementation. It also indicates that the infrastructure overhead is not negligible: the load/store units, semaphore system, memory system, and RISC-V core collectively consume substantial logic even though the systolic array remains the principal computational resource.

Evaluation of end-to-end performance

The principal performance comparison is between execution entirely on the embedded RISC-V baseline and execution on the EAAC configuration combining the RISC-V core with the systolic GEMM unit.

Configuration Measured cycles
RISC-V-only baseline 3,591,489
EAAC accelerator total 127,348
GEMM computation within EAAC 2,092
EAAC RISC-V boot 24,711
EAAC RISC-V computation 102,057

The integrated accelerator reduces total execution time from 3,591,489 to 127,348 cycles, corresponding to a reported 28x speedup. The result demonstrates successful end-to-end compilation and execution through the EAAC stack rather than merely measuring an isolated GEMM kernel.

The detailed breakdown also qualifies the result. GEMM consumes only 2,092 cycles, or 1.64% of total accelerator execution time. The embedded RISC-V computation accounts for 102,057 cycles, while binary-transfer and boot overhead account for an additional 24,711 cycles. Thus, the measured speedup is not primarily a consequence of the GEMM unit dominating total execution time; it reflects the combination of hardware offload, workload partitioning, and the weakness of the intentionally simple RISC-V baseline.

This distinction is important for interpreting the result. The 28x improvement establishes that EAAC can connect an MLIR program, static memory planning, DMA orchestration, semaphore synchronization, generated RISC-V code, and a custom accelerator into a functioning system. It does not establish that the current architecture provides high utilization of the systolic array. In fact, the low GEMM fraction shows that the prototype is substantially limited by non-GEMM work and software initialization overhead.

The execution-time comparison also includes instruction-memory transfer costs for both configurations. The accelerator’s binary transfer is longer because the internal EAAC-to-LLVM lowering path is not optimized. Consequently, the reported result includes a compiler/runtime overhead that is unfavorable to the accelerator and may decrease as the lowering pipeline improves. At the same time, the evaluation uses a small synthetic workload and a weak scalar baseline, so the result should not be interpreted as a comparison against a high-performance CPU, vector processor, GPU, or production accelerator.

Semaphore scalability

The paper’s second evaluation examines compiler behavior under increasingly large programs with intentionally unfavorable buffer lifetimes. In the worst case, the number of distinct hardware semaphores grows linearly with the number of operations.

Figure 4

Figure 4: Worst-case semaphore demand as a function of total operation count, before and after hardware-aware pruning.

This result exposes the primary current scalability concern in EAAC. A virtual semaphore is initially associated with each data-flow edge, and long-lived buffers preserve synchronization state across many operations. Without sufficiently aggressive pruning, larger programs can exhaust the finite semaphore bank and become infeasible even when the underlying computation is structurally suitable for acceleration.

The paper makes a deliberately strong observation: for the prototype, much of this semaphore demand is unnecessary because only one GEMM unit is present, so GEMM operations are intrinsically serialized. Hardware-aware analysis breaks the linear growth observed in the conservative configuration by removing semaphores between operations whose order is already guaranteed. Additional pruning is possible when different units have known run-ahead or ordering properties.

This result supports the paper’s co-design thesis. Semaphore allocation cannot be optimized effectively using compiler information alone; it depends on the number and behavior of functional units, their execution ordering, and the degree of internal serialization. Conversely, the hardware exposes enough structure for the compiler to remove synchronization that would otherwise be required by a conservative data-flow model. The remaining question is whether this approach remains effective for architectures with multiple instances of the same unit, dynamic resource contention, or genuinely concurrent streaming pipelines.

Limitations and open questions

EAAC is explicitly restricted to static data-flow workloads. Programs with dynamic allocation, irregular memory accesses, substantial control flow, or data-dependent iteration are outside the demonstrated compilation model. The architecture’s absence of a program counter and its reliance on statically analyzable lifetimes simplify hardware control but constrain applicability.

The prototype evaluation is also narrow. It uses one synthetic dense-layer workload, one systolic GEMM unit, a small RISC-V core, and an FPGA implementation whose critical path lies in the control processor. The 28x speedup is therefore a validation of integration rather than a broad architectural comparison. No energy measurements, area-normalized performance analysis, comparison with an optimized CPU or GPU, or evaluation across multiple real workloads is reported.

The compiler remains dependent on target-specific hardware knowledge. Its initial virtual-semaphore construction can scale poorly, and the selected workload was reduced because larger programs produced infeasible semaphore requirements under the available optimization. This is a direct limitation rather than a peripheral implementation detail: semaphore allocation currently constrains the size of compilable workloads.

Memory planning also assumes well-defined static buffer lifetimes and single-tier residency. The spilling heuristic selects the buffer with the latest next use, but the paper does not provide a comparative evaluation against alternative multi-tier allocation algorithms, nor does it quantify DMA traffic, spill-induced latency, or memory-energy costs. Similarly, the correctness and performance behavior of semaphore chaining and generation-tag reuse are demonstrated conceptually and through the prototype, but not evaluated under a broad range of concurrent producer–consumer topologies.

The paper leaves open how EAAC would scale when the hardware contains multiple parallel instances of a functional unit, when synchronization latency becomes significant relative to coarse-grained computation, and when the compiler must jointly optimize tiling, memory placement, DMA scheduling, and accelerator occupancy. It also leaves open whether a more expressive MLIR dialect could make hardware-unit integration less dependent on manually implemented code-generation passes.

Conclusion

EAAC presents a deliberately minimal framework for compiler–hardware co-development. Its contribution is the integration of MLIR-based compilation, static multi-tier memory planning, explicit DMA orchestration, distributed hardware semaphores, and heterogeneous functional units under a common memory-mapped interface. The prototype demonstrates a complete compilation-to-execution path and reports a 28x speedup over a small RISC-V-only baseline.

The evaluation simultaneously identifies the framework’s principal constraint: synchronization-resource demand can grow linearly with instruction count under conservative allocation. Hardware-aware semaphore pruning mitigates this behavior, but the result depends on detailed knowledge of accelerator-unit ordering and remains an important compiler optimization target. EAAC therefore offers a credible infrastructure layer for statically scheduled accelerator systems, while leaving scalability across larger, more concurrent, and less regular workloads as an open empirical question.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper presents EAAC, which stands for Extensible Accelerator Architecture. EAAC is a system that helps engineers build special computer hardware called accelerators and connect that hardware to software more easily.

An accelerator is like a special-purpose machine inside a computer. Instead of asking a general-purpose processor to do every task, an accelerator can be designed to do one type of job very quickly. For example, a graphics processor is built to perform many calculations at the same time, which is useful for games and artificial intelligence.

The main idea of the paper is:

Hardware and software should be designed together so that new accelerators can be tested and used more quickly.

The authors create both:

  • a hardware design that can work with different accelerator units, and
  • a compiler that translates programs into instructions for that hardware.

2. What questions are the researchers asking?

The researchers are mainly trying to answer these questions:

  • Can a flexible hardware system make it easier to add new accelerator units?
  • Can a compiler automatically manage where data is stored and when it should be moved?
  • Can different hardware units safely share memory without interfering with one another?
  • How much faster is the accelerator system than using only a small general-purpose processor?
  • Does the compiler use a reasonable number of synchronization tools as programs become larger?

The authors are not trying to create a completely automatic system that designs all hardware. Instead, they want to provide a useful starting point that hardware designers can extend.

3. How did they do the research?

The overall design

EAAC uses several important ideas.

MLIR is used as the basis of the compiler. MLIR is a flexible way of describing programs before they are turned into machine instructions. It is similar to writing a recipe in a form that can later be translated into instructions for different kinds of machines.

This is useful because different programs, such as machine-learning programs or image-processing programs, can potentially be converted into MLIR and then compiled for EAAC.

Managing memory

EAAC uses several levels, or tiers, of memory:

  • Tier 0 is small and fast memory used directly by the accelerator units.
  • Tiers 1 and 2 are larger but slower memories.

This is similar to keeping frequently used school supplies on your desk while storing less-used supplies in a cupboard. If the desk becomes full, some items are moved to the cupboard and brought back when needed.

The compiler decides:

  • where each piece of data should be stored,
  • when data should be moved between memory tiers, and
  • when memory can be reused.

It uses DMA, or direct memory access, to move data without making the main processor handle every individual transfer.

Keeping hardware units synchronized

Several hardware units may need to use the same data. They must not read data before it is ready or overwrite data that another unit is still using.

EAAC uses hardware semaphores to solve this problem. A semaphore works like a traffic light or a permission token:

  • a hardware unit waits if the data is not ready,
  • it receives permission when the data becomes available,
  • and it can then safely read or write the memory.

The compiler studies the program and inserts these synchronization instructions where necessary. It also removes some of them when it can prove that they are not needed.

The prototype accelerator

To test EAAC, the authors built a prototype containing three main parts:

  1. A GEMM accelerator, which quickly performs matrix multiplication.
  2. A small RISC-V processor, which handles other calculations and control tasks.
  3. Load and store units, which move data between the accelerator and the host computer.

GEMM means general matrix-matrix multiplication. A matrix is a rectangular grid of numbers. Matrix multiplication is used heavily in neural networks, image processing, and scientific calculations.

The GEMM unit contains a systolic array. This is a group of small calculators arranged like a grid. Numbers flow through the grid, allowing many multiplications and additions to happen at the same time—similar to an assembly line where each worker performs one step of a larger task.

The researchers tested a workload based on a small neural-network layer. It performed 24 separate 16×1616 \times 16 matrix multiplications, followed by additional operations such as adding biases and applying ReLU, a common neural-network function.

4. What did they find?

The accelerator was much faster

The version using only the RISC-V processor took:

  • 3,591,489 clock cycles

The version using the EAAC accelerator took:

  • 127,348 clock cycles

This means the accelerator was about 28 times faster for the tested workload.

Implementation Execution time
RISC-V processor only 3,591,489 cycles
EAAC accelerator 127,348 cycles

This is important because it shows that a specialized hardware unit can perform a calculation much more quickly than a small general-purpose processor.

Most of the accelerator's time was spent outside the main calculation

The actual GEMM calculation took only 2,092 cycles, which was about 1.64% of the accelerator's total execution time.

Most of the time was spent by the RISC-V processor doing extra integer calculations, such as preparing data and finishing the workload. This tells the researchers that the GEMM unit itself is very fast, but the surrounding software and data-management tasks still need improvement.

In other words, building a fast engine is not enough if loading, preparing, and organizing the materials takes much longer.

The compiler's semaphore use can become a problem

The researchers also tested larger and more difficult programs. In the worst case, the number of semaphores needed increased roughly in a straight-line relationship with the number of instructions.

That means a program with twice as many instructions might need about twice as many synchronization resources.

This could become a problem because hardware can provide only a limited number of semaphores. In the prototype, there were 16 hardware semaphores.

However, the researchers found that they could reduce semaphore use by examining more details about the hardware. For example, if two operations are already guaranteed to happen in order, the compiler does not need to add another semaphore between them.

The prototype used many hardware resources

The GEMM unit used most of the FPGA resources, especially the digital signal-processing units used for fast arithmetic. The prototype was designed mainly to prove that the EAAC idea works, so it was not yet highly optimized.

The maximum measured clock speed was about 104.6 MHz. The slowest path through the system was inside the RISC-V processor.

5. Why are these results important?

The results show that EAAC can support the complete process from a high-level program to working hardware:

  1. A program is represented using MLIR.
  2. The compiler assigns memory locations.
  3. The compiler plans data movement.
  4. The compiler adds synchronization instructions.
  5. Hardware units execute the resulting instructions.

This is valuable because creating a custom accelerator normally requires engineers to build many separate pieces of software and hardware. That can take a long time and make experimentation difficult.

EAAC provides a shared structure so that different kinds of accelerators—whether written by hand or produced by other design tools—can communicate in a consistent way.

6. What could this mean in the future?

EAAC could make it easier for researchers and engineers to design new accelerators for tasks such as:

  • artificial intelligence,
  • image and video processing,
  • scientific calculations,
  • digital signal processing, and
  • other jobs with predictable data movement.

The system is especially suitable for programs where the compiler can predict what data will be needed and when. It may be less suitable for programs with many unpredictable branches, loops, or irregular memory accesses.

The paper also shows several areas for future improvement:

  • reducing the number of semaphores required,
  • improving the compiler's handling of the RISC-V processor,
  • making data movement faster,
  • reducing hardware resource use, and
  • testing larger and more realistic workloads.

Simple conclusion

The paper introduces EAAC, a framework that helps custom hardware and software work together. Its compiler organizes memory and synchronization, while its hardware provides special-purpose units for fast calculations.

In the experiment, an EAAC system with a matrix-multiplication accelerator was 28 times faster than using only a small RISC-V processor. The work suggests that combining compiler design with hardware design can make accelerator development easier and faster, although the system still needs improvements before it can handle much larger and more complicated programs.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • Limited workload evaluation: The architecture is evaluated primarily on one synthetic dense-layer workload consisting of 24 16×16 int8 GEMM tiles; its effectiveness on representative CNNs, transformers, DSP pipelines, image-processing workloads, and non-machine-learning applications remains untested.
  • No comparison with relevant alternatives: The reported 28× speedup is compared only with the embedded, educational RISC-V baseline. Comparisons against a stronger RISC-V processor, vector extensions, CPU implementations, GPUs, FPGAs, SNAX, CGRAs, or HLS-generated accelerators are absent.
  • Unclear contribution of the accelerator versus the baseline choice: Because the baseline uses a simple single-issue Sodor-derived core, the measured speedup may substantially reflect the weakness of the baseline rather than the general benefit of EAAC. The paper does not quantify this effect using more capable processor configurations.
  • No energy or power evaluation: Although energy efficiency motivates accelerator architectures in the introduction, the paper reports neither dynamic power, static power, energy per operation, nor energy comparisons with the baseline or alternative architectures.
  • No ASIC-area or technology-portability analysis: Resource results are reported for one Kintex-7 FPGA, with no ASIC synthesis, technology scaling analysis, or evaluation on other FPGA families. The area, timing, and power implications of the architecture outside this target are unresolved.
  • Unvalidated claims about extensibility: The prototype contains only a GEMM unit, a load/store unit, and one RISC-V core. The effort required to integrate heterogeneous units such as convolution, pooling, DSP, or irregular control-oriented accelerators is not demonstrated.
  • No quantitative integration-effort study: The paper claims that EAAC lowers hardware/software co-development cost, but does not measure development time, code size, number of required compiler changes, or engineering effort against existing frameworks.
  • Insufficient evidence for frontend interoperability: Although MLIR is presented as enabling multiple frontends, the evaluation does not demonstrate compilation from TensorFlow, PyTorch, ONNX, or other independent MLIR-producing toolchains.
  • Restricted support for dynamic programs: EAAC assumes statically allocated memrefs, predictable lifetimes, and statically analyzable dataflow. The paper does not characterize which dynamic shapes, indirect accesses, data-dependent control flow, loops, recursion, or runtime allocation patterns are unsupported or how they could be handled.
  • Unclear correctness guarantees for synchronization: The paper describes semaphore chaining, generation tags, alias analysis, and buffer reuse, but provides no formal proof, model checking, litmus tests, or systematic validation that the compiler prevents all data hazards and deadlocks.
  • Deadlock and liveness behavior is not analyzed: The effects of semaphore recycling, spin-waiting, chained predicates, and resource exhaustion on deadlock, starvation, or bounded completion are not established.
  • Semaphore scalability remains unresolved: The worst-case number of required semaphores grows linearly with instruction count, and the evaluated workload was deliberately reduced because larger workloads exceeded practical semaphore resources. The maximum scalable program size and hardware requirements are therefore unknown.
  • Semaphore optimization is evaluated only synthetically: The compiler evaluation uses generated worst-case buffer lifetimes and does not report results for a broad suite of real application graphs. It is unclear how often the proposed pruning strategies reduce synchronization overhead in practice.
  • No systematic exploration of hardware parameters: The effects of the number of semaphores, generation-bit width, memory-tier sizes, bus widths, DMA bandwidth, array dimensions, and hardware-unit count are not evaluated.
  • No synchronization-overhead breakdown: The paper does not separately measure semaphore programming, acquisition, chaining, spin-waiting, and synchronization contention costs, making it difficult to identify the dominant scalability bottlenecks.
  • Memory-planning quality is not benchmarked: The modified linear-scan spilling strategy is described, but there is no comparison with optimal allocation, ILP/CP-based allocation, alternative heuristics, or other scratchpad-management techniques in terms of transfer volume and execution time.
  • DMA and memory-hierarchy performance is underexplored: The evaluation does not report DMA latency, bandwidth utilization, tier-by-tier traffic, overlap efficiency, or sensitivity to memory-access patterns. Consequently, the benefit of the multi-tier hierarchy is not isolated.
  • Compiler compilation time is not reported: EAAC is motivated partly by rapid iteration, but compilation latency, memory consumption, scalability with program size, and the cost of semaphore allocation are not quantified.
  • The RISC-V code-generation path is a known bottleneck: Instruction-memory transfer and LLVM lowering account for significant overhead, yet no optimization or alternative code-generation path is evaluated. It remains unclear whether the reported performance would persist after realistic compiler/runtime improvements.
  • The role of host initialization is ambiguous: Both measurements include binary-transfer time, but the paper does not provide results excluding initialization or distinguish one-time setup costs from steady-state execution, which limits interpretation for repeated inference or streaming workloads.
  • No repeated-execution or throughput study: The evaluation appears to measure a single invocation. It does not assess amortized performance across batches, continuous streams, repeated kernel launches, or workloads where initialization and configuration can be reused.
  • Limited treatment of contention and concurrency: The prototype contains a single GEMM unit, so it cannot demonstrate contention among multiple instances of the same accelerator or among several concurrent hardware units. The claimed partial out-of-order execution and distributed scheduling model therefore remain only partially validated.
  • No robustness analysis under resource exhaustion: The compiler’s behavior when semaphore capacity, scratchpad capacity, instruction memory, data memory, or DMA bandwidth is insufficient is not fully specified. Feasibility diagnostics, fallback strategies, and compile-time failure modes remain open questions.
  • Hardware interface assumptions are not standardized: The paper states that custom units need only a memory interface and an instruction-receive method, but does not provide a formal interface specification, timing protocol, verification methodology, or compatibility tests for independently developed units.
  • Security and isolation are unexplored: Shared scratchpad memory, memory-mapped semaphores, DMA, and embedded processors are not analyzed for protection, privilege separation, faulty accelerator behavior, or isolation between mutually untrusted workloads.
  • Fault handling is absent: The architecture does not discuss recovery from malformed instructions, DMA errors, hardware-unit faults, semaphore corruption, or RISC-V exceptions.
  • The claimed flexibility–efficiency position is not empirically established: EAAC is argued to occupy a middle ground between fixed HLS accelerators and general-purpose CGRAs, but no quantitative comparison of programmability, performance, area, energy, or supported workload diversity is provided.
  • Data-type and numerical-format support is unclear: The prototype uses int8 inputs and int32 accumulation, but the compiler and hardware framework’s support for floating point, mixed precision, sparsity, quantization variants, and wider or irregular element types is not demonstrated.
  • No end-to-end application-quality validation: The synthetic workload does not establish accuracy, numerical correctness, or model-level behavior for a complete deployed application, particularly after bias addition, ReLU, and requantization.
  • Reproducibility is incomplete: Although repositories are linked, the paper does not provide sufficient details about compiler versions, synthesis constraints, generated binaries, exact benchmark inputs, simulation versus FPGA execution methodology, or scripts needed to reproduce all reported results.
  • The effect of hardware frequency limitations is unresolved: The prototype’s critical path lies in the RISC-V core, but the paper does not evaluate how replacing or optimizing that core changes system-level performance, resource use, or the balance between computation and orchestration overhead.

Practical Applications

The paper’s main practical contribution is a reusable hardware–compiler integration framework for statically analyzable, data-intensive workloads. Its demonstrated 28× speedup is promising, but it was measured on a synthetic dense-layer workload and a prototype FPGA implementation; therefore, applications should be understood as opportunities rather than validated production results.

Immediate Applications

  • Rapid prototyping of FPGA accelerators for machine-learning inference
    • Sector: AI hardware, embedded systems, edge computing.
    • Teams can use EAAC’s MLIR-based pipeline to integrate custom units for GEMM, convolution-like kernels, ReLU, pooling, quantization, and other tensor operations without developing a complete memory-management and synchronization stack from scratch.
    • A practical workflow would be:
    • 1. Export a model or kernel into MLIR.
    • 2. Bufferize tensor operations into statically allocated memrefs.
    • 3. Map supported operations to EAAC hardware units.
    • 4. Compile DMA transfers, scratchpad placement, semaphore dependencies, and embedded RISC-V control code.
    • 5. Package the resulting FlatBuffer and load it through the runtime.
    • Feasibility dependencies: The workload must have predictable memory access and mostly static data dependencies. Dynamic control flow, irregular sparse operations, and data-dependent memory access may require additional compiler and hardware support.
  • Edge-AI devices for low-latency classification and anomaly detection
    • Sector: Industrial monitoring, predictive maintenance, IoT, robotics, automotive.
    • The prototype’s combination of an int8 systolic GEMM unit and an embedded RISC-V core could support compact neural-network inference close to sensors, reducing transfers to a CPU, GPU, or cloud service.
    • Potential products include FPGA-based sensor hubs, vibration-anomaly detectors, small vision systems, and embedded controllers for factory equipment.
    • The RISC-V core can execute control, activation, and requantization operations while the accelerator performs matrix multiplication.
    • Feasibility dependencies: Production deployment would require validation across representative models, power measurements, real sensor workloads, fault handling, and comparison against commercial microcontrollers, NPUs, and FPGA IP.
  • Reusable accelerator IP blocks for FPGA-based system-on-chip designs
    • Sector: Semiconductor design, embedded computing, industrial electronics.
    • EAAC’s memory-mapped data, instruction, and semaphore interfaces allow independently developed RTL, Chisel, or HLS-generated units to coexist in one accelerator fabric.
    • Hardware teams could create a library of compatible units—for example, GEMM, FFT, image filters, compression, encryption, or signal-processing blocks—and connect them through the shared scratchpad and DMA hierarchy.
    • This could reduce integration work when replacing one accelerator implementation with another.
    • Feasibility dependencies: Each unit must implement the EAAC instruction and memory-interface conventions, and the compiler must contain suitable code-generation support. Area, timing, and bus-contention limits remain design-specific.
  • Compiler-assisted DMA and scratchpad-memory management
    • Sector: Embedded software, FPGA programming, high-performance computing.
    • EAAC’s static allocation and lifetime-based spilling can be used immediately for applications whose working sets exceed the fastest on-chip memory.
    • The compiler can determine when buffers should move between Tier 0, Tier 1, and Tier 2 memory, then emit explicit DMA operations. This provides a repeatable alternative to manually inserting transfers and buffer reuse logic.
    • Likely early adopters include developers of image pipelines, DSP kernels, and fixed-size neural-network inference engines.
    • Feasibility dependencies: The memory sizes, transfer bandwidths, and buffer lifetimes must be accurately represented. Poor allocation or excessive spilling can eliminate the benefit of acceleration.
  • Education and research infrastructure for hardware–software co-design
    • Sector: Academia and engineering training.
    • The open-source compiler, accelerator, and minimal RISC-V core can serve as a laboratory platform for courses and research on MLIR, FPGA architecture, DMA scheduling, dataflow execution, hardware semaphores, and RISC-V integration.
    • Students can add a functional unit and study the effect of memory hierarchy, synchronization, bus width, or systolic-array dimensions without implementing an entire toolchain.
    • Feasibility dependencies: The existing prototype is FPGA-oriented and includes educationally simple components. Additional documentation, tests, simulators, and reproducible build flows would improve classroom usability.
  • Benchmarking and experimentation with synchronization strategies
    • Sector: Computer architecture and compiler research.
    • EAAC provides a concrete testbed for evaluating hardware semaphores, dependency pruning, semaphore reuse, generation tags, and data-driven scheduling.
    • Researchers can compare compiler policies by measuring semaphore count, execution time, memory traffic, FPGA resource use, and synchronization stalls.
    • This is especially useful because the paper identifies semaphore allocation as a current scalability bottleneck.
    • Feasibility dependencies: Results depend strongly on the target hardware’s number of semaphore entries, execution overlap, memory bandwidth, and instruction-granularity assumptions.
  • Specialized DSP and image-processing pipelines
    • Sector: Telecommunications, medical imaging, industrial vision, audio processing.
    • Regular operations such as FIR filtering, transforms, matrix operations, pixel-wise functions, and fixed pipelines can be implemented as EAAC functional units connected by statically scheduled buffers.
    • The semaphore and DMA system can support producer–consumer streaming between stages, including buffer broadcasting to multiple consumers.
    • Feasibility dependencies: The current evaluation focuses on GEMM rather than DSP or image-processing kernels. New functional units, numerical formats, and domain-specific code-generation passes would be required.

Long-Term Applications

  • A portable accelerator ecosystem spanning multiple hardware implementations
    • Sector: Semiconductor platforms, cloud acceleration, embedded systems.
    • With a sufficiently standardized EAAC interface, the same MLIR-level program could target different FPGA or ASIC configurations containing different combinations of functional units.
    • Possible outcomes include accelerator modules for linear algebra, cryptography, signal processing, and machine learning that share compiler and runtime infrastructure while differing in implementation technology.
    • This would resemble an accelerator-oriented software ecosystem in which hardware vendors expose compatible memory, instruction, and synchronization abstractions.
    • Dependencies and risks: Hardware units currently require application-specific code generation. Portability would require a stable ISA or operation specification, capability discovery, versioning, formal interface definitions, and validation across devices.
  • ASIC accelerators for energy-efficient edge inference
    • Sector: Healthcare devices, robotics, autonomous systems, industrial IoT.
    • EAAC’s scratchpad-centric architecture, explicit data reuse, and avoidance of general-purpose instruction execution for large kernels could be adapted into ASICs optimized for low power and predictable latency.
    • An ASIC implementation could combine fixed-function systolic arrays with small programmable RISC-V controllers, potentially supporting multiple related model families rather than only one neural network.
    • Dependencies and risks: The paper validates only an FPGA prototype. ASIC feasibility requires physical-design studies, power and thermal analysis, manufacturing cost justification, security review, and evidence that the added flexibility does not compromise efficiency.
  • Scalable multi-accelerator dataflow systems
    • Sector: Robotics, autonomous vehicles, high-performance embedded computing.
    • Multiple accelerator units could execute partially out of order, communicating through shared buffers and chained hardware semaphores. A larger system might support pipelines such as sensor preprocessing → feature extraction → matrix computation → decision logic.
    • This could provide predictable latency while allowing independent stages to overlap.
    • Dependencies and risks: The current prototype has limited semaphore resources, and worst-case semaphore requirements grow linearly with instruction count. Larger systems will need better dependency analysis, hierarchical synchronization, semaphore virtualization, or local synchronization within functional units.
  • Automatic hardware-aware scheduling and cost modeling
    • Sector: Compiler technology and electronic design automation.
    • Future EAAC versions could incorporate detailed models of memory bandwidth, DMA latency, functional-unit occupancy, semaphore pressure, and energy consumption.
    • The compiler could then select buffer placements, operation ordering, semaphore elimination, tile sizes, and hardware-unit assignments based on an explicit performance or energy objective.
    • Such a tool could support design-space exploration before committing to an FPGA or ASIC implementation.
    • Dependencies and risks: Accurate cost models require hardware measurements or simulation. Optimization must avoid trading away correctness when operations overlap or reuse memory addresses.
  • Automatic integration of neural-network frontends
    • Sector: AI software infrastructure.
    • Since MLIR can receive representations from frameworks such as TensorFlow and PyTorch, a mature EAAC toolchain could compile quantized model subgraphs directly into EAAC programs.
    • A possible product would be an edge-inference compiler that partitions a model between EAAC units and the embedded RISC-V core, inserts DMA transfers, and reports unsupported operations.
    • Dependencies and risks: The present work demonstrates a synthetic fully connected layer, not complete models. Real deployment requires support for convolutions, attention, normalization, dynamic shapes, model partitioning, numerical calibration, and robust fallback behavior.
  • Support for real-time robotics and control pipelines
    • Sector: Robotics, drones, autonomous vehicles, industrial automation.
    • Deterministic data-driven scheduling and explicit synchronization could be useful for fixed-rate pipelines involving image or lidar preprocessing, sensor fusion, control-law evaluation, and actuator command preparation.
    • Hardware units could be specialized for perception or signal processing while the RISC-V core handles supervisory control.
    • Dependencies and risks: Real-time use requires bounded worst-case execution time, interrupt handling, fault recovery, priority management, and guarantees against deadlock or semaphore starvation. These properties are not established by the current prototype.
  • Acceleration of privacy-sensitive or safety-critical local computation
    • Sector: Healthcare, finance, defense, and public infrastructure.
    • Local FPGA or ASIC accelerators could process medical signals, biometric data, fraud-detection features, or industrial telemetry without sending raw data to cloud systems.
    • Static memory allocation and constrained data movement may also simplify certain forms of security auditing and information-flow analysis.
    • Dependencies and risks: The paper does not evaluate confidentiality, side-channel resistance, secure boot, isolation, or formal verification. These would be required before deployment in regulated or safety-critical environments.
  • Compiler and architecture standards for open hardware research
    • Sector: Academia, open-source EDA, public research policy.
    • EAAC could contribute to a broader open ecosystem in which researchers publish accelerator units together with MLIR dialects, interface descriptions, runtime packages, and reproducible FPGA configurations.
    • Funding agencies and universities could use such platforms to make accelerator results more comparable and easier to reproduce than designs tied to proprietary toolchains.
    • Dependencies and risks: This requires sustained maintenance, standardized benchmarks, compatibility testing, licensing clarity, and independent replication. The current evaluation is limited in workload size and hardware diversity, so broader claims about general-purpose accelerator performance would be premature.

Glossary

  • Accelerator architecture: A hardware design specialized to speed up particular computational workloads. “Heterogeneous accelerator architectures offer an efficient path to performance for compute-intensive workloads.”
  • Address aliasing: A condition in which different names or references resolve to the same memory address. “If the compiler finds an instance of address aliasing within a sliding window, it locates the semaphores guarding the given buffers and chains them in order of use.”
  • Atomic operation: An indivisible operation that completes without interference from concurrent operations. “To accelerate atomic access to the semaphore pairs, EAAC utilizes hardware semaphores.”
  • AXI4-Stream: A widely used streaming interconnect protocol for transferring data between hardware components. “A load and store unit that provides fully bidirectional data transfers to and from the host using AXI4-Stream interfaces.”
  • Bufferization: A compiler transformation that converts abstract tensor values into explicit memory buffers. “To convert tensors into a memory representation that is closer to hardware, bufferization on tensors is used.”
  • Calyx: A hardware compiler infrastructure and accelerator design framework. “Existing frameworks address this primarily by raising the abstraction level of accelerator implementation through high-level synthesis (HLS), accelerator design languages, or hardware compiler infrastructures such as Allo, Calyx, and CIRCT.”
  • Chisel: A hardware-construction language embedded in Scala that supports parameterized hardware generators. “This allows accelerators produced through different methods, including hand-written RTL, Chisel generators, and HLS-generated designs, to coexist within a shared framework.”
  • Coarse-grained reconfigurable array (CGRA): A reconfigurable architecture composed of relatively large programmable processing elements. “Coarse-Grained Reconfigurable Arrays (CGRAs) occupy a different point in the design space.”
  • Compiler backend: The compiler component that translates an intermediate representation into target-specific code or instructions. “Domain-specific hardware acceleration has historically complicated compiler backend development.”
  • Compiler–hardware co-design: Joint development of compiler software and hardware so that each is optimized for the capabilities of the other. “We propose EAAC (Extensible Accelerator ArChitecture), a framework for hardware-compiler co-development focused on reducing the cost of hardware/software integration.”
  • Control and status register (CSR): A processor register used to configure hardware or report its state. “SNAX utilizes RISC-V processor cores for synchronization, configuration, and control of the hardware accelerator units, specifically through control and status registers in the RISC-V core.”
  • Data dependency: A relationship in which one operation requires data produced or modified by another operation. “EAAC is specifically designed for workloads that are optimal for acceleration, often characterized by regular and predictable memory access patterns, minimal control flow, and statically analyzable data dependencies.”
  • Data-flow graph: A graph representing operations as nodes and data dependencies as edges. “To ensure correctness, the EAAC compiler will allocate a virtual semaphore for every edge between a producer and consumer node in the dataflow graph.”
  • Data-driven scheduling: Scheduling in which an operation executes when its required input data becomes available. “EAAC operates exclusively using data-driven scheduling; unlike a typical von Neumann architecture, it does not have a program counter, branching, or looping.”
  • Direct memory access (DMA): Hardware-supported transfer of data between memory regions without requiring continuous processor involvement. “The EAAC compiler is capable of static memory allocation and manages a multi-tiered memory hierarchy using explicit direct memory access (DMA) control between layers.”
  • Decoupled access-execute: An architecture that separates data movement or access operations from computation operations so they can proceed independently. “A flexible hardware architecture based on decoupled access-execute, with phased execution allowing for data transfers to overlap with execution.”
  • Dialect: In MLIR, an extensible collection of operations, types, and attributes for a particular abstraction or domain. “Developers can define custom IR operations and transformation passes to support compiling for targets other than classical assembly languages.”
  • Digital signal processing: Computational manipulation of sampled signals such as audio, video, or sensor data. “Examples include GEMM, digital signal processing, and image processing.”
  • Embedded processor: A processor integrated into a larger hardware system to perform local control or computation. “We demonstrate this design through a prototype GEMM accelerator combining a systolic array unit with an embedded RISC-V core.”
  • Fan-out: The number of destination components receiving a signal or data value from one source. “This mechanism also extends to support broadcasting behavior for instruction outputs with a fan-out of more than one.”
  • FlatBuffer: A serialized binary data format designed for efficient access without requiring unpacking or parsing the entire object. “The final output from the compiler is in the form of a FlatBuffer, which is imported into an assembler and runtime that configures and programs the accelerator.”
  • Formal description: A precise, machine-processable specification of a system’s behavior or interface. “Act dissolves tensor operations into a set of IR-like primitive tensor operations; ISA instructions are then defined in terms of these primitives.”
  • General matrix multiplication (GEMM): Matrix multiplication, usually expressed as C=ABC = AB or C=αAB+βCC = \alpha AB + \beta C, and a core workload in machine learning. “We validate this approach with a prototype GEMM accelerator combining a systolic array unit with an embedded RISC-V core.”
  • Generation tag: Metadata distinguishing different uses or allocations of the same physical synchronization or storage resource. “To prevent aliasing, the compiler appends a semaphore generation tag to each new allocation.”
  • High-level synthesis (HLS): The automated generation of hardware descriptions from high-level algorithmic code. “Existing frameworks address this primarily by raising the abstraction level of accelerator implementation through high-level synthesis (HLS), accelerator design languages, or hardware compiler infrastructures such as Allo, Calyx, and CIRCT.”
  • Intermediate representation (IR): A structured compiler representation between source code and target machine code. “MLIR is intended to function as an intermediate representation supported by multiple frontends.”
  • Instruction set architecture (ISA): The programmer-visible specification of a processor’s instructions, registers, data types, and behavior. “Act dissolves tensor operations into a set of IR-like primitive tensor operations; ISA instructions are then defined in terms of these primitives.”
  • Linear scan: A compiler algorithm commonly used for register allocation and interval-based resource assignment. “Data movement scheduling is integrated into the memory allocation algorithm using a modified approach inspired by linear-scan.”
  • Lowering: A compiler transformation that translates a higher-level representation into a lower-level representation. “In this pipeline, operations are first lowered into the LLVM dialect, which is later compiled into RISV-V assembly.”
  • Memory-mapped: Organized so that hardware devices or control registers are accessed through ordinary memory addresses. “The architecture additionally allows for a degree of out-of-order processing between different hardware units.”
  • MemRef: An MLIR type representing a reference to a memory buffer together with metadata such as shape and memory space. “Memrefs are MLIR's core dialect for memory references, and act as pointers to buffers in memory with additional information about element types, shapes, and memory space.”
  • Memory hierarchy: A system of memory levels with differing capacities, latencies, and access costs. “EAAC organizes its memories into a configurably tiered hierarchy interleaved with programmable DMA units.”
  • MLIR: A compiler infrastructure and extensible intermediate-representation framework for heterogeneous and domain-specific systems. “MLIR is a compiler framework designed with heterogeneous and domain-specific computing in mind.”
  • Out-of-order execution: Execution in which operations may run in an order different from their original program order when their dependencies permit it. “The architecture additionally allows for a degree of out-of-order processing between different hardware units.”
  • Predicate: A Boolean condition used to determine whether an action should occur. “These events are logged and can be used to create predicates, such as "semaphore x generated event y".”
  • Read-after-write (RAW) hazard: A dependency hazard in which a read must not occur before an earlier write to the same location completes. “In the subsequent pass, the compiler analyzes the address space of the accelerator (SPM Tier 0) to identify cases where read-after-write (RAW) errors due to address-space reuse are likely to occur.”
  • ReLU: The rectified linear unit activation function, defined as max⁡(0,x)\max(0,x). “24 independent 16×1616\times16 int8 matmul tiles are summed in int32, then bias-added, ReLU'd, and requantized back to int8.”
  • Requantization: Conversion of numerical values from one quantized precision or scale to another. “24 independent 16×1616\times16 int8 matmul tiles are summed in int32, then bias-added, ReLU'd, and requantized back to int8.”
  • Register renaming: A processor technique that maps architectural registers to distinct physical registers to eliminate false dependencies. “The use of generation tags can be compared to register renaming in Tomasulo's algorithm, where multiple versions of the same architectural register can coexist and are distinguished by different tags.”
  • RISC-V: An open instruction set architecture designed around a modular and extensible specification. “We validate this approach with a prototype GEMM accelerator combining a systolic array unit with an embedded RISC-V core.”
  • Scratchpad memory (SPM): Programmer- or compiler-managed on-chip memory that provides predictable, low-latency access. “Because individual instructions represent large units of work, we chose to construct the architecture as a memory-to-memory system---with the Tier 0 SPM being the working memory for the hardware units.”
  • Semaphore: A synchronization mechanism based on a counter or state that regulates access to shared resources. “The synchronization primitives chosen are counting or binary semaphore pairs.”
  • Systolic array: A regular grid of processing elements that passes data rhythmically between neighboring elements for highly parallel computation. “We validate this approach with a prototype GEMM accelerator combining a systolic array unit with an embedded RISC-V core.”
  • Tensor: A multidimensional numerical data structure commonly used to represent machine-learning data. “EAAC and Act both treat tensors as first-class datatypes.”
  • TileLink: A hardware interconnect protocol used to connect processor cores and memory-mapped devices. “Sodor has been modified with a TileLink data-memory interface that allows it to interact with the hardware semaphores and the main Tier 0 scratchpad memory.”
  • Triggered instruction: An instruction whose execution is enabled by satisfaction of a specified event or condition. “Semaphore initializations are issued when a given set of predicates is satisfied, in a fashion inspired by triggered instructions.”
  • Virtual semaphore: A compiler-level synchronization object that has not yet been assigned to a physical hardware semaphore. “In practice, the compiler manages the state of buffers using virtual semaphores, which are semaphores not assigned to physical hardware semaphores yet.”
  • Von Neumann architecture: A computer architecture in which instructions and data are stored in memory and execution is generally controlled by a program counter. “EAAC operates exclusively using data-driven scheduling; unlike a typical von Neumann architecture, it does not have a program counter, branching, or looping.”

Open Problems

We're still in the process of identifying open problems mentioned in this paper. Please check back in a few minutes.

Tweets

Sign up for free to view the 1 tweet with 167 likes about this paper.