Papers
Topics
Authors
Recent
Search
2000 character limit reached

RV-IM100: Quantifying ISA Extension, Datapath Width, and Pipeline Depth Trade-offs in RISC-V Microarchitectures

Published 3 May 2026 in cs.AR | (2605.01902v2)

Abstract: While functional RISC-V implementations are readily available in academia, controlled empirical studies that extend a single baseline architecture along multiple design axes and quantify the resulting trade-offs at each step remain scarce. This paper presents RV-IM100, a family of 10 incremental FPGA-implemented microarchitectures derived from a common 5-stage pipeline baseline, systematically varying datapath width from RV32 to RV64, instruction set from I to IM, and pipeline depth from 5 to 8 stages under controlled conditions. The I-to-IM extension produced strongly benchmark-dependent effects at the 5-stage level: CoreMark throughput more than doubled while Dhrystone throughput decreased marginally despite improved per-MHz efficiency. Within the RV32IM configuration, an iterative timing-closure methodology combined with pipeline deepening from 5 to 8 stages raised max frequency from 43 to 126MHz, increasing both Dhrystone and CoreMark throughput by 71%, while per-MHz efficiency decreased by 41%. The 6-to-7-stage transition caused throughput regression in RV64 despite higher frequency, revealing that the outcome depends on available frequency headroom. Cross-width comparison showed RV32 outperforming RV64 in absolute throughput, with per-MHz efficiency diverging by benchmark: RV64 led by 2.3% in DMIPS/MHz while RV32 led by 4.6% in CoreMark/MHz. At 8 stages, RV32 required 59% fewer LUTs, 51% fewer FFs, and 80% fewer DSPs, indicating that the resource cost of width extension substantially exceeds the modest efficiency differences. These results provide a quantitative reference for design-space exploration in RISC-V microarchitectures. All RTL sources and benchmark configurations are publicly available.

Authors (1)

Summary

  • The paper presents a controlled, empirical study that isolates the impact of ISA extensions, datapath width, and pipeline depth on RISC-V performance.
  • Key findings include significant frequency improvements with deep pipelines, non-monotonic throughput trends, and notable resource and power overheads in wider datapaths.
  • Methodological innovations such as parametric evolution and iterative RTL optimization offer reproducible quantitative insights for future microarchitectural research.

Authoritative Technical Summary of "RV-IM100: Quantifying ISA Extension, Datapath Width, and Pipeline Depth Trade-offs in RISC-V Microarchitectures" (2605.01902)

Systematic Microarchitectural Variation in RISC-V: Motivation and Approach

RV-IM100 addresses a core deficiency in the literature: absence of controlled, empirical studies where a single baseline microarchitecture is incrementally modified and measured across orthogonal design axes—datapath width (RV32 vs. RV64), instruction set extension (I vs. IM), and pipeline depth (5–8 stages). By constructing ten variants derived from a common 5-stage RV32I core, the work enables quantitative isolation of architectural effects that are otherwise confounded in cross-core or ecosystem-level comparisons.

The design methodology adheres to strict parametric evolution, preserving baseline logic unless modifications are structurally necessary. RV32 and RV64 variants utilize identical control logic, forwarding, and hazard paths; the datapath is widened for RV64, with necessary ISA-driven changes—W-suffix opcodes, 6-bit shamt, dual-width ALU and memory enable logic. For IM extension, multi-cycle multiplier and divider units are added, structured to exploit FPGA DSP primitives, with pipeline stalls on multi-cycle operations. Pipeline deepening follows a stepwise insertion of architectural stages targeting critical path segmentation; notably, the EXR stage in 8SP decomposes hazard/forwarding logic from ALU computation.

Figure 1

Figure 1: Simplified block diagram of the 46F5SP base microarchitecture.

Impact of Structural, Width, and ISA Extension: Frequency, Throughput, and Efficiency

Frequency Scaling

Pipeline deepening produces substantial frequency improvements across both RV32 and RV64. Frequency rises from 43 MHz (RV32IM 5SP) and 38 MHz (RV64IM 5SP) to 126 MHz (RV32IM 8SP) and 101 MHz (RV64IM 8SP), representing 192% and 165% increases, respectively. The EX-BR split and memory migration to synchronous BRAM in 6SP and 7SP configurations contribute to timing improvement, although frequency gains are contingent on successful segmentation of the critical path, notably the forwarding-ALU chain.

Figure 2

Figure 2: Signal-level block diagram of 72F8SP, the final RV64IM 8-stage pipeline microarchitecture (IF, IO, ID, EXR, EX, BR, MEM, WB).

Figure 3

Figure 3: SoC maximum operating frequency across pipeline variants.

Benchmark Throughput

Absolute throughput (performance at realized frequency) and per-MHz efficiency (IPC) reveal non-monotonic trends. The addition of M extension is strongly benchmark-dependent: CoreMark throughput more than doubles, while Dhrystone declines marginally due to frequency penalty and negligible multiplication/division in its workload.

For deep pipeline variants, throughput regresses in RV64 between 6SP and 7SP despite higher frequency, attributed to increased branch flush and two-cycle front-end refill latency; absolute throughput increases only when frequency improvement is sufficient to compensate for IPC loss. At 8SP, combined structural and RTL-level optimizations yield the steepest throughput gains.

Figure 4

Figure 4

Figure 4: Benchmark absolute throughput across pipeline variants.

Relative Efficiency (IPC per MHz)

Per-MHz efficiency declines monotonically with pipeline deepening due to increased misprediction penalty and load-use hazards. The 8SP configuration achieves lower per-MHz metrics but the frequency and absolute throughput remain superior to shallower configurations. Cross-width analysis shows RV64 outperforming RV32 in DMIPS/MHz by 2.3% (at 8SP), whereas RV32 leads in CoreMark/MHz by 4.6%. This divergence suggests non-trivial IPC effects inherent to wider datapaths, possibly due to compiler-generated instruction mix and control overhead, warranting instruction-level profiling in future work.

Figure 5

Figure 5

Figure 5: Benchmark relative efficiency across pipeline variants.

Resource Utilization and Power

Resource utilization exhibits super-linear scaling with datapath width. At 8SP, RV32 requires 59% fewer LUTs, 51% fewer FFs, and 80% fewer DSPs than RV64, underscoring the prominent resource cost of width extension relative to modest efficiency gains. Pipeline stage insertion generally converts combinational LUT logic to sequential FFs, reducing LUT count and shifting resource profile.

Power estimation mirrors resource trends: RV64 variants consume 17% more power at 8SP. Populating pipeline stages with FFs, as opposed to LUT-based combinational critical paths, mitigates power increases despite deeper pipelines.

Figure 6

Figure 6

Figure 6: Core-only resource utilization across pipeline variants.

Figure 7

Figure 7: Core-only estimated total on-chip power across pipeline variants.

Timing Closure and RTL Optimization

The timing closure methodology is central to achieving frequency targets. Iterative RTL optimization eliminates critical path bottlenecks, especially forwarding control fan-out. The insertion of the EXR stage in 8SP is necessitated by irreducible hazard/forwarding/ALU chain, achieving workload-balance across pipeline stages and enabling a 100 MHz RV64IM core.

Limitations, Implications, and Future Directions

The study omits cache hierarchies and uses a simple 2-bit branch predictor. The structural and ISA variations are thus isolated from memory interaction complexity and advanced branch speculation impacts. Empirical results establish that pipeline deepening can yield higher throughput only if frequency headroom absorbs increased IPC penalties; widened datapaths have disproportionate resource and power costs versus relatively small efficiency gains.

Practically, for resource-constrained FPGAs, RV32 variants offer better cost-performance trade-offs unless 64-bit address/data is mandatory. Theoretically, the study validates that architectural evolution must jointly consider interactions in pipeline structure, field width, and optional extensions rather than single-axis extrapolation.

Future work should incorporate configurable cache hierarchies, stronger branch predictors, and instruction-level profiling to model hazard/refill effects precisely. Open-sourced RTL and benchmarks provide a reproducible foundation for subsequent microarchitectural exploration.

Figure 8

Figure 8: Signal-level block diagram of the RV-IM100 SoC.

Conclusion

RV-IM100 provides the first controlled, empirical quantification of trade-offs across ISA, width, and pipeline depth in a single RISC-V processor lineage. The results demonstrate that throughput, efficiency, resource, and power outcomes are shaped by complex interactions between structural, datapath, and ISA extensions. Deep pipelines are only justified where frequency gains absorb increased control penalties; bit-width extension incurs disproportionate resource and power overhead. The released RTL and benchmarks serve as quantitative references and enable reproducible evaluation for future architectural research and practical deployment.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.