---
title: BASIC_RV32s RISC-V Microarchitecture Evolution
url: https://www.emergentmind.com/topics/basic_rv32s
type: topic
---

# BASIC_RV32s RISC-V Microarchitecture Evolution

Searching arXiv for the primary BASIC_RV32s paper and closely related RV32/RISC-V educational microarchitecture references.
BASIC_RV32s is an open-source educational framework and microarchitectural roadmap for implementing a RISC-V RV32I processor. It is defined not as a single reference core, but as an explicitly documented evolution of several cores, beginning with a textbook-style single-cycle implementation and culminating in a relatively complete 5-stage pipelined core with full hazard forwarding, dynamic branch prediction, and exception handling, verified on FPGA within a System-on-Chip (SoC) that supports UART-based debugging and reports 1.09 DMIPS/MHz at 50 MHz. By releasing Register-Transfer Level (RTL) source code, signal-level logic block diagrams, and development logs under the MIT license, the project is positioned as a reproducible bridge between architectural theory and synthesizable hardware practice [2510.15887].

## 1. Concept, goals, and architectural scope

BASIC_RV32s is organized around a step-by-step, Patterson & Hennessy–style path from theory to working hardware. Its stated goals are to expose the full progression from simple to more capable cores, to document not only the final design but also intermediate versions together with engineering decisions and debugging history, and to make the framework reproducible and modifiable by releasing RTL, diagrams, and development logs under the MIT license [2510.15887].

The architectural target is the RISC-V RV32I base integer ISA, augmented over the course of the roadmap by the Zicsr extension and a minimal trap-oriented subset consisting of `ECALL`, `EBREAK`, and `MRET`. The supported instruction classes include integer arithmetic and logic, loads and stores, conditional branches, jumps, `LUI`, `AUIPC`, CSR access instructions such as `CSRRW`, `CSRRS`, `CSRRC` and their immediate forms, and machine-mode trap-related instructions in later variants. The privileged scope is explicitly machine-mode oriented, and the trap controller and exception detector are structured to be compatible with the machine-mode trap model, while full ISA-compliant exception handling is reserved for future work [2510.15887].

| Core variant | ISA / scope | Main addition |
|---|---|---|
| RV32I37F | 37 instructions from base RV32I | Baseline single-cycle datapath and control |
| RV32I43F | RV32I37F + Zicsr | `CSRFile` and CSR control paths |
| RV32I46F | RV32I43F + `ECALL`, `EBREAK`, `MRET` | Trap Controller and Exception Detector |
| RV32I46F_5SP | Same ISA coverage as RV32I46F | 5-stage pipeline, hazard logic, branch prediction |

A common misconception is to treat BASIC_RV32s as only the final pipelined core. The project description rejects that interpretation explicitly: it is not just a single reference core, but a documented microarchitectural evolution. This gives it a dual identity as both an instructional reference and a realistic, synthesizable RISC-V microarchitecture [2510.15887].

## 2. Evolution from single-cycle execution to pipelined execution

The roadmap proceeds through four concrete designs: a single-cycle RV32I37F, a CSR-capable RV32I43F, a trap-capable RV32I46F, and the final 5-stage pipelined RV32I46F_5SP. In the initial single-cycle form, the Program Counter, Instruction Memory, Register File, ALU, and Data Memory are each used once per instruction, and fetch, decode, execute, memory access, and write-back all complete within one clock cycle. The control path is centralized and decodes opcode and function fields into ALU operation, register write enable, memory control, and PC selection signals, closely matching the canonical single-cycle organization described in Patterson and Hennessy [2510.15887].

The transition to RV32I43F introduces a dedicated `CSRFile` together with control logic for CSR addressing, operand selection, and updates consistent with Zicsr semantics. RV32I46F then adds a Trap Controller and Exception Detector, enabling handling of `ECALL`, `EBREAK`, simple synchronous exceptions, trap entry via machine CSRs such as `mepc` and `mcause`, and trap return through `MRET`. These additions remain embedded in a single-cycle microarchitecture, but they establish the state and control machinery later reused in the pipelined core [2510.15887].

The final RV32I46F_5SP reorganizes the design into the classic 5-stage pipeline: IF, ID, EX, MEM, and WB. The paper’s signal-level block diagram identifies a PC Controller, a 2-bit dynamic Branch Predictor in IF, a Hazard Unit, a Forward Unit, pipeline-integrated CSR logic, an Exception Detector and Trap Controller, and a Debug Interface exporting signals such as `dbg_pc`, `dbg_instruction`, `dbg_ALU_result`, `mcycle`, and `minstret`. This organization places BASIC_RV32s within the same broad microarchitectural lineage as other educational and FPGA-oriented five-stage RV32 cores, including RVCoreP, which likewise centers optimization around branch prediction, hazard handling, and FPGA timing paths [2002.03568].

This staged progression is significant because it preserves continuity between abstract textbook datapaths and concrete implementation artifacts. Each added feature is presented not as an isolated enhancement but as a structural transformation of the preceding core, allowing datapath, control, and debugging considerations to remain visible across versions [2510.15887].

## 3. Hazard handling, branch prediction, and control-flow behavior

The pipelined RV32I46F_5SP includes explicit Hazard Unit and Forwarding Unit logic. The primary data hazard class is RAW, while WAR and WAW are stated not to occur in this simple in-order pipeline. Data hazards are resolved through forwarding from the execution, memory access, and write-back stages, and a dependent load-use case is handled by stalling the pipeline for one cycle, holding the PC and IF/ID state while injecting a bubble into ID/EX. This is presented as the standard Patterson & Hennessy solution for a five-stage in-order machine [2510.15887].

Control hazards are managed through a 2-bit dynamic branch predictor located in the IF stage, branch resolution in EX, and pipeline flush on misprediction. On each branch, the predictor supplies a taken or not-taken decision; EX later evaluates the actual condition and target. If the prediction is wrong, incorrect-path instructions are flushed and the PC is redirected, while the predictor’s saturating counter state is updated. Structural hazards are not emphasized; the design description suggests they are avoided by design through separate or logically partitioned instruction and data memories, and implementation-level memory mapping conflicts were resolved by bypass paths between instruction and data memory [2510.15887].

The performance objective of the pipelined core is near-1 CPI on straight-line code, but the measured Dhrystone average is 1.61 cycles per instruction, reflecting remaining stalls and branch mispredictions. The same section associates the reported 1.09 DMIPS/MHz with the combined impact of pipelining, forwarding, and dynamic prediction [2510.15887].

In broader context, BASIC_RV32s occupies an intermediate point between minimal branch handling and advanced predictor research. The Hardcaml RV32IM branch-prediction work, for example, moves from static decode-stage prediction to BTB, RAS, bimodal prediction, and BATAGE in a small in-order pipeline [2312.10426]. BASIC_RV32s adopts the more modest but practically instructive choice of a 2-bit dynamic predictor, which is consistent with its role as a documented microarchitectural roadmap rather than a maximal predictor study.

## 4. CSR support, traps, and the exception model

CSR support enters BASIC_RV32s at the RV32I43F stage through the `CSRFile` module and associated control logic. CSR reads can feed the ALU or write-back path, while CSR writes take data from the ALU or register file. This extends the baseline datapath from pure integer execution toward machine control state manipulation, and it is the prerequisite for the later trap machinery [2510.15887].

Trap and exception support arrives in RV32I46F through two dedicated modules: the Trap Controller and the Exception Detector. The supported mechanisms include `ECALL`, `EBREAK`, basic exception detection such as illegal instruction and possibly misaligned access in early form, trap entry through saving PC and cause information into machine CSRs such as `mepc` and `mcause`, PC redirection to a trap handler, and trap return through `MRET`. In the pipelined RV32I46F_5SP, these mechanisms are integrated with flush control so that traps can remove partially executed instructions and redirect execution cleanly [2510.15887].

The design is explicitly framed around a subset of the RISC-V privileged architecture, centered on machine mode. Full ISA-compliant exception handling is deferred to future work. This is important for interpretation: BASIC_RV32s provides usable trap infrastructure and a simplified machine-mode model, but it does not claim full privileged-architecture completeness [2510.15887].

This suggests a productive connection to formal ISA-level work. The ACL2 model “RV32I in ACL2” provides an operational-semantics RV32I simulator that covers the 37 non-environment RV32I instructions and separates decoding functions from semantic counterparts [2507.19009]. A plausible implication is that BASIC_RV32s supplies a pedagogically rich RTL pathway, while formal models of this kind supply a complementary architectural reference for future refinement or proof-oriented verification.

## 5. RTL organization, SoC realization, verification, and measured results

The RTL is written in Verilog and distributed in two forms: “clean” versions with concise RTL and annotated versions with detailed comments explaining rationale and behavior. The design principles emphasize streamlined I/O, clear module roles, and performance-oriented organization through pipelining, forwarding, and dynamic branch prediction. For FPGA verification, the pipelined core is integrated into `46F5SP_SoC`, which includes on-chip instruction and data memory, GPIO, a Debug UART, and a Benchmark Controller [2510.15887].

The SoC targets the Digilent Nexys Video board with a Xilinx Artix-7 XC7A200T-1SBG484C FPGA. The system runs at 50 MHz and includes a Clock Enable input to support single-step execution. Physical controls include 5 push buttons, 8 LEDs, and one reset button. The UART transmitter connects to the FPGA’s FTDI USB-UART and is used with a terminal for benchmark output and debugging, while debug signals such as `dbg_pc`, `dbg_instruction`, `dbg_ALU_result`, `dbg_mcycle`, and `dbg_minstret` enable instruction-level observation [2510.15887].

| Item | Reported figure |
|---|---|
| Standalone RV32I46F_5SP | 3,010 LUTs, 998 Flip-Flops |
| Full 46F5SP_SoC | 11,660 LUTs, 2,383 Flip-Flops |
| Clock | 50 MHz |
| Dhrystone version | 2.1 |
| Iterations | 2,000 |
| Instructions | 646,640 |
| Cycles | 1,043,092 |
| Average CPI | 1.61 |
| Performance | 1.09 DMIPS/MHz |

Verification combines module-level and core-level Verilog testbenches, stepwise validation of each core version, CSR-based cross-checking through `mcycle` and `minstret`, FPGA execution of Dhrystone, UART logging, and button-driven single-step debugging. The resulting environment is notable because it joins architectural pedagogy to practical FPGA observability: the same platform that measures benchmark performance also supports instruction-by-instruction tracing of pipeline behavior [2510.15887].

## 6. Position within educational and open-source RV32 research

The paper compares RV32I46F_5SP against NEORV32 and DarkRISCV and reports 1.09 DMIPS/MHz for BASIC_RV32s, versus 0.42 DMIPS/MHz for NEORV32 and 0.66 DMIPS/MHz for DarkRISCV. The authors attribute the performance advantage primarily to 5-stage pipelining and dynamic branch prediction. At the same time, the project differentiates itself less by raw benchmark position than by documentation style: it is explicitly framed as a microarchitectural roadmap, accompanied by signal-level block diagrams, development logs, and clean versus annotated RTL [2510.15887].

This placement is clearer when BASIC_RV32s is read against neighboring educational efforts. The Logisim RV32I core is a near-complete teaching implementation centered on a simple, non-pipelined or conceptually single-cycle datapath [2312.01455]. The dynamic-clock single-cycle design with mixed 16-bit and 32-bit instructions explores partial RV32IC support and instruction-dependent timing in simulation rather than a pipelined FPGA-ready roadmap [2211.04455]. BASIC_RV32s differs from both by ending in an FPGA-verified 5-stage core with forwarding, prediction, CSR support, and trap handling, while still preserving the incremental educational pathway [2510.15887].

The project is also recognized as an Intermediate-Level Learning Resource in the official RISC-V `riscv/learn` GitHub repository. That status is consistent with its position between very simple educational toy cores and more feature-rich, production-oriented open-source designs. A recurrent misunderstanding is to equate educational scope with architectural triviality; BASIC_RV32s instead shows that an instructional platform can still incorporate realistic microarchitectural mechanisms such as branch prediction, forwarding, trap entry, and FPGA-based SoC integration [2510.15887].

## 7. Limitations, extensions, and research relevance

The current limitations are explicit. BASIC_RV32s does not yet implement full ISA-compliant exception handling across the privileged architecture; its ISA remains limited to RV32I with Zicsr and minimal trap support; it has no M, C, F, or D extensions; its memory system has no caches or multi-level hierarchy; and while UART debugging and step-by-step execution are available, no standard RISC-V debug module is described [2510.15887].

The planned extensions are correspondingly broad: full privileged-architecture exceptions, hierarchical cache memory, expansion to RV64G, and a more comprehensive guide to modern processor design including multi-issue or superscalar pipelines, more advanced branch predictors, and memory protection or virtual memory. The framework’s modular structure and incremental methodology are presented as the basis for these additions [2510.15887].

A plausible extension path is illustrated by RVCoreP-32IM, which adds RV32M to a five-stage RV32I soft processor through a fork-join multi-cycle execute structure and reports substantial benchmark gains with controlled resource overhead [2010.16171]. BASIC_RV32s does not implement that mechanism, but its staged evolution and clear module boundaries suggest that similar feature growth—adding RV32M, caches, or more advanced branch prediction one unit at a time—fits the project’s design philosophy [2510.15887].

In that sense, BASIC_RV32s is best understood not merely as an endpoint design but as a structured substrate for research and teaching. Its distinctive contribution lies in making the intermediate states of processor construction first-class artifacts: RTL, diagrams, debugging history, FPGA bring-up, and measurable performance are all exposed as part of the architecture itself [2510.15887].

Source: https://www.emergentmind.com/topics/basic_rv32s