BASIC_RV32s RISC-V Microarchitecture Evolution
- BASIC_RV32s is an open-source educational framework that defines a documented roadmap for evolving a RISC-V RV32I processor from a simple single-cycle to an advanced 5-stage pipelined core.
- It progressively integrates critical features such as CSR support, trap handling, dynamic branch prediction, and hazard resolution to enhance both instructional value and practical performance.
- The framework provides detailed RTL code, block diagrams, and debugging logs that bridge theoretical architectural design with applied FPGA-based implementation.
Searching arXiv for the primary BASIC_RV32s paper and closely related RV32/RISC-V educational microarchitecture references. BASIC_RV32s is an open-source educational framework and microarchitectural roadmap for implementing a RISC-V RV32I processor. It is defined not as a single reference core, but as an explicitly documented evolution of several cores, beginning with a textbook-style single-cycle implementation and culminating in a relatively complete 5-stage pipelined core with full hazard forwarding, dynamic branch prediction, and exception handling, verified on FPGA within a System-on-Chip (SoC) that supports UART-based debugging and reports 1.09 DMIPS/MHz at 50 MHz. By releasing Register-Transfer Level (RTL) source code, signal-level logic block diagrams, and development logs under the MIT license, the project is positioned as a reproducible bridge between architectural theory and synthesizable hardware practice (Kang et al., 4 Sep 2025).
1. Concept, goals, and architectural scope
BASIC_RV32s is organized around a step-by-step, Patterson & Hennessy–style path from theory to working hardware. Its stated goals are to expose the full progression from simple to more capable cores, to document not only the final design but also intermediate versions together with engineering decisions and debugging history, and to make the framework reproducible and modifiable by releasing RTL, diagrams, and development logs under the MIT license (Kang et al., 4 Sep 2025).
The architectural target is the RISC-V RV32I base integer ISA, augmented over the course of the roadmap by the Zicsr extension and a minimal trap-oriented subset consisting of ECALL, EBREAK, and MRET. The supported instruction classes include integer arithmetic and logic, loads and stores, conditional branches, jumps, LUI, AUIPC, CSR access instructions such as CSRRW, CSRRS, CSRRC and their immediate forms, and machine-mode trap-related instructions in later variants. The privileged scope is explicitly machine-mode oriented, and the trap controller and exception detector are structured to be compatible with the machine-mode trap model, while full ISA-compliant exception handling is reserved for future work (Kang et al., 4 Sep 2025).
| Core variant | ISA / scope | Main addition |
|---|---|---|
| RV32I37F | 37 instructions from base RV32I | Baseline single-cycle datapath and control |
| RV32I43F | RV32I37F + Zicsr | CSRFile and CSR control paths |
| RV32I46F | RV32I43F + ECALL, EBREAK, MRET |
Trap Controller and Exception Detector |
| RV32I46F_5SP | Same ISA coverage as RV32I46F | 5-stage pipeline, hazard logic, branch prediction |
A common misconception is to treat BASIC_RV32s as only the final pipelined core. The project description rejects that interpretation explicitly: it is not just a single reference core, but a documented microarchitectural evolution. This gives it a dual identity as both an instructional reference and a realistic, synthesizable RISC-V microarchitecture (Kang et al., 4 Sep 2025).
2. Evolution from single-cycle execution to pipelined execution
The roadmap proceeds through four concrete designs: a single-cycle RV32I37F, a CSR-capable RV32I43F, a trap-capable RV32I46F, and the final 5-stage pipelined RV32I46F_5SP. In the initial single-cycle form, the Program Counter, Instruction Memory, Register File, ALU, and Data Memory are each used once per instruction, and fetch, decode, execute, memory access, and write-back all complete within one clock cycle. The control path is centralized and decodes opcode and function fields into ALU operation, register write enable, memory control, and PC selection signals, closely matching the canonical single-cycle organization described in Patterson and Hennessy (Kang et al., 4 Sep 2025).
The transition to RV32I43F introduces a dedicated CSRFile together with control logic for CSR addressing, operand selection, and updates consistent with Zicsr semantics. RV32I46F then adds a Trap Controller and Exception Detector, enabling handling of ECALL, EBREAK, simple synchronous exceptions, trap entry via machine CSRs such as mepc and mcause, and trap return through MRET. These additions remain embedded in a single-cycle microarchitecture, but they establish the state and control machinery later reused in the pipelined core (Kang et al., 4 Sep 2025).
The final RV32I46F_5SP reorganizes the design into the classic 5-stage pipeline: IF, ID, EX, MEM, and WB. The paper’s signal-level block diagram identifies a PC Controller, a 2-bit dynamic Branch Predictor in IF, a Hazard Unit, a Forward Unit, pipeline-integrated CSR logic, an Exception Detector and Trap Controller, and a Debug Interface exporting signals such as dbg_pc, dbg_instruction, dbg_ALU_result, mcycle, and minstret. This organization places BASIC_RV32s within the same broad microarchitectural lineage as other educational and FPGA-oriented five-stage RV32 cores, including RVCoreP, which likewise centers optimization around branch prediction, hazard handling, and FPGA timing paths (Miyazaki et al., 2020).
This staged progression is significant because it preserves continuity between abstract textbook datapaths and concrete implementation artifacts. Each added feature is presented not as an isolated enhancement but as a structural transformation of the preceding core, allowing datapath, control, and debugging considerations to remain visible across versions (Kang et al., 4 Sep 2025).
3. Hazard handling, branch prediction, and control-flow behavior
The pipelined RV32I46F_5SP includes explicit Hazard Unit and Forwarding Unit logic. The primary data hazard class is RAW, while WAR and WAW are stated not to occur in this simple in-order pipeline. Data hazards are resolved through forwarding from the execution, memory access, and write-back stages, and a dependent load-use case is handled by stalling the pipeline for one cycle, holding the PC and IF/ID state while injecting a bubble into ID/EX. This is presented as the standard Patterson & Hennessy solution for a five-stage in-order machine (Kang et al., 4 Sep 2025).
Control hazards are managed through a 2-bit dynamic branch predictor located in the IF stage, branch resolution in EX, and pipeline flush on misprediction. On each branch, the predictor supplies a taken or not-taken decision; EX later evaluates the actual condition and target. If the prediction is wrong, incorrect-path instructions are flushed and the PC is redirected, while the predictor’s saturating counter state is updated. Structural hazards are not emphasized; the design description suggests they are avoided by design through separate or logically partitioned instruction and data memories, and implementation-level memory mapping conflicts were resolved by bypass paths between instruction and data memory (Kang et al., 4 Sep 2025).
The performance objective of the pipelined core is near-1 CPI on straight-line code, but the measured Dhrystone average is 1.61 cycles per instruction, reflecting remaining stalls and branch mispredictions. The same section associates the reported 1.09 DMIPS/MHz with the combined impact of pipelining, forwarding, and dynamic prediction (Kang et al., 4 Sep 2025).
In broader context, BASIC_RV32s occupies an intermediate point between minimal branch handling and advanced predictor research. The Hardcaml RV32IM branch-prediction work, for example, moves from static decode-stage prediction to BTB, RAS, bimodal prediction, and BATAGE in a small in-order pipeline (Saveau, 2023). BASIC_RV32s adopts the more modest but practically instructive choice of a 2-bit dynamic predictor, which is consistent with its role as a documented microarchitectural roadmap rather than a maximal predictor study.
4. CSR support, traps, and the exception model
CSR support enters BASIC_RV32s at the RV32I43F stage through the CSRFile module and associated control logic. CSR reads can feed the ALU or write-back path, while CSR writes take data from the ALU or register file. This extends the baseline datapath from pure integer execution toward machine control state manipulation, and it is the prerequisite for the later trap machinery (Kang et al., 4 Sep 2025).
Trap and exception support arrives in RV32I46F through two dedicated modules: the Trap Controller and the Exception Detector. The supported mechanisms include ECALL, EBREAK, basic exception detection such as illegal instruction and possibly misaligned access in early form, trap entry through saving PC and cause information into machine CSRs such as mepc and mcause, PC redirection to a trap handler, and trap return through MRET. In the pipelined RV32I46F_5SP, these mechanisms are integrated with flush control so that traps can remove partially executed instructions and redirect execution cleanly (Kang et al., 4 Sep 2025).
The design is explicitly framed around a subset of the RISC-V privileged architecture, centered on machine mode. Full ISA-compliant exception handling is deferred to future work. This is important for interpretation: BASIC_RV32s provides usable trap infrastructure and a simplified machine-mode model, but it does not claim full privileged-architecture completeness (Kang et al., 4 Sep 2025).
This suggests a productive connection to formal ISA-level work. The ACL2 model “RV32I in ACL2” provides an operational-semantics RV32I simulator that covers the 37 non-environment RV32I instructions and separates decoding functions from semantic counterparts (Kwan, 25 Jul 2025). A plausible implication is that BASIC_RV32s supplies a pedagogically rich RTL pathway, while formal models of this kind supply a complementary architectural reference for future refinement or proof-oriented verification.
5. RTL organization, SoC realization, verification, and measured results
The RTL is written in Verilog and distributed in two forms: “clean” versions with concise RTL and annotated versions with detailed comments explaining rationale and behavior. The design principles emphasize streamlined I/O, clear module roles, and performance-oriented organization through pipelining, forwarding, and dynamic branch prediction. For FPGA verification, the pipelined core is integrated into 46F5SP_SoC, which includes on-chip instruction and data memory, GPIO, a Debug UART, and a Benchmark Controller (Kang et al., 4 Sep 2025).
The SoC targets the Digilent Nexys Video board with a Xilinx Artix-7 XC7A200T-1SBG484C FPGA. The system runs at 50 MHz and includes a Clock Enable input to support single-step execution. Physical controls include 5 push buttons, 8 LEDs, and one reset button. The UART transmitter connects to the FPGA’s FTDI USB-UART and is used with a terminal for benchmark output and debugging, while debug signals such as dbg_pc, dbg_instruction, dbg_ALU_result, dbg_mcycle, and dbg_minstret enable instruction-level observation (Kang et al., 4 Sep 2025).
| Item | Reported figure |
|---|---|
| Standalone RV32I46F_5SP | 3,010 LUTs, 998 Flip-Flops |
| Full 46F5SP_SoC | 11,660 LUTs, 2,383 Flip-Flops |
| Clock | 50 MHz |
| Dhrystone version | 2.1 |
| Iterations | 2,000 |
| Instructions | 646,640 |
| Cycles | 1,043,092 |
| Average CPI | 1.61 |
| Performance | 1.09 DMIPS/MHz |
Verification combines module-level and core-level Verilog testbenches, stepwise validation of each core version, CSR-based cross-checking through mcycle and minstret, FPGA execution of Dhrystone, UART logging, and button-driven single-step debugging. The resulting environment is notable because it joins architectural pedagogy to practical FPGA observability: the same platform that measures benchmark performance also supports instruction-by-instruction tracing of pipeline behavior (Kang et al., 4 Sep 2025).
6. Position within educational and open-source RV32 research
The paper compares RV32I46F_5SP against NEORV32 and DarkRISCV and reports 1.09 DMIPS/MHz for BASIC_RV32s, versus 0.42 DMIPS/MHz for NEORV32 and 0.66 DMIPS/MHz for DarkRISCV. The authors attribute the performance advantage primarily to 5-stage pipelining and dynamic branch prediction. At the same time, the project differentiates itself less by raw benchmark position than by documentation style: it is explicitly framed as a microarchitectural roadmap, accompanied by signal-level block diagrams, development logs, and clean versus annotated RTL (Kang et al., 4 Sep 2025).
This placement is clearer when BASIC_RV32s is read against neighboring educational efforts. The Logisim RV32I core is a near-complete teaching implementation centered on a simple, non-pipelined or conceptually single-cycle datapath (Patil et al., 2023). The dynamic-clock single-cycle design with mixed 16-bit and 32-bit instructions explores partial RV32IC support and instruction-dependent timing in simulation rather than a pipelined FPGA-ready roadmap (Chen et al., 2022). BASIC_RV32s differs from both by ending in an FPGA-verified 5-stage core with forwarding, prediction, CSR support, and trap handling, while still preserving the incremental educational pathway (Kang et al., 4 Sep 2025).
The project is also recognized as an Intermediate-Level Learning Resource in the official RISC-V riscv/learn GitHub repository. That status is consistent with its position between very simple educational toy cores and more feature-rich, production-oriented open-source designs. A recurrent misunderstanding is to equate educational scope with architectural triviality; BASIC_RV32s instead shows that an instructional platform can still incorporate realistic microarchitectural mechanisms such as branch prediction, forwarding, trap entry, and FPGA-based SoC integration (Kang et al., 4 Sep 2025).
7. Limitations, extensions, and research relevance
The current limitations are explicit. BASIC_RV32s does not yet implement full ISA-compliant exception handling across the privileged architecture; its ISA remains limited to RV32I with Zicsr and minimal trap support; it has no M, C, F, or D extensions; its memory system has no caches or multi-level hierarchy; and while UART debugging and step-by-step execution are available, no standard RISC-V debug module is described (Kang et al., 4 Sep 2025).
The planned extensions are correspondingly broad: full privileged-architecture exceptions, hierarchical cache memory, expansion to RV64G, and a more comprehensive guide to modern processor design including multi-issue or superscalar pipelines, more advanced branch predictors, and memory protection or virtual memory. The framework’s modular structure and incremental methodology are presented as the basis for these additions (Kang et al., 4 Sep 2025).
A plausible extension path is illustrated by RVCoreP-32IM, which adds RV32M to a five-stage RV32I soft processor through a fork-join multi-cycle execute structure and reports substantial benchmark gains with controlled resource overhead (Islam et al., 2020). BASIC_RV32s does not implement that mechanism, but its staged evolution and clear module boundaries suggest that similar feature growth—adding RV32M, caches, or more advanced branch prediction one unit at a time—fits the project’s design philosophy (Kang et al., 4 Sep 2025).
In that sense, BASIC_RV32s is best understood not merely as an endpoint design but as a structured substrate for research and teaching. Its distinctive contribution lies in making the intermediate states of processor construction first-class artifacts: RTL, diagrams, debugging history, FPGA bring-up, and measurable performance are all exposed as part of the architecture itself (Kang et al., 4 Sep 2025).