Papers
Topics
Authors
Recent
Search
2000 character limit reached

Floorline Performance Model in Neuromorphic Accelerators

Updated 4 December 2025
  • Floorline Performance Model is a framework that characterizes and bounds neuromorphic accelerator performance based on per-core maximum synaptic, compute, and traffic loads.
  • The model employs a two-stage optimization methodology using sparsity-aware training and floorline-informed partitioning to pinpoint bottlenecks and enhance performance.
  • Empirical validations on platforms like Loihi 2 demonstrate significant speedup and energy improvements by transitioning workloads among memory-bound, compute-bound, and traffic-bound states.

A floorline performance model is a theoretical and empirical framework for characterizing, bounding, and optimizing execution time and energy efficiency in neuromorphic accelerators. In contrast to conventional roofline models, which primarily address memory and compute ceilings, the floorline model exposes three competitive bottleneck states—memory-bound, compute-bound, and traffic-bound—and leverages per-core maximum statistics for synaptic, compute, and message loads to predict actual performance. The model enables pinpoint allocation of optimization effort and informs a two-stage methodology for workload restructuring on real neuromorphic hardware (Yik et al., 26 Nov 2025).

1. Theoretical Basis and Bottleneck States

Neuromorphic accelerators, comprised of arrays of neurocores, execute ML inference using event-driven dataflows and spatially expanded architectures that co-locate memory and computation. In each timestep, every core carries out three main operations:

  • Synaptic operations (synops): Retrieving weights from local memory and performing accumulations.
  • Neuron activation computations (act-comp): Applying activation nonlinearities and updating internal neuronal state.
  • NoC message traffic (msg): Emitting sparse activation spikes over the on-chip network.

Because all relevant state (weights, neuron activations, messages) resides on-chip, the cost of each operation may be comparable, and any of the three may dominate step time. This yields three distinct bottleneck states:

  • M1. Memory-bound (synops-bound): The slowest core dictates the timestep duration due to its peak synop load. Time scales with the maximum per-core synaptic operations.
  • M2. Compute-bound: When architectural choices or sparsity reduce synop load, the bottleneck can shift to neuron activation computations.
  • M3. Traffic-bound: Design choices (e.g., high core utilization or layer partitioning) can provoke NoC congestion; synchronization waits for the most spike-active core.

A critical observation is that per-core maximum (not global aggregate) synop, act-comp, or spike loads control which bottleneck state is active.

2. Mathematical Formulation of Performance Bounds

Let each neurocore cc perform ScS_c synops, CcC_c activation computations, and McM_c outgoing spike messages per timestep. Define maximum intensities:

Smax⁡=max⁡cSc,Cmax⁡=max⁡cCc,Mmax⁡=max⁡cMcS_{\max} = \max_c S_c, \quad C_{\max} = \max_c C_c, \quad M_{\max} = \max_c M_c

Architectural peak rates are:

  • BmemB_{\rm mem}: peak synop bandwidth (synops/sec)
  • FpeakF_{\rm peak}: peak neuron-compute (activations/sec)
  • BnocB_{\rm noc}: peak NoC bandwidth (spikes/sec)

Lower bounds on timestep time:

  • Memory-bound: Tmem=Smax⁡Bmem\displaystyle T_{\rm mem} = \frac{S_{\max}}{B_{\rm mem}}
  • Compute-bound: Tcomp=Cmax⁡Fpeak\displaystyle T_{\rm comp} = \frac{C_{\max}}{F_{\rm peak}}
  • Traffic-bound: ScS_c0

Actual performance is:

ScS_c1

Due to barrier synchronization, the highest of these three sets the realized timestep duration.

3. Floorline Visualization and Interpretation

The floorline plot visualizes per-core synop intensity ScS_c2 on the x-axis versus measured or predicted timestep time ScS_c3 on the y-axis, typically as a log–log plot. The main analytic boundaries are:

  • Memory-bound slope: ScS_c4 (linear with ScS_c5, slope +1 in log–log, 45°)
  • Compute floor: As ScS_c6, ScS_c7 plateaus at ScS_c8
  • Traffic ceiling: Workloads above both boundaries are set by ScS_c9

Each workload-adapted network, once mapped and partitioned, yields a point CcC_c0. The point’s relationship to the slope and floor precisely indicates which bottleneck is active and which mapping, partitioning, or sparsity transformation is likely to yield performance improvement.

4. Analytical and Empirical Model Validation

Analytical modeling predicts scaling trends for synops, activation computes, and message loads as functions of activation sparsity CcC_c1, weight sparsity CcC_c2, network width CcC_c3, and number of cores CcC_c4:

  • Synops: CcC_c5
  • Activation computes: CcC_c6
  • Messages (for downstream layer with CcC_c7 cores): CcC_c8

Microbenchmarks on three neuromorphic platforms—Brainchip AKD1000 (80 cores), Synsense Speck (9 cores), and Intel Loihi 2 (120 cores)—were used for quantitative parameter calibration and validation. Benchmarks sweep network sparsity, partitioning, and core mapping:

  • Workloads trace the analytic memory-bound line CcC_c9 until McM_c0 is reduced to point where the compute floor McM_c1 dominates.
  • Aggressive partitioning further lowers McM_c2 and compute floor, but increases power.
  • Extensive core utilization may push workloads above the slope–floor envelope to traffic-bound; strided placement heuristics (on Loihi 2) restore memory-bound scaling.
  • These empirical findings yielded precise parameters for McM_c3, McM_c4, and McM_c5, confirming the model’s fidelity.

5. Floorline-Guided Two-Stage Optimization Methodology

Optimization follows a two-stage workflow:

Stage 1: Sparsity-Aware Training

  • Apply activation and synop-count regularizers targeting spike-based sparsity (e.g., Transformed McM_c6 on AKD1000 and PilotNet, synop-count penalties on Speck).
  • Employ one-shot pruning with fine-tuning for Loihi 2.
  • Adjust per-layer sparsity schedules to balance core loads, preventing per-core bottleneck conditions.

Stage 2: Floorline-Informed Partitioning & Mapping

  • For workloads on memory-bound slope, partition core with peak McM_c7; for compute floor, partition core with peak McM_c8.
  • For traffic-bound workloads, remap using “strided” core assignment to mitigate McM_c9.
  • Only retain transformations which lower Smax⁡=max⁡cSc,Cmax⁡=max⁡cCc,Mmax⁡=max⁡cMcS_{\max} = \max_c S_c, \quad C_{\max} = \max_c C_c, \quad M_{\max} = \max_c M_c0, with explicit backtracking to avoid unnecessary power escalation.
  • Stop once the workload lies on the floorline envelope for the given network and hardware.

6. Case Study: PilotNet Optimization on Loihi 2

Applying the methodology to the PilotNet CNN on Loihi 2 yields concrete improvements:

  • Hardware limits: Smax⁡=max⁡cSc,Cmax⁡=max⁡cCc,Mmax⁡=max⁡cMcS_{\max} = \max_c S_c, \quad C_{\max} = \max_c C_c, \quad M_{\max} = \max_c M_c1 synops/s, Smax⁡=max⁡cSc,Cmax⁡=max⁡cCc,Mmax⁡=max⁡cMcS_{\max} = \max_c S_c, \quad C_{\max} = \max_c C_c, \quad M_{\max} = \max_c M_c2 activations/s, Smax⁡=max⁡cSc,Cmax⁡=max⁡cCc,Mmax⁡=max⁡cMcS_{\max} = \max_c S_c, \quad C_{\max} = \max_c C_c, \quad M_{\max} = \max_c M_c3 spikes/s.
  • Initial trained network: Smax⁡=max⁡cSc,Cmax⁡=max⁡cCc,Mmax⁡=max⁡cMcS_{\max} = \max_c S_c, \quad C_{\max} = \max_c C_c, \quad M_{\max} = \max_c M_c4, Smax⁡=max⁡cSc,Cmax⁡=max⁡cCc,Mmax⁡=max⁡cMcS_{\max} = \max_c S_c, \quad C_{\max} = \max_c C_c, \quad M_{\max} = \max_c M_c5, Smax⁡=max⁡cSc,Cmax⁡=max⁡cCc,Mmax⁡=max⁡cMcS_{\max} = \max_c S_c, \quad C_{\max} = \max_c C_c, \quad M_{\max} = \max_c M_c6.
  • Compute step times:
    • Smax⁡=max⁡cSc,Cmax⁡=max⁡cCc,Mmax⁡=max⁡cMcS_{\max} = \max_c S_c, \quad C_{\max} = \max_c C_c, \quad M_{\max} = \max_c M_c7 ms
    • Smax⁡=max⁡cSc,Cmax⁡=max⁡cCc,Mmax⁡=max⁡cMcS_{\max} = \max_c S_c, \quad C_{\max} = \max_c C_c, \quad M_{\max} = \max_c M_c8 ms
    • Smax⁡=max⁡cSc,Cmax⁡=max⁡cCc,Mmax⁡=max⁡cMcS_{\max} = \max_c S_c, \quad C_{\max} = \max_c C_c, \quad M_{\max} = \max_c M_c9 ms
    • BmemB_{\rm mem}0 ms (compute-bound)
  • Partitioning to more cores: ~BmemB_{\rm mem}1 reduces stepwise from compute-bound to traffic-bound, then remapping returns to compute floor, achieving BmemB_{\rm mem}2 ms, a 4BmemB_{\rm mem}3 speedup, with energy per step improved roughly 3BmemB_{\rm mem}4.

These results demonstrate stepwise optimization along the floorline axes, moving from compute to traffic-bound to improved compute-bound via partitioning and core remapping.

7. Significance and Generalization

The floorline performance model extends roofline analysis principles to account for neuromorphic architectures featuring tightly interleaved on-chip memory, computation, and message traffic. By positioning workloads in BmemB_{\rm mem}5 space, practitioners can immediately identify bottlenecks and allocate optimization effort (sparsity, partitioning, remapping) to approach theoretical performance bounds. The two-stage optimization regime yields multi-fold improvements in runtime and energy efficiency for diverse workloads across current neuromorphic architectures and is extensible to future chip designs (Yik et al., 26 Nov 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Floorline Performance Model.