Benchmarking methods that account for fundamental architectural differences of neuromorphic computing

Develop benchmarking methodologies that rigorously account for the fundamental architectural differences of neuromorphic computing relative to conventional and GPU-based systems, including a principled decomposition of energy and performance costs across neuron dynamics, spike generation, synaptic operations, and spike communication.

Background

The authors call for empirical validation of their theoretical framework but note that existing neuromorphic benchmarks are largely application-driven and do not capture architectural differences. The distinct cost structure of neuromorphic systems—especially communication and event-driven computation—complicates apples-to-apples comparisons.

A comprehensive, architecture-aware benchmarking methodology is needed to evaluate neuromorphic systems fairly and to guide algorithm design, yet how to construct such benchmarks remains unresolved.

References

The benchmarking of NMC has been a small but growing endeavor , however these early efforts have been more application-driven as is typical in machine learning and it remains an open question how to account for the fundamental architectural differences of NMC.

Neuromorphic Computing: A Theoretical Framework for Time, Space, and Energy Scaling  (2507.17886 - Aimone, 23 Jul 2025) in Section 7.2 Limitations of this analysis

Based on these preliminary observations, we outline below a number of open questions in future work. First, energy-efficiency evaluation requires a unified and physically meaningful accounting rule. Directly comparing a heuristic spike-count estimate of an SNN with a FLOP-style approximation of an ANN is insufficient. Both models should be analyzed under an operation-level formulation---ANN inference is measured in multiply--accumulate operations, and event-driven SNN inference is measured in accumulate operations.

LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service  (2608.13144 - Xiang et al., 13 Aug 2026) in Section 6.2, subsection “Exploration: An SNN-Based GuardNet”

Two limitations remain open: a clear language-modeling gap, most visible on LAMBADA, and the absence of measurements on physical neuromorphic hardware, which is required before any system-level energy claim can be made.

Large Language Models with At Most One Spike per Neuron  (2609.05151 - Zhao et al., 4 Sep 2026) in Section 5, “Energy Consumption Analysis”; Section 6, “Conclusion”