YOLE: Neuromorphic Object Detection Baseline
- YOLE is a neuromorphic object detector that converts asynchronous events into a temporally decaying leaky surface for CNN-based detection.
- The model employs a standard YOLO loss and architecture adapted to event-based inputs, striking a balance between frame-based and asynchronous processing.
- Empirical results show YOLE outperforms fcYOLE across multiple datasets, establishing it as a robust baseline for event-based object detection.
Searching arXiv for the primary YOLE paper and nearby event-based object detection work to ground the article. YOLE, short for “You Only Look at Events,” is a neuromorphic object detector that adapts the YOLO detection formulation to event-based cameras by operating on an integrated leaky surface rather than on conventional image frames. It was introduced as the baseline detector in “Asynchronous Convolutional Networks for Object Detection in Neuromorphic Cameras” (Cannici et al., 2018), where it serves both as a practical event-camera detection model and as the reference point for comparison with the paper’s fully asynchronous extension, fcYOLE. In that formulation, YOLE is explicitly characterized as YOLO + leaky surface: a frame-based convolutional detector applied to temporally decayed event accumulations, rather than a network that propagates raw asynchronous events through event-native layers.
1. Event-based sensing and the rationale for YOLE
Event cameras, also called neuromorphic cameras, do not emit conventional frames at a fixed rate. Instead, they output an event only when a pixel undergoes a brightness change, represented as
where and are pixel coordinates, is the timestamp, and is the polarity of the brightness change (Cannici et al., 2018). The properties emphasized for these sensors are microsecond temporal resolution, low power consumption, low bandwidth, high temporal sparsity, and reduced data redundancy.
These sensing characteristics make event cameras attractive for tasks such as tracking, optical flow, and odometry, but object detection is more demanding because it requires both category assignment and localization. YOLE addresses this setting by translating asynchronous event streams into a representation that can be processed by standard CNN machinery. The central motivation is practical: standard CNNs can be trained efficiently, whereas deep SNNs remain difficult to train for complex tasks; at the same time, a temporally decaying event surface preserves temporal ordering better than naive frame binning and provides a simple baseline for comparing conventional frame-style processing with explicitly asynchronous architectures (Cannici et al., 2018).
A common misconception is to treat YOLE as an event-native asynchronous detector. It is not. The essential design decision is event integration first, CNN second. This distinguishes it from methods that attempt to exploit sparsity throughout the network.
2. Leaky-surface input representation
YOLE consumes a leaky surface, a dense 2D array that is updated as events arrive. For an event at time , the surface update is
with , leak rate , increment 0, and 1 enforcing non-negativity (Cannici et al., 2018). The paper fixes 2 and varies only 3 depending on the dataset.
Operationally, every event adds activation at its pixel while the full surface decays over time. This yields a frame-like representation that still carries temporal dynamics. The stated reasons for preferring the leaky surface over simple binary frame accumulation are that it avoids treating all events equally, maintains a notion of temporal decay, and handles noise more smoothly (Cannici et al., 2018).
This suggests that YOLE’s representation is best understood as a compromise between two incompatible desiderata: strict event-driven computation and compatibility with dense CNN training. It is temporally informed, but not event-native in the strict asynchronous sense.
3. Detection architecture and YOLO formulation
YOLE is a standard CNN detector trained with the YOLO loss, operating on leaky surfaces rather than RGB images (Cannici et al., 2018). The input surface size is 4. The field of view is divided into a grid, and each region predicts 5 bounding boxes together with class probabilities over 6 classes. For the MNIST-style detection setup, the paper uses a 7 grid on a 8 surface; for N-Caltech101, it uses a 9 grid.
The network predicts the standard YOLO-style outputs: objectness, bounding-box coordinates, bounding-box size, and class probabilities. The hidden layers use Leaky ReLU, while the output layer uses linear activation. The exact architecture depends on the dataset. For MNIST-based experiments, the design is inspired by LeNet and uses 6 convolution/pooling layers for feature extraction. For N-Caltech101, the network is inspired by VGG16, but simplified to one layer per convolutional block (Cannici et al., 2018).
The paper is explicit that YOLE is not the full original YOLO network architecture. In this context, “YOLO” primarily denotes the training and detection formulation, not a verbatim reuse of the original layer-by-layer design. That distinction is important for interpreting the model: YOLE inherits the one-stage regression-style detection philosophy, while adapting the backbone to the constraints of event-derived inputs.
YOLE is trained with the standard multi-objective YOLO loss, covering localization error, confidence or objectness error, classification error, and no-object penalty. For some experiments, better performance was obtained with modified hyperparameters:
- 0
- 1
4. Training protocol and dataset construction
The common training setup uses Adam with learning rate 2, 3, 4, and 5 (Cannici et al., 2018). The first four convolutional layers are initialized from a recognition network pretrained to classify the target objects, while the remaining layers use Glorot initialization. Early stopping is applied using validation sets of the same size as the test sets.
Batch sizes are dataset-specific and chosen to fit GPU memory:
- Shifted N-MNIST: 10
- Shifted MNIST-DVS: 40
- N-Caltech101: 40
- Blackboard MNIST: 25
- OD-Poker-DVS: 35
A substantial part of the work consists of extending or creating event-based object detection datasets (Cannici et al., 2018). The datasets used are N-Caltech101, Shifted N-MNIST, Shifted MNIST-DVS, OD-Poker-DVS, and Blackboard MNIST.
Shifted N-MNIST is built from N-MNIST by placing one or two digits in non-overlapping positions in a larger field of view. Bounding boxes are extracted by integrating events into a frame, removing noise, clustering with DBSCAN-like grouping, and shifting boxes to the final digit position. The supplementary material specifies noise filtering with 6, radius 7, and minimum area threshold 8.
Shifted MNIST-DVS follows a similar procedure using MNIST-DVS samples at multiple scales: scale4, scale8, and scale16.
OD-Poker-DVS extends Poker-DVS to object detection by using tracking to annotate pips with bounding boxes in the original recordings; samples are split into short time windows of approximately 1.5 ms.
Blackboard MNIST is a synthetic dataset generated with the DAVIS simulator, using chalk-like white digits on a blackboard with random positions, scales, and camera trajectories. It is organized into EASY, MEDIUM, and HARD levels, and introduces varying scale, multiple objects, motion-induced event bursts, and partial visibility filtering (Cannici et al., 2018).
5. Empirical performance and comparison with fcYOLE
The paper evaluates YOLE using accuracy, mean Average Precision (mAP), and per-class Average Precision (AP) on N-Caltech101. Accuracy is computed by matching ground-truth boxes to predicted boxes with highest IoU (Cannici et al., 2018).
The principal quantitative comparison is between YOLE and fcYOLE, the paper’s asynchronous fully convolutional extension.
| Dataset | fcYOLE | YOLE |
|---|---|---|
| Shifted MNIST-DVS | 94.0 acc / 87.4 mAP | 96.1 acc / 92.0 mAP |
| Blackboard MNIST | 88.5 acc / 84.7 mAP | 90.4 acc / 87.4 mAP |
| OD-Poker-DVS | 79.10 acc / 78.69 mAP | 87.3 acc / 82.2 mAP |
| N-Caltech101 | 57.1 acc / 26.9 mAP | 64.9 acc / 39.8 mAP |
YOLE outperforms fcYOLE on all reported datasets (Cannici et al., 2018). The paper attributes this difference not to an intrinsic weakness of event-based layers, but to the fact that fcYOLE removes fully connected layers and therefore has less expressive power. That is a notable interpretive point: the comparison does not imply that asynchronous computation is inherently inferior, only that the particular fully convolutional simplification imposed a representational cost.
The paper also studies robustness on Shifted N-MNIST under progressively harder variants.
| Variant | Accuracy | mAP |
|---|---|---|
| v1 | 94.9 | 91.3 |
| v2 | 91.7 | 87.9 |
| v2* | 94.7 | 90.5 |
| v2fr | 88.6 | 81.5 |
| v2fr+ns | 85.5 | 77.4 |
The progression shows that YOLE can detect multiple objects, but performance declines with clutter and noise. The paper links part of the recovered performance to the modified YOLO loss weights 9 and 0 (Cannici et al., 2018).
For N-Caltech101, the paper reports that YOLE achieves much better AP than fcYOLE on many classes, especially those with many training examples and less intra-class variability. Performance remains weak for classes with few training samples, high variability, and class imbalance. The paper nevertheless notes strong AP on categories such as motorbikes, airplanes, faces_easy, watch, minaret, and umbrellas (Cannici et al., 2018).
6. Relation to asynchronous detection, limitations, and significance
YOLE is best understood in relation to fcYOLE. Whereas YOLE converts events into a single leaky surface and then applies a conventional dense CNN, fcYOLE is a fully convolutional, asynchronous event-based detector that introduces e-conv, e-max-pool, layer-wise internal states, update matrices 1, and positional indices 2 for pooling (Cannici et al., 2018). In conceptual terms, the distinction is straightforward:
- YOLE: event integration first, CNN second
- fcYOLE: event-driven computation throughout the network
This difference also appears in timing behavior. Events were grouped into 10 ms batches, and timings were averaged over 1000 runs. On Shifted N-MNIST, the event-based approach achieves 22.6 ms per batch and about 2× speedup. On Blackboard MNIST, the event-based approach takes 43.2 ms per batch, while the conventional network takes 34.6 ms per batch (Cannici et al., 2018). The authors therefore conclude that asynchronous CNNs are advantageous when changes are localized and sparse, but can be slower when changes are widespread and noisy. They further note that their asynchronous implementation is beneficial only up to about 80% event sparsity.
Within that comparison, YOLE’s advantages are explicit: it is a simple and practical baseline for event-camera object detection, uses standard CNN training, works well on several datasets, preserves some temporal information via leaky integration, and achieves better performance than fcYOLE in the reported experiments (Cannici et al., 2018). Its limitations are equally clear: it does not exploit event sparsity directly, recomputes a conventional CNN on a dense surface, can degrade significantly under heavy noise or clutter, and depends on the quality of leaky-surface parameters.
The broader significance of YOLE lies in its bridging role. It demonstrates that standard object detection can be adapted to neuromorphic cameras through temporally decaying event integration, thereby linking the mature ecosystem of frame-based deep detection with the emerging domain of event-based sensing (Cannici et al., 2018). A plausible implication is that YOLE’s enduring value is methodological as much as empirical: it establishes a strong baseline against which truly asynchronous detectors can be judged, while showing that event-camera object detection need not begin with a fully event-native learning stack.