---
title: 'YOLE: Neuromorphic Object Detection Baseline'
url: https://www.emergentmind.com/topics/yole
type: topic
---

# YOLE: Neuromorphic Object Detection Baseline

Searching arXiv for the primary YOLE paper and nearby event-based object detection work to ground the article.
YOLE, short for **“You Only Look at Events,”** is a neuromorphic object detector that adapts the YOLO detection formulation to **event-based cameras** by operating on an **integrated leaky surface** rather than on conventional image frames. It was introduced as the baseline detector in “Asynchronous Convolutional Networks for Object Detection in Neuromorphic Cameras” [1805.07931], where it serves both as a practical event-camera detection model and as the reference point for comparison with the paper’s fully asynchronous extension, **fcYOLE**. In that formulation, YOLE is explicitly characterized as **YOLO + leaky surface**: a **frame-based** convolutional detector applied to temporally decayed event accumulations, rather than a network that propagates raw asynchronous events through event-native layers.

## 1. Event-based sensing and the rationale for YOLE

Event cameras, also called **neuromorphic cameras**, do not emit conventional frames at a fixed rate. Instead, they output an event only when a pixel undergoes a **brightness change**, represented as

\[
\mathbf{e} = \langle x, y, ts, p \rangle
\]

where \(x\) and \(y\) are pixel coordinates, \(ts\) is the timestamp, and \(p \in \{1,-1\}\) is the polarity of the brightness change [1805.07931]. The properties emphasized for these sensors are **microsecond temporal resolution**, **low power consumption**, **low bandwidth**, **high temporal sparsity**, and reduced data redundancy.

These sensing characteristics make event cameras attractive for tasks such as tracking, optical flow, and odometry, but object detection is more demanding because it requires both category assignment and localization. YOLE addresses this setting by translating asynchronous event streams into a representation that can be processed by standard CNN machinery. The central motivation is practical: standard CNNs can be trained efficiently, whereas deep SNNs remain difficult to train for complex tasks; at the same time, a temporally decaying event surface preserves temporal ordering better than naive frame binning and provides a simple baseline for comparing conventional frame-style processing with explicitly asynchronous architectures [1805.07931].

A common misconception is to treat YOLE as an event-native asynchronous detector. It is not. The essential design decision is **event integration first, CNN second**. This distinguishes it from methods that attempt to exploit sparsity throughout the network.

## 2. Leaky-surface input representation

YOLE consumes a **leaky surface**, a dense 2D array that is updated as events arrive. For an event at time \(ts^t\), the surface update is

\[
q_{x_s, y_s}^t = \max\left(p_{x_s, y_s}^{t-1} - \lambda \cdot \Delta_{ts}, 0\right)
\]

\[
p_{x_s, y_s}^t =
\begin{cases}
q_{x_s, y_s}^t + \Delta_{incr} & \text{if } (x_s, y_s)^t = (x_e, y_e)^t \\
q_{x_s, y_s}^t & \text{otherwise}
\end{cases}
\]

with \(\Delta_{ts} = ts^t - ts^{t-1}\), leak rate \(\lambda\), increment \(\Delta_{incr}\), and \(\max(\cdot,0)\) enforcing non-negativity [1805.07931]. The paper fixes \(\Delta_{incr}=1\) and varies only \(\lambda\) depending on the dataset.

Operationally, every event adds activation at its pixel while the full surface decays over time. This yields a frame-like representation that still carries temporal dynamics. The stated reasons for preferring the leaky surface over simple binary frame accumulation are that it avoids treating all events equally, maintains a notion of temporal decay, and handles noise more smoothly [1805.07931].

This suggests that YOLE’s representation is best understood as a compromise between two incompatible desiderata: strict event-driven computation and compatibility with dense CNN training. It is temporally informed, but not event-native in the strict asynchronous sense.

## 3. Detection architecture and YOLO formulation

YOLE is a **standard CNN detector trained with the YOLO loss**, operating on leaky surfaces rather than RGB images [1805.07931]. The input surface size is **\(128 \times 128\)**. The field of view is divided into a grid, and each region predicts **\(B=2\)** bounding boxes together with class probabilities over **\(C\)** classes. For the MNIST-style detection setup, the paper uses a **\(4 \times 4\)** grid on a \(128 \times 128\) surface; for **N-Caltech101**, it uses a **\(5 \times 7\)** grid.

The network predicts the standard YOLO-style outputs: objectness, bounding-box coordinates, bounding-box size, and class probabilities. The hidden layers use **Leaky ReLU**, while the output layer uses **linear activation**. The exact architecture depends on the dataset. For MNIST-based experiments, the design is inspired by **LeNet** and uses **6 convolution/pooling layers for feature extraction**. For **N-Caltech101**, the network is inspired by **VGG16**, but simplified to one layer per convolutional block [1805.07931].

The paper is explicit that YOLE is **not** the full original YOLO network architecture. In this context, “YOLO” primarily denotes the **training and detection formulation**, not a verbatim reuse of the original layer-by-layer design. That distinction is important for interpreting the model: YOLE inherits the one-stage regression-style detection philosophy, while adapting the backbone to the constraints of event-derived inputs.

YOLE is trained with the standard **multi-objective YOLO loss**, covering localization error, confidence or objectness error, classification error, and no-object penalty. For some experiments, better performance was obtained with modified hyperparameters:

- \(\lambda_{coord} = 25.0\)
- \(\lambda_{noobj} = 0.25\)

[1805.07931]

## 4. Training protocol and dataset construction

The common training setup uses **Adam** with learning rate \(10^{-4}\), \(\beta_1 = 0.9\), \(\beta_2 = 0.999\), and \(\epsilon = 10^{-8}\) [1805.07931]. The first four convolutional layers are initialized from a **recognition network pretrained to classify the target objects**, while the remaining layers use **Glorot initialization**. **Early stopping** is applied using validation sets of the same size as the test sets.

Batch sizes are dataset-specific and chosen to fit GPU memory:

- **Shifted N-MNIST:** **10**
- **Shifted MNIST-DVS:** **40**
- **N-Caltech101:** **40**
- **Blackboard MNIST:** **25**
- **OD-Poker-DVS:** **35**

A substantial part of the work consists of extending or creating event-based object detection datasets [1805.07931]. The datasets used are **N-Caltech101**, **Shifted N-MNIST**, **Shifted MNIST-DVS**, **OD-Poker-DVS**, and **Blackboard MNIST**.

**Shifted N-MNIST** is built from N-MNIST by placing one or two digits in non-overlapping positions in a larger field of view. Bounding boxes are extracted by integrating events into a frame, removing noise, clustering with DBSCAN-like grouping, and shifting boxes to the final digit position. The supplementary material specifies noise filtering with \(\rho = 3\), radius \(R = 2\), and minimum area threshold \(\mathrm{min}_{area} = 10\).

**Shifted MNIST-DVS** follows a similar procedure using MNIST-DVS samples at multiple scales: **scale4**, **scale8**, and **scale16**.

**OD-Poker-DVS** extends Poker-DVS to object detection by using tracking to annotate pips with bounding boxes in the original recordings; samples are split into short time windows of approximately **1.5 ms**.

**Blackboard MNIST** is a synthetic dataset generated with the **DAVIS simulator**, using chalk-like white digits on a blackboard with random positions, scales, and camera trajectories. It is organized into **EASY**, **MEDIUM**, and **HARD** levels, and introduces varying scale, multiple objects, motion-induced event bursts, and partial visibility filtering [1805.07931].

## 5. Empirical performance and comparison with fcYOLE

The paper evaluates YOLE using **accuracy**, **mean Average Precision (mAP)**, and **per-class Average Precision (AP)** on N-Caltech101. Accuracy is computed by matching ground-truth boxes to predicted boxes with highest **IoU** [1805.07931].

The principal quantitative comparison is between YOLE and **fcYOLE**, the paper’s asynchronous fully convolutional extension.

| Dataset | fcYOLE | YOLE |
|---|---:|---:|
| Shifted MNIST-DVS | 94.0 acc / 87.4 mAP | 96.1 acc / 92.0 mAP |
| Blackboard MNIST | 88.5 acc / 84.7 mAP | 90.4 acc / 87.4 mAP |
| OD-Poker-DVS | 79.10 acc / 78.69 mAP | 87.3 acc / 82.2 mAP |
| N-Caltech101 | 57.1 acc / 26.9 mAP | 64.9 acc / 39.8 mAP |

YOLE outperforms fcYOLE on all reported datasets [1805.07931]. The paper attributes this difference not to an intrinsic weakness of event-based layers, but to the fact that fcYOLE removes fully connected layers and therefore has less expressive power. That is a notable interpretive point: the comparison does not imply that asynchronous computation is inherently inferior, only that the particular fully convolutional simplification imposed a representational cost.

The paper also studies robustness on **Shifted N-MNIST** under progressively harder variants.

| Variant | Accuracy | mAP |
|---|---:|---:|
| v1 | 94.9 | 91.3 |
| v2 | 91.7 | 87.9 |
| v2* | 94.7 | 90.5 |
| v2fr | 88.6 | 81.5 |
| v2fr+ns | 85.5 | 77.4 |

The progression shows that YOLE can detect multiple objects, but performance declines with clutter and noise. The paper links part of the recovered performance to the modified YOLO loss weights \(\lambda_{coord} = 25.0\) and \(\lambda_{noobj} = 0.25\) [1805.07931].

For **N-Caltech101**, the paper reports that YOLE achieves much better AP than fcYOLE on many classes, especially those with many training examples and less intra-class variability. Performance remains weak for classes with few training samples, high variability, and class imbalance. The paper nevertheless notes strong AP on categories such as **motorbikes**, **airplanes**, **faces_easy**, **watch**, **minaret**, and **umbrellas** [1805.07931].

## 6. Relation to asynchronous detection, limitations, and significance

YOLE is best understood in relation to **fcYOLE**. Whereas YOLE converts events into a single leaky surface and then applies a conventional dense CNN, fcYOLE is a **fully convolutional, asynchronous event-based detector** that introduces **e-conv**, **e-max-pool**, layer-wise internal states, update matrices \(\mathit{F}_{(n)}\), and positional indices \(\mathit{I}_{(n)}\) for pooling [1805.07931]. In conceptual terms, the distinction is straightforward:

- **YOLE:** event integration first, CNN second
- **fcYOLE:** event-driven computation throughout the network

This difference also appears in timing behavior. Events were grouped into **10 ms batches**, and timings were averaged over **1000 runs**. On **Shifted N-MNIST**, the event-based approach achieves **22.6 ms per batch** and about **2× speedup**. On **Blackboard MNIST**, the event-based approach takes **43.2 ms per batch**, while the conventional network takes **34.6 ms per batch** [1805.07931]. The authors therefore conclude that asynchronous CNNs are advantageous when changes are localized and sparse, but can be slower when changes are widespread and noisy. They further note that their asynchronous implementation is beneficial only up to about **80% event sparsity**.

Within that comparison, YOLE’s advantages are explicit: it is a **simple and practical baseline for event-camera object detection**, uses **standard CNN training**, works well on several datasets, preserves some temporal information via leaky integration, and achieves better performance than fcYOLE in the reported experiments [1805.07931]. Its limitations are equally clear: it does **not** exploit event sparsity directly, recomputes a conventional CNN on a dense surface, can degrade significantly under heavy noise or clutter, and depends on the quality of leaky-surface parameters.

The broader significance of YOLE lies in its bridging role. It demonstrates that standard object detection can be adapted to neuromorphic cameras through temporally decaying event integration, thereby linking the mature ecosystem of frame-based deep detection with the emerging domain of event-based sensing [1805.07931]. A plausible implication is that YOLE’s enduring value is methodological as much as empirical: it establishes a strong baseline against which truly asynchronous detectors can be judged, while showing that event-camera object detection need not begin with a fully event-native learning stack.

Source: https://www.emergentmind.com/topics/yole