Papers
Topics
Authors
Recent
Search
2000 character limit reached

YOLE: Neuromorphic Object Detection Baseline

Updated 13 July 2026
  • YOLE is a neuromorphic object detector that converts asynchronous events into a temporally decaying leaky surface for CNN-based detection.
  • The model employs a standard YOLO loss and architecture adapted to event-based inputs, striking a balance between frame-based and asynchronous processing.
  • Empirical results show YOLE outperforms fcYOLE across multiple datasets, establishing it as a robust baseline for event-based object detection.

Searching arXiv for the primary YOLE paper and nearby event-based object detection work to ground the article. YOLE, short for “You Only Look at Events,” is a neuromorphic object detector that adapts the YOLO detection formulation to event-based cameras by operating on an integrated leaky surface rather than on conventional image frames. It was introduced as the baseline detector in “Asynchronous Convolutional Networks for Object Detection in Neuromorphic Cameras” (Cannici et al., 2018), where it serves both as a practical event-camera detection model and as the reference point for comparison with the paper’s fully asynchronous extension, fcYOLE. In that formulation, YOLE is explicitly characterized as YOLO + leaky surface: a frame-based convolutional detector applied to temporally decayed event accumulations, rather than a network that propagates raw asynchronous events through event-native layers.

1. Event-based sensing and the rationale for YOLE

Event cameras, also called neuromorphic cameras, do not emit conventional frames at a fixed rate. Instead, they output an event only when a pixel undergoes a brightness change, represented as

e=⟨x,y,ts,p⟩\mathbf{e} = \langle x, y, ts, p \rangle

where xx and yy are pixel coordinates, tsts is the timestamp, and p∈{1,−1}p \in \{1,-1\} is the polarity of the brightness change (Cannici et al., 2018). The properties emphasized for these sensors are microsecond temporal resolution, low power consumption, low bandwidth, high temporal sparsity, and reduced data redundancy.

These sensing characteristics make event cameras attractive for tasks such as tracking, optical flow, and odometry, but object detection is more demanding because it requires both category assignment and localization. YOLE addresses this setting by translating asynchronous event streams into a representation that can be processed by standard CNN machinery. The central motivation is practical: standard CNNs can be trained efficiently, whereas deep SNNs remain difficult to train for complex tasks; at the same time, a temporally decaying event surface preserves temporal ordering better than naive frame binning and provides a simple baseline for comparing conventional frame-style processing with explicitly asynchronous architectures (Cannici et al., 2018).

A common misconception is to treat YOLE as an event-native asynchronous detector. It is not. The essential design decision is event integration first, CNN second. This distinguishes it from methods that attempt to exploit sparsity throughout the network.

2. Leaky-surface input representation

YOLE consumes a leaky surface, a dense 2D array that is updated as events arrive. For an event at time tstts^t, the surface update is

qxs,yst=max⁡(pxs,yst−1−λ⋅Δts,0)q_{x_s, y_s}^t = \max\left(p_{x_s, y_s}^{t-1} - \lambda \cdot \Delta_{ts}, 0\right)

pxs,yst={qxs,yst+Δincrif (xs,ys)t=(xe,ye)t qxs,ystotherwisep_{x_s, y_s}^t = \begin{cases} q_{x_s, y_s}^t + \Delta_{incr} & \text{if } (x_s, y_s)^t = (x_e, y_e)^t \ q_{x_s, y_s}^t & \text{otherwise} \end{cases}

with Δts=tst−tst−1\Delta_{ts} = ts^t - ts^{t-1}, leak rate λ\lambda, increment xx0, and xx1 enforcing non-negativity (Cannici et al., 2018). The paper fixes xx2 and varies only xx3 depending on the dataset.

Operationally, every event adds activation at its pixel while the full surface decays over time. This yields a frame-like representation that still carries temporal dynamics. The stated reasons for preferring the leaky surface over simple binary frame accumulation are that it avoids treating all events equally, maintains a notion of temporal decay, and handles noise more smoothly (Cannici et al., 2018).

This suggests that YOLE’s representation is best understood as a compromise between two incompatible desiderata: strict event-driven computation and compatibility with dense CNN training. It is temporally informed, but not event-native in the strict asynchronous sense.

3. Detection architecture and YOLO formulation

YOLE is a standard CNN detector trained with the YOLO loss, operating on leaky surfaces rather than RGB images (Cannici et al., 2018). The input surface size is xx4. The field of view is divided into a grid, and each region predicts xx5 bounding boxes together with class probabilities over xx6 classes. For the MNIST-style detection setup, the paper uses a xx7 grid on a xx8 surface; for N-Caltech101, it uses a xx9 grid.

The network predicts the standard YOLO-style outputs: objectness, bounding-box coordinates, bounding-box size, and class probabilities. The hidden layers use Leaky ReLU, while the output layer uses linear activation. The exact architecture depends on the dataset. For MNIST-based experiments, the design is inspired by LeNet and uses 6 convolution/pooling layers for feature extraction. For N-Caltech101, the network is inspired by VGG16, but simplified to one layer per convolutional block (Cannici et al., 2018).

The paper is explicit that YOLE is not the full original YOLO network architecture. In this context, “YOLO” primarily denotes the training and detection formulation, not a verbatim reuse of the original layer-by-layer design. That distinction is important for interpreting the model: YOLE inherits the one-stage regression-style detection philosophy, while adapting the backbone to the constraints of event-derived inputs.

YOLE is trained with the standard multi-objective YOLO loss, covering localization error, confidence or objectness error, classification error, and no-object penalty. For some experiments, better performance was obtained with modified hyperparameters:

  • yy0
  • yy1

(Cannici et al., 2018)

4. Training protocol and dataset construction

The common training setup uses Adam with learning rate yy2, yy3, yy4, and yy5 (Cannici et al., 2018). The first four convolutional layers are initialized from a recognition network pretrained to classify the target objects, while the remaining layers use Glorot initialization. Early stopping is applied using validation sets of the same size as the test sets.

Batch sizes are dataset-specific and chosen to fit GPU memory:

  • Shifted N-MNIST: 10
  • Shifted MNIST-DVS: 40
  • N-Caltech101: 40
  • Blackboard MNIST: 25
  • OD-Poker-DVS: 35

A substantial part of the work consists of extending or creating event-based object detection datasets (Cannici et al., 2018). The datasets used are N-Caltech101, Shifted N-MNIST, Shifted MNIST-DVS, OD-Poker-DVS, and Blackboard MNIST.

Shifted N-MNIST is built from N-MNIST by placing one or two digits in non-overlapping positions in a larger field of view. Bounding boxes are extracted by integrating events into a frame, removing noise, clustering with DBSCAN-like grouping, and shifting boxes to the final digit position. The supplementary material specifies noise filtering with yy6, radius yy7, and minimum area threshold yy8.

Shifted MNIST-DVS follows a similar procedure using MNIST-DVS samples at multiple scales: scale4, scale8, and scale16.

OD-Poker-DVS extends Poker-DVS to object detection by using tracking to annotate pips with bounding boxes in the original recordings; samples are split into short time windows of approximately 1.5 ms.

Blackboard MNIST is a synthetic dataset generated with the DAVIS simulator, using chalk-like white digits on a blackboard with random positions, scales, and camera trajectories. It is organized into EASY, MEDIUM, and HARD levels, and introduces varying scale, multiple objects, motion-induced event bursts, and partial visibility filtering (Cannici et al., 2018).

5. Empirical performance and comparison with fcYOLE

The paper evaluates YOLE using accuracy, mean Average Precision (mAP), and per-class Average Precision (AP) on N-Caltech101. Accuracy is computed by matching ground-truth boxes to predicted boxes with highest IoU (Cannici et al., 2018).

The principal quantitative comparison is between YOLE and fcYOLE, the paper’s asynchronous fully convolutional extension.

Dataset fcYOLE YOLE
Shifted MNIST-DVS 94.0 acc / 87.4 mAP 96.1 acc / 92.0 mAP
Blackboard MNIST 88.5 acc / 84.7 mAP 90.4 acc / 87.4 mAP
OD-Poker-DVS 79.10 acc / 78.69 mAP 87.3 acc / 82.2 mAP
N-Caltech101 57.1 acc / 26.9 mAP 64.9 acc / 39.8 mAP

YOLE outperforms fcYOLE on all reported datasets (Cannici et al., 2018). The paper attributes this difference not to an intrinsic weakness of event-based layers, but to the fact that fcYOLE removes fully connected layers and therefore has less expressive power. That is a notable interpretive point: the comparison does not imply that asynchronous computation is inherently inferior, only that the particular fully convolutional simplification imposed a representational cost.

The paper also studies robustness on Shifted N-MNIST under progressively harder variants.

Variant Accuracy mAP
v1 94.9 91.3
v2 91.7 87.9
v2* 94.7 90.5
v2fr 88.6 81.5
v2fr+ns 85.5 77.4

The progression shows that YOLE can detect multiple objects, but performance declines with clutter and noise. The paper links part of the recovered performance to the modified YOLO loss weights yy9 and tsts0 (Cannici et al., 2018).

For N-Caltech101, the paper reports that YOLE achieves much better AP than fcYOLE on many classes, especially those with many training examples and less intra-class variability. Performance remains weak for classes with few training samples, high variability, and class imbalance. The paper nevertheless notes strong AP on categories such as motorbikes, airplanes, faces_easy, watch, minaret, and umbrellas (Cannici et al., 2018).

6. Relation to asynchronous detection, limitations, and significance

YOLE is best understood in relation to fcYOLE. Whereas YOLE converts events into a single leaky surface and then applies a conventional dense CNN, fcYOLE is a fully convolutional, asynchronous event-based detector that introduces e-conv, e-max-pool, layer-wise internal states, update matrices tsts1, and positional indices tsts2 for pooling (Cannici et al., 2018). In conceptual terms, the distinction is straightforward:

  • YOLE: event integration first, CNN second
  • fcYOLE: event-driven computation throughout the network

This difference also appears in timing behavior. Events were grouped into 10 ms batches, and timings were averaged over 1000 runs. On Shifted N-MNIST, the event-based approach achieves 22.6 ms per batch and about 2× speedup. On Blackboard MNIST, the event-based approach takes 43.2 ms per batch, while the conventional network takes 34.6 ms per batch (Cannici et al., 2018). The authors therefore conclude that asynchronous CNNs are advantageous when changes are localized and sparse, but can be slower when changes are widespread and noisy. They further note that their asynchronous implementation is beneficial only up to about 80% event sparsity.

Within that comparison, YOLE’s advantages are explicit: it is a simple and practical baseline for event-camera object detection, uses standard CNN training, works well on several datasets, preserves some temporal information via leaky integration, and achieves better performance than fcYOLE in the reported experiments (Cannici et al., 2018). Its limitations are equally clear: it does not exploit event sparsity directly, recomputes a conventional CNN on a dense surface, can degrade significantly under heavy noise or clutter, and depends on the quality of leaky-surface parameters.

The broader significance of YOLE lies in its bridging role. It demonstrates that standard object detection can be adapted to neuromorphic cameras through temporally decaying event integration, thereby linking the mature ecosystem of frame-based deep detection with the emerging domain of event-based sensing (Cannici et al., 2018). A plausible implication is that YOLE’s enduring value is methodological as much as empirical: it establishes a strong baseline against which truly asynchronous detectors can be judged, while showing that event-camera object detection need not begin with a fully event-native learning stack.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to YOLE.