fcYOLE: Asynchronous Event Detector
- fcYOLE is an asynchronous event-based object detector that leverages leaky surface representations and e-conv layers to process sparse event data.
- The model reformulates YOLO-style detection by replacing fully connected layers with 1x1 convolutions, updating only regions affected by new events.
- Through analytic propagation of decay and sparse updates, fcYOLE reduces redundant computation, though it may sacrifice global context needed for high accuracy.
fcYOLE is a fully convolutional, asynchronous event-driven object detector for event-based or neuromorphic cameras, introduced in "Asynchronous Convolutional Networks for Object Detection in Neuromorphic Cameras" (Cannici et al., 2018). It addresses object detection from event-camera data, where the sensor emits events rather than frames, and each event is represented as , with pixel location , timestamp , and polarity of a brightness change. Within the paper, fcYOLE is defined relative to a simpler baseline, YOLE, which integrates events into leaky surfaces and then applies a conventional frame-style detector. fcYOLE retains the same detection objective and training philosophy as YOLO and YOLE, but reformulates inference with event-based convolutional and max-pooling layers so that only output locations affected by incoming events or activation-regime changes are recomputed, while unaffected regions are updated analytically through a leak term.
1. Position within event-based object detection
fcYOLE is presented for the problem setting of object detection from event-camera data. Unlike conventional RGB cameras, an event camera does not emit images at a fixed rate. Instead, it reports sparse spatiotemporal changes, and the detector must therefore operate on sparse changes rather than on dense frames (Cannici et al., 2018).
The paper distinguishes fcYOLE from both YOLE and ordinary YOLO-style detectors. YOLE, "You Only Look at Events," is the baseline that converts event streams into leaky surfaces and processes those surfaces synchronously with a standard CNN detector. It is event-based only at the level of input representation. It does not exploit event sparsity inside the network, because each surface is treated like a normal image.
fcYOLE differs in two explicit ways. First, it is fully convolutional: fully connected detection layers are removed and replaced by convolutional layers. Second, it is asynchronous at inference time: each layer maintains internal state from the previous time step and recomputes only those output locations whose receptive fields are affected by new events or by changes in activation regime. Conventional frame-based YOLO-style detectors, by contrast, assume dense frames and recompute every feature map at every frame time, even when only a small part of the scene has changed.
A common misconception is to treat fcYOLE as merely YOLO applied to event frames. The paper explicitly frames it differently. fcYOLE starts from the same leaky surface representation as YOLE, but its forward pass is reformulated to exploit the asynchronous, sparse, and localized character of event streams. In that sense, it is not just a frame-style detector operating on event-derived images.
2. Input representation and leaky event surfaces
fcYOLE does not operate on raw unprocessed events directly in the sense of a spiking network. It begins from a leaky surface representation, shared with YOLE, but embeds that representation as the first layer of an asynchronous network (Cannici et al., 2018).
For each incoming event at time , all pixels are decayed and the event pixel is incremented:
where is the current surface value, , 0 is the leak rate, and 1 is the event increment. The paper fixes 2 and varies 3. For readability, it defines
4
This representation is central because it yields a simple linear dependence between consecutive surfaces. The paper’s motivation is that this linear relation can be propagated analytically through asynchronous layers.
The sparsity exploited by fcYOLE has two forms. Input sparsity refers to the fact that only a small set of pixels receives events at each update. Propagation sparsity refers to the fact that only output locations whose receptive fields are influenced by those changed pixels, or whose activation slope changes, must be locally recomputed. The leaky surface layer also forwards auxiliary information downstream: the list of incoming events, 5, and the list of surface pixels reset to zero by the 6 operation. This matters because downstream layers must track not only where the signal changed, but also where the truncation altered the update dynamics.
3. Architecture and detection parameterization
The paper describes fcYOLE as a fully convolutional conversion of YOLE. The network is trained first with standard layers, and then, at inference time, standard convolution and max-pooling layers are replaced by event-based versions, e-conv and e-max-pool, using the same learned weights (Cannici et al., 2018).
For MNIST-style datasets, YOLE and fcYOLE use an architecture inspired by LeNet, with 6 conv-pool layers for feature extraction, and differ only in the final regression and classification part. For N-Caltech101, they use a different, slightly VGG-inspired architecture with one conv layer per conv group; fcYOLE remains the fully convolutional event-based variant.
The structural facts given for fcYOLE are explicit:
- input is a leaky surface, typically 7;
- all standard convolution layers are replaced by e-conv;
- all standard max-pooling layers are replaced by e-max-pool;
- all fully connected layers are replaced by 8 e-conv layers;
- hidden activations are Leaky ReLU;
- the last layer uses linear activation;
- training uses the YOLO multi-objective loss from Redmon et al.
For the main detector described in Section 3.4 of the paper, the field of view is divided into a 9 grid for 0 input, each grid cell predicts 1 bounding boxes, and objects belong to 2 classes. The last 3 e-conv layer maps feature vectors to the detection parameters. The paper also states that the last layer maps feature vectors into a set of 20 values defining the predicted bounding boxes. For the MNIST-like case, this is consistent with 4 and 5, so that each cell outputs
6
values: for each of the 7 boxes, 8, plus 9 class scores. The resulting detection tensor over the 0 grid therefore has shape 1. For N-Caltech101, the surface is divided into a 2 grid with 3 boxes per cell.
The paper does not provide a complete layer-by-layer table of kernel counts and feature-map dimensions for every stage in the main text, so a complete exact reconstruction is not possible from the provided material. What is explicit is the intended decomposition into a leaky surface layer, a feature extractor made of alternating e-conv and e-max-pool layers with Leaky ReLU after hidden convolutions, and a detection head composed of 4 e-conv layers.
Because fcYOLE is fully convolutional, the same subnetwork can be tiled across larger surfaces. The paper explicitly notes that subnetworks processing 5 regions are parameter-shared and can be stacked to process larger surfaces without redesign or retraining. A plausible implication is that the model trades some global contextual capacity for locality and parameter sharing.
4. Asynchronous computation and event-based layers
The core technical contribution of fcYOLE is the event-driven reformulation of convolution and max-pooling (Cannici et al., 2018). Each layer maintains internal state consisting of the previous output feature map 6, an update matrix 7, and, for pooling, an index matrix 8.
For the first convolutional layer, ordinary computation is written as
9
When a new event arrives, pixels not directly hit by the event decay as
0
The first-layer output then becomes
1
The derivation assumes piecewise linear hidden activations, such as ReLU or Leaky ReLU,
2
and assumes that the update does not cross into another linear segment. Under that condition,
3
The interpretation given in the paper is that, if nothing local happened except global decay and the activation remains on the same linear branch, the output can be updated by an additive correction rather than by recomputing a full convolution.
For deeper layers, the same principle yields the recursive update
4
For the first layer, incorporating the surface truncation as a ReLU gives
5
where
6
In the paper’s interpretation, 7 captures how the uniform leak in the input propagates through the network at each spatial location, taking current activation slopes into account. If nothing in a region receives an event and no activation branch changes, then the local feature can be updated by
8
The e-conv algorithm stores the previous feature map and update matrix. Initialization is done by a full forward pass on a blank surface; this is the only time the entire network is computed densely. On each new event batch, the layer locally updates 9, analytically updates unaffected output positions, recomputes 0 locally where receptive fields are affected by incoming events, and forwards updated features and generated events to the next layer. The paper thus distinguishes exact local recomputation where needed from cheap state updates elsewhere.
Event-based max pooling follows the same stateful logic. Each e-max-pool layer stores a positional matrix 1 containing the position of the max value in each receptive field. On event arrival, receptive fields affected by events have their maxima recomputed; the output feature map is then built by reading the input at the positions indicated by 2, and the update matrix 3 is fetched from the previous layer at those max locations. The paper also highlights a subtlety: even if no event occurs in a receptive field, different inputs may decay at different rates because their 4 values differ, so the identity of the maximum can change due purely to leak. To avoid unnecessary recomputation, a sufficient condition is used: if an input feature 5 is both the current maximum in receptive field 6 and the one with the minimum update rate 7 in 8, then the pooled output remains maximum under decay and its index need not be recomputed until a new event enters 9.
These derivations rely on piecewise linear activations, unchanged linear-segment membership during the cheap update, local support of convolution kernels, and sparse input changes. When an activation crosses into a different linear branch, the corresponding 0 entries must be updated locally and an event is emitted to downstream layers.
5. Training protocol, datasets, and empirical results
The training configuration reported for fcYOLE uses the YOLO multi-objective loss, the Adam optimizer, learning rate 1, 2, 3, 4, early stopping, and dataset-dependent batch sizes (Cannici et al., 2018). The first four conv layers are initialized from a pretrained recognition network; later layers use Glorot initialization. The paper also reports tuning two YOLO hyperparameters, 5 and 6, which improved some results over the original YOLO defaults.
Evaluation is conducted on several event-based detection datasets. Shifted N-MNIST is an extension of N-MNIST for detection, with 7 training and 8 testing samples; variants include v1, v2, v2fr, and v2fr+ns. Shifted MNIST-DVS is a detection version of MNIST-DVS with mixed scales and random placement in a 9 field of view, with 0 samples. OD-Poker-DVS is a detection extension of Poker-DVS with 1 training and 2 testing samples. Blackboard MNIST is a synthetic dataset created with the DAVIS simulator, reported in the main paper as 3 training and 4 testing samples, while the supplementary gives a full combined collection of 5 total split into training, testing, and validation. N-Caltech101 uses a stratified 80/20 train/test split, and the background class is excluded because it lacks box annotations. Metrics are accuracy, computed by matching each ground-truth box with the predicted box of highest IoU, and mAP.
The main quantitative comparison is summarized below.
| Dataset | fcYOLE | YOLE |
|---|---|---|
| Shifted MNIST-DVS | accuracy 6, mAP 7 | accuracy 8, mAP 9 |
| Blackboard MNIST | accuracy 0, mAP 1 | accuracy 2, mAP 3 |
| OD-Poker-DVS | accuracy 4, mAP 5 | accuracy 6, mAP 7 |
| N-Caltech101 | accuracy 8, mAP 9 | accuracy 0, mAP 1 |
Across these datasets, fcYOLE consistently trails YOLE in accuracy and mAP. The paper explicitly states that this drop is not attributed to approximation in the event-based layers; the claim is that the event layers produce the same outputs as conventional layers. Instead, the reduction is attributed to the fully convolutional design, which replaces fully connected layers and thereby reduces expressive power. The paper explains the weakness as a locality issue: each fcYOLE region predicts using only information from its own local portion of the field of view. If an object straddles a region boundary, the network may have to infer size or class from partial evidence, whereas YOLE’s FC layers can exploit more global context.
For N-Caltech101, the classwise AP tables show that fcYOLE performs best on categories with many training samples and/or low intraclass variability, including Motorbikes 2, airplanes 3, Faces_easy 4, watch 5, dollar_bill 6, car_side 7, grand_piano 8, and menorah 9. The paper states that performance degrades strongly on rare or high-variability classes, with many classes near zero.
6. Computational profile, advantages, and limitations
The computational rationale for fcYOLE is the reduction of redundant computation. A synchronous frame-style detector recomputes every feature map for every integrated frame, even when only a few pixels changed. fcYOLE instead updates only changed receptive fields locally, uses stored state and update matrices elsewhere, produces outputs only as a consequence of incoming events, and can update under pure time decay without explicit new frames (Cannici et al., 2018).
The paper reports runtime measurements on Shifted N-MNIST and Blackboard MNIST, grouping events into batches of 10 ms and averaging over 1000 runs. On Shifted N-MNIST, the event-based approach achieved a 00 speedup, with 22.6 ms per batch. On Blackboard MNIST, fcYOLE was slower, at 43.2 ms per batch, than the conventional network at 34.6 ms per batch. The explanation given is that Blackboard MNIST contains more spatially widespread changes, so the sparse-update assumption is less favorable. The paper further states that the implementation is not optimized for noisy scenes and that asynchronous CNNs were faster only up to about 80% event sparsity, where sparsity denotes the percentage of changed pixels in the reconstructed image.
This leads to a conditional interpretation of the model’s utility. The paper indicates that fcYOLE is preferable when the sensor is an event camera, scene changes are sparse and localized, low-latency asynchronous inference is desired, computation should scale with event activity rather than frame rate, and deployment benefits from fewer parameters and a fully convolutional architecture. By contrast, YOLE is preferable when maximum detection accuracy matters more than asynchronous efficiency, objects may span multiple regions and require global context, scene changes are broad or noisy so sparse updates do not save compute, or synchronous processing of integrated event surfaces is acceptable.
The main innovations identified in the paper are asynchronous convolutional inference for event-based detection, the formalization of event-based convolution with state 01 and update matrix 02, the formalization of event-based max-pooling with positional state 03, the conversion of a YOLO-like detector into a fully convolutional event-driven network, and the ability to run a CNN trained on integrated event surfaces in an event-driven manner without changing learned weights by replacing standard layers with asynchronous equivalents at inference. The limitations identified are reduced expressive power due to removal of FC layers, speed benefits only when changes are sparse, absence of event-based FC layers in the current formulation, lack of optimization for all sparsity regimes, and the fact that the method is realized in general-purpose hardware and software rather than in an ad hoc hardware implementation. The authors explicitly suggest that dedicated hardware could better realize the method’s potential and enable fairer comparison with hardware-accelerated SNNs.
Taken together, these properties define fcYOLE as a detector whose central contribution is not only event-based input handling but asynchronous CNN inference itself. Mathematically, the model is organized around the recursively defined update matrix 04 and the use of piecewise linear activations to propagate leak-driven state changes without dense recomputation. Empirically, it demonstrates that this formulation can reduce computation when event activity is sufficiently sparse, even though the fully convolutional architecture generally yields lower detection accuracy than YOLE.