ReYOLOv8: Advanced YOLOv8 Redesign
- ReYOLOv8 is a specialized variant of the YOLOv8 detector that integrates recurrent modeling and refined spatial attention for real-time perception in event-based and autonomous driving scenarios.
- It leverages efficient event encoding with VTEI and employs techniques like Random Polarity Suppression to boost mAP while reducing latency on benchmarks such as GEN1 and PEDRo.
- The design preserves the YOLOv8 detection head while reworking the backbone and neck, illustrating a flexible pattern that adapts to both recurrent and refined feature extraction needs.
ReYOLOv8 denotes a set of YOLOv8-derived redesigns in which the standard YOLOv8 detector is retained as a baseline but its feature extraction, temporal modeling, or task structure is reworked for a specialized real-time perception problem. In the most explicit usage, ReYOLOv8 is the recurrent event-based detector introduced for event-camera object detection, where YOLOv8 is augmented with ConvLSTM-based spatiotemporal modeling, Volume of Ternary Event Images (VTEI), and Random Polarity Suppression (Silva et al., 2024). In a separate autonomous-driving context, the same label is a useful shorthand for a refined YOLOv8 variant that replaces C2f with C2f_RFAConv, inserts Triplet Attention, and adds a P2 detection head for small objects (Ling et al., 2024). This suggests that ReYOLOv8 is not a single canonical release, but a broader pattern of re-architected YOLOv8 systems specialized for latency-sensitive vision.
1. Terminological scope and relation to standard YOLOv8
YOLOv8, as analyzed in the recent literature, is an anchor-free, single-stage detector built around a CSPNet-based backbone, an FPN+PAN neck, and a dense prediction head that regresses boxes and class scores across multiple feature scales (Yaseen, 2024). The standard pipeline is image input backbone neck head, with multi-scale feature maps typically denoted and their fused counterparts (Yaseen, 2024). In the mainstream Ultralytics formulation assumed by the autonomous-driving refinement paper, the backbone uses stem convolutions and multiple C2f stages, the neck is FPN/PAN-like, and the head is decoupled and predicts class probability, bounding box regression using Distribution Focal Loss, and objectness or quality (Ling et al., 2024).
Within that baseline, ReYOLOv8-style work preserves the YOLOv8 detection head and training losses but relocates innovation into the backbone and neck. In the event-based formulation, the essential redesign is temporal: frame-based feature extraction is retained structurally, but recurrent blocks are inserted so that prediction depends on event history rather than on a single encoded window (Silva et al., 2024). In the autonomous-driving refinement, the essential redesign is spatial: conventional C2f internals are replaced by receptive-field-aware convolution, Triplet Attention is injected, and an additional P2 head is introduced to strengthen high-resolution detection (Ling et al., 2024).
A common misconception is that ReYOLOv8 necessarily denotes recurrence. The published record does not support that narrower usage. One paper uses ReYOLOv8 to name a recurrent event-based detector, whereas another supports a refined or redesigned YOLOv8 interpretation without recurrent computation (Silva et al., 2024). A plausible implication is that the prefix “Re” functions as “recurrent” in one line of work and “refined” or “redesigned” in another.
2. ReYOLOv8 as a recurrent event-based detector
The event-based ReYOLOv8 framework addresses a sensing regime in which the input is not an RGB frame but an asynchronous stream of events , emitted when the log-brightness change exceeds a threshold,
This sensing model is motivated by motion blur, limited dynamic range, and frame-rate latency in standard RGB cameras, particularly in autonomous vehicles and robotics (Silva et al., 2024).
Architecturally, the detector keeps the overall YOLOv8 decomposition into backbone, neck, and detection head, but replaces frame input with an event tensor and inserts recurrent blocks after C2f feature blocks. The pipeline is event stream VTEI modified YOLOv8 backbone YOLOv8-style PANet neck 0 three-scale detection head (Silva et al., 2024). For the small model, the backbone begins with a 1-channel VTEI tensor, uses alternating Conv2D, C2f, and ConvLSTM stages, and ends in SPPF; the neck upsamples and concatenates deeper recurrent features; the detector predicts on feature levels with channels 2 (Silva et al., 2024).
The recurrent component is a ConvLSTM inserted after C2f blocks at multiple resolutions. Given current feature map 3 and previous hidden and cell states 4, the gates are computed by 5 convolutions on 6, followed by
7
The hidden state 8 becomes the recurrently refined feature map passed forward in the backbone (Silva et al., 2024). Because recurrence is applied at several scales, temporal context is accumulated in both high-resolution and low-resolution representations.
The model family is defined at nano, small, and medium scales. ReYOLOv8n, ReYOLOv8s, and ReYOLOv8m contain 9M, 0M, and 1M parameters, respectively (Silva et al., 2024). The same paper reports that, relative to similarly scaled event-based baselines, the models improve mean Average Precision on the GEN1 dataset by 2, 3, and 4 across nano, small, and medium scales, respectively, while reducing the number of trainable parameters by an average of 5 and maintaining real-time processing speeds between 6 ms and 7 ms (Silva et al., 2024).
3. Event encoding, augmentation, and training protocol
A defining component of recurrent ReYOLOv8 is VTEI, a dense event representation designed for low latency and low memory. For a time window 8 and 9 temporal bins, each event timestamp is normalized to a bin index,
0
and a tensor 1 is filled so that each voxel stores the polarity of the last event observed at that pixel and bin (Silva et al., 2024). The resulting values are ternary: 2, 3, or 4. The encoding is intentionally compact, because only the last polarity is stored per pixel-bin.
The paper compares VTEI with MDES, SHist, and voxel grids. For GEN1 with 5, 6, and 7, VTEI uses 8 bytes per non-zero entry in COO form, and at a 9 ms window with an average of about 0k events it achieves 1 ms encoding latency and 2 Mev/s event processing rate (Silva et al., 2024). The same analysis reports compression ratios of 3–4, with better latency than SHist, MDES, and voxel grids, and lower bandwidth than SHist and voxel grids (Silva et al., 2024). The detector’s temporal modeling therefore rests on an encoding that is efficient enough not to erase the latency advantage of event sensing.
Training uses truncated backpropagation through time. Sequence length is 5 for GEN1 and 6 for PEDRo; hidden and cell states are reset between clips during training and at sequence boundaries during validation and test (Silva et al., 2024). The optimizer is SGD with momentum 7; weight decay is 8 for GEN1 and 9 for PEDRo; there is a 0-epoch warm-up with momentum 1 and bias learning rate 2; and the image sizes are 3 for GEN1 and 4 for PEDRo (Silva et al., 2024). Losses remain YOLOv8 defaults with box loss 5, class loss 6, and DFL 7 (Silva et al., 2024).
The event-specific augmentation is Random Polarity Suppression. Given a VTEI tensor 8, the transformation either leaves 9 unchanged or removes one polarity across all temporal bins: 0 Here 1 is the suppression probability and 2 the polarity-selection probability (Silva et al., 2024). The reported sweeps show that small suppression probabilities, typically 3, are preferable, and that balanced suppression with 4 is usually best (Silva et al., 2024). The gains are modest on GEN1 but substantially larger on PEDRo, where ReYOLOv8n improves from 5 to 6 mAP and ReYOLOv8m from 7 to 8 mAP under the selected settings (Silva et al., 2024).
4. Empirical performance on GEN1 and PEDRo
The event-based ReYOLOv8 paper evaluates on two benchmarks with distinct operating regimes: Prophesee GEN1 for automotive perception and PEDRo for robotics (Silva et al., 2024). GEN1 contains urban driving scenes with cars and pedestrians at 9 resolution and about 0k bounding boxes; PEDRo contains DAVIS346 handheld robotics recordings at 1 resolution and about 2k pedestrian boxes (Silva et al., 2024). The primary metric is COCO-style mAP@[0.5:0.95].
On PEDRo, ReYOLOv8n, ReYOLOv8s, and ReYOLOv8m obtain 3, 4, and 5 mAP@[0.5:0.95], with runtimes of 6 ms, 7 ms, and 8 ms, respectively (Silva et al., 2024). The YOLOv8x baseline adapted to events achieves 9 mAP@[0.5:0.95] with 0M parameters and 1 ms runtime, whereas ReYOLOv8m reaches higher accuracy with 2M parameters and ReYOLOv8n is about 3 smaller (Silva et al., 2024). The paper summarizes this as a 4 to 5 mAP improvement over the YOLOv8x-based baseline, with 6 and 7 smaller models and an average speed enhancement of 8 on PEDRo (Silva et al., 2024).
On GEN1, the reported gains are scale-matched against event-based competitors. At nano scale, ReYOLOv8n reaches 9 mAP with 0M parameters and 1 ms runtime, compared with RVT-T at 2 mAP and 3M parameters (Silva et al., 2024). At small scale, ReYOLOv8s reaches 4 mAP with 5M parameters; the paper highlights a 6 mAP gain over HMNet-L1 while using about 7 fewer parameters (Silva et al., 2024). At medium scale, ReYOLOv8m reaches 8 mAP with 9M parameters, exceeding SAST-CB at 00 and DTSDNet-M at 01 mAP (Silva et al., 2024).
The qualitative interpretation offered by the study is that recurrence stabilizes detection under sparse or noisy events, while VTEI keeps the end-to-end pipeline latency compatible with real-time use (Silva et al., 2024). Failure cases remain concentrated in extremely small distant pedestrians and cluttered backgrounds whose event patterns resemble noise (Silva et al., 2024). Those limitations are consistent with the fact that temporal integration can compensate for sparsity only up to the point where the object generates too little signal.
5. ReYOLOv8-style refinement for frame-based autonomous driving
A separate line of work applies the ReYOLOv8 idea to frame-based detection for autonomous driving rather than to event streams. In this formulation, the baseline is mainstream Ultralytics YOLOv8, but all positions using C2f are replaced internally by C2f_RFAConv blocks, Triplet Attention modules are inserted around key backbone and neck feature maps, and a new P2 detection head is added for small-object detection (Ling et al., 2024). The paper keeps the detection head and the training loss of YOLOv8 basically intact, so the redesign is concentrated in feature extraction (Ling et al., 2024).
The key operation is Receptive Field Attention Convolution. Instead of applying the same kernel 02 at every spatial position, the method introduces attention weights 03 so that the effective kernel becomes
04
The paper describes the resulting mechanism qualitatively as “Kernel1 = 05, Kernel2 = 06, Kernel3 = 07” and so on, so that each receptive field sees a different kernel adapted to local content (Ling et al., 2024). The C2f_RFAConv block preserves the global split–multi-block–concat–fuse layout of C2f while replacing the Bottleneck internals with RFAConv-based computation (Ling et al., 2024).
Triplet Attention is used in the form described in “Rotate to Attend: Convolutional Triplet Attention Module,” with three branches that model channel–height–width interactions through permutation, Z-pooling, 08 convolution, sigmoid gating, and aggregation (Ling et al., 2024). In the reported configuration, these modules are inserted into the improved YOLOv8 network structure after C2f_RFAConv is already in place (Ling et al., 2024). The additional P2 head extends the detection scale hierarchy upward to higher-resolution features, with the explicit aim of strengthening small-object detection (Ling et al., 2024).
The reported gains are numerical and stepwise. On the COCO dataset referred to in the paper, the original YOLOv8 baseline obtains [email protected] 09 and [email protected]:0.95 10 (Ling et al., 2024). After C2f_RFAConv only, the reported all-classes [email protected] is about 11 (Ling et al., 2024). After C2f_RFAConv, Triplet Attention, and P2 are combined, the final model reaches [email protected] 12 and [email protected]:0.95 13, corresponding to absolute gains of 14 and 15 over the baseline (Ling et al., 2024). The same source reports per-class PR improvements, for example pedestrian from 16 to 17 and van from 18 to 19 (Ling et al., 2024).
This version of ReYOLOv8 does not provide explicit FLOPs or FPS, but it states that RFAConv and Triplet Attention are designed to be lightweight and that the model continues to satisfy the real-time requirement for autonomous driving (Ling et al., 2024). The paper also notes that experiments are done on COCO and that generalization to autonomous-driving benchmarks such as KITTI, BDD100K, nuScenes, or Waymo is not explicitly shown (Ling et al., 2024).
6. Related redesign patterns, deployment logic, and limitations
The broader YOLOv8 literature shows that ReYOLOv8 sits within a wider pattern of task-aligned redesign. In distracted-driving classification, P-YOLOv8 adapts YOLOv8n-cls through transfer learning and reports 20 accuracy with 21 parameters and a 22 MB model size on the State Farm dataset (Elshamy et al., 2024). That model is not called ReYOLOv8, but it demonstrates the same design principle: begin from a small YOLOv8 backbone, specialize the head to the target label space, and exploit pretraining to preserve accuracy under a strict parameter budget (Elshamy et al., 2024). In autonomous-driving multi-task learning, A-YOLOM keeps a YOLOv8-style backbone and detection head, adds separate segmentation necks with an Adaptive Concatenation Module, and uses lightweight segmentation heads; the nano variant has 23M parameters and runs at 24 FPS while jointly performing detection, drivable-area segmentation, and lane-line segmentation (Wang et al., 2023). In UAV localization, YoloTag uses a lightweight YOLOv8 detector in a closed-loop system with EPnP and a Butterworth filter, and reports 25 FPS on a Quadro P2200 (Raxit et al., 2024). Collectively, these works place ReYOLOv8 in a larger engineering tradition of YOLOv8 specialization for edge or real-time systems.
Several design regularities recur across these derivatives. First, the YOLOv8 backbone–neck–head decomposition is typically preserved, even when individual blocks are replaced or auxiliary heads are added (Yaseen, 2024). Second, the smallest viable configuration is often preferred for embedded deployment, with task-specific heads or recurrent blocks added only where the target application justifies the overhead (Elshamy et al., 2024). Third, real-time use cases repeatedly push modifications into the feature extractor and leave the high-level training and deployment stack recognizable as YOLOv8 (Ling et al., 2024).
The main limitation of the term itself is semantic rather than algorithmic: there is no single standardized ReYOLOv8 architecture in the literature. One established usage denotes the recurrent event-based detector on GEN1 and PEDRo, while another denotes a refined frame-based detector built from C2f_RFAConv, Triplet Attention, and a P2 head (Silva et al., 2024). A second limitation is evidential scope. The event-based formulation is validated on two event datasets, and the frame-based autonomous-driving refinement is evaluated on COCO-style traffic classes rather than on a broad suite of driving benchmarks (Ling et al., 2024). The frame-based paper also does not report detailed FLOPs or FPS, and the event-based paper notes the latency cost of ConvLSTM relative to the fastest non-recurrent baselines (Silva et al., 2024).
A plausible implication is that ReYOLOv8 should be treated as a research pattern rather than a fixed model name. In that pattern, YOLOv8 supplies the scalable real-time scaffold, while domain-specific pressures—event sparsity, small-object traffic detection, multi-task scene understanding, or extreme parameter limits—determine whether the redesign is recurrent, attention-based, receptive-field-aware, or head-specialized.