---
title: 'ReYOLOv8: Advanced YOLOv8 Redesign'
url: https://www.emergentmind.com/topics/reyolov8
type: topic
---

# ReYOLOv8: Advanced YOLOv8 Redesign

ReYOLOv8 denotes a set of YOLOv8-derived redesigns in which the standard YOLOv8 detector is retained as a baseline but its feature extraction, temporal modeling, or task structure is reworked for a specialized real-time perception problem. In the most explicit usage, ReYOLOv8 is the recurrent event-based detector introduced for event-camera object detection, where YOLOv8 is augmented with ConvLSTM-based spatiotemporal modeling, Volume of Ternary Event Images (VTEI), and Random Polarity Suppression [2408.05321]. In a separate autonomous-driving context, the same label is a useful shorthand for a refined YOLOv8 variant that replaces C2f with C2f_RFAConv, inserts Triplet Attention, and adds a P2 detection head for small objects [2407.09530]. This suggests that ReYOLOv8 is not a single canonical release, but a broader pattern of re-architected YOLOv8 systems specialized for latency-sensitive vision.

## 1. Terminological scope and relation to standard YOLOv8

YOLOv8, as analyzed in the recent literature, is an anchor-free, single-stage detector built around a CSPNet-based backbone, an FPN+PAN neck, and a dense prediction head that regresses boxes and class scores across multiple feature scales [2408.15857]. The standard pipeline is image input \(\rightarrow\) backbone \(\rightarrow\) neck \(\rightarrow\) head, with multi-scale feature maps typically denoted \(F_3, F_4, F_5\) and their fused counterparts \(\tilde{F}_3, \tilde{F}_4, \tilde{F}_5\) [2408.15857]. In the mainstream Ultralytics formulation assumed by the autonomous-driving refinement paper, the backbone uses stem convolutions and multiple C2f stages, the neck is FPN/PAN-like, and the head is decoupled and predicts class probability, bounding box regression using Distribution Focal Loss, and objectness or quality [2407.09530].

Within that baseline, ReYOLOv8-style work preserves the YOLOv8 detection head and training losses but relocates innovation into the backbone and neck. In the event-based formulation, the essential redesign is temporal: frame-based feature extraction is retained structurally, but recurrent blocks are inserted so that prediction depends on event history rather than on a single encoded window [2408.05321]. In the autonomous-driving refinement, the essential redesign is spatial: conventional C2f internals are replaced by receptive-field-aware convolution, Triplet Attention is injected, and an additional P2 head is introduced to strengthen high-resolution detection [2407.09530].

A common misconception is that ReYOLOv8 necessarily denotes recurrence. The published record does not support that narrower usage. One paper uses ReYOLOv8 to name a recurrent event-based detector, whereas another supports a refined or redesigned YOLOv8 interpretation without recurrent computation [2408.05321]. A plausible implication is that the prefix “Re” functions as “recurrent” in one line of work and “refined” or “redesigned” in another.

## 2. ReYOLOv8 as a recurrent event-based detector

The event-based ReYOLOv8 framework addresses a sensing regime in which the input is not an RGB frame but an asynchronous stream of events \(e_k = (x_k, y_k, p_k, t_k)\), emitted when the log-brightness change exceeds a threshold,
\[
\Delta L(x_k, y_k, t_k) \ge p_k C.
\]
This sensing model is motivated by motion blur, limited dynamic range, and frame-rate latency in standard RGB cameras, particularly in autonomous vehicles and robotics [2408.05321].

Architecturally, the detector keeps the overall YOLOv8 decomposition into backbone, neck, and detection head, but replaces frame input with an event tensor and inserts recurrent blocks after C2f feature blocks. The pipeline is event stream \(\rightarrow\) VTEI \(\rightarrow\) modified YOLOv8 backbone \(\rightarrow\) YOLOv8-style PANet neck \(\rightarrow\) three-scale detection head [2408.05321]. For the small model, the backbone begins with a \(5\)-channel VTEI tensor, uses alternating Conv2D, C2f, and ConvLSTM stages, and ends in SPPF; the neck upsamples and concatenates deeper recurrent features; the detector predicts on feature levels with channels \([88, 176, 344]\) [2408.05321].

The recurrent component is a ConvLSTM inserted after C2f blocks at multiple resolutions. Given current feature map \(x_t\) and previous hidden and cell states \((h_{t-1}, c_{t-1})\), the gates are computed by \(1\times1\) convolutions on \([x_t, h_{t-1}]\), followed by
\[
c_t = r \odot c_{t-1} + c \odot i, \qquad
h_t = o \odot \tanh(c_t).
\]
The hidden state \(h_t\) becomes the recurrently refined feature map passed forward in the backbone [2408.05321]. Because recurrence is applied at several scales, temporal context is accumulated in both high-resolution and low-resolution representations.

The model family is defined at nano, small, and medium scales. ReYOLOv8n, ReYOLOv8s, and ReYOLOv8m contain \(4.7\)M, \(8.4\)M, and \(18.1\)M parameters, respectively [2408.05321]. The same paper reports that, relative to similarly scaled event-based baselines, the models improve mean Average Precision on the GEN1 dataset by \(5\%\), \(2.8\%\), and \(2.5\%\) across nano, small, and medium scales, respectively, while reducing the number of trainable parameters by an average of \(4.43\%\) and maintaining real-time processing speeds between \(9.2\) ms and \(15.5\) ms [2408.05321].

## 3. Event encoding, augmentation, and training protocol

A defining component of recurrent ReYOLOv8 is VTEI, a dense event representation designed for low latency and low memory. For a time window \([t_0, t_N]\) and \(B=5\) temporal bins, each event timestamp is normalized to a bin index,
\[
T_k = \frac{t_k - t_0}{t_N - t_0} B,
\]
and a tensor \(I \in \mathbb{R}^{B \times H \times W}\) is filled so that each voxel stores the polarity of the last event observed at that pixel and bin [2408.05321]. The resulting values are ternary: \(-1\), \(0\), or \(+1\). The encoding is intentionally compact, because only the last polarity is stored per pixel-bin.

The paper compares VTEI with MDES, SHist, and voxel grids. For GEN1 with \(B=5\), \(H=240\), and \(W=304\), VTEI uses \(3\) bytes per non-zero entry in COO form, and at a \(50\) ms window with an average of about \(192\)k events it achieves \(1.5\) ms encoding latency and \(128\) Mev/s event processing rate [2408.05321]. The same analysis reports compression ratios of \(2.53\)–\(2.95\times\), with better latency than SHist, MDES, and voxel grids, and lower bandwidth than SHist and voxel grids [2408.05321]. The detector’s temporal modeling therefore rests on an encoding that is efficient enough not to erase the latency advantage of event sensing.

Training uses truncated backpropagation through time. Sequence length is \(11\) for GEN1 and \(5\) for PEDRo; hidden and cell states are reset between clips during training and at sequence boundaries during validation and test [2408.05321]. The optimizer is SGD with momentum \(0.937\); weight decay is \(0.011\) for GEN1 and \(0.005\) for PEDRo; there is a \(3\)-epoch warm-up with momentum \(0.8\) and bias learning rate \(0.1\); and the image sizes are \(320\times224\) for GEN1 and \(352\times288\) for PEDRo [2408.05321]. Losses remain YOLOv8 defaults with box loss \(7.5\), class loss \(0.5\), and DFL \(1.5\) [2408.05321].

The event-specific augmentation is Random Polarity Suppression. Given a VTEI tensor \(I\), the transformation either leaves \(I\) unchanged or removes one polarity across all temporal bins:
\[
I' =
\begin{cases}
I, & \text{if } r_1 \ge s,\\
I_p, & \text{if } r_1 < s \text{ and } r_2 \ge p,\\
I_n, & \text{otherwise}.
\end{cases}
\]
Here \(s\) is the suppression probability and \(p\) the polarity-selection probability [2408.05321]. The reported sweeps show that small suppression probabilities, typically \(s \le 0.125\), are preferable, and that balanced suppression with \(p=0.5\) is usually best [2408.05321]. The gains are modest on GEN1 but substantially larger on PEDRo, where ReYOLOv8n improves from \(59.0\) to \(63.9\) mAP and ReYOLOv8m from \(66.5\) to \(69.1\) mAP under the selected settings [2408.05321].

## 4. Empirical performance on GEN1 and PEDRo

The event-based ReYOLOv8 paper evaluates on two benchmarks with distinct operating regimes: Prophesee GEN1 for automotive perception and PEDRo for robotics [2408.05321]. GEN1 contains urban driving scenes with cars and pedestrians at \(304\times240\) resolution and about \(255\)k bounding boxes; PEDRo contains DAVIS346 handheld robotics recordings at \(346\times260\) resolution and about \(43\)k pedestrian boxes [2408.05321]. The primary metric is COCO-style mAP@[0.5:0.95].

On PEDRo, ReYOLOv8n, ReYOLOv8s, and ReYOLOv8m obtain \(63.9\), \(65.5\), and \(69.1\) mAP@[0.5:0.95], with runtimes of \(9.2\) ms, \(10.4\) ms, and \(12.3\) ms, respectively [2408.05321]. The YOLOv8x baseline adapted to events achieves \(58.6\) mAP@[0.5:0.95] with \(68.2\)M parameters and \(17.6\) ms runtime, whereas ReYOLOv8m reaches higher accuracy with \(18.1\)M parameters and ReYOLOv8n is about \(14.5\times\) smaller [2408.05321]. The paper summarizes this as a \(9\%\) to \(18\%\) mAP improvement over the YOLOv8x-based baseline, with \(14.5\times\) and \(3.8\times\) smaller models and an average speed enhancement of \(1.67\times\) on PEDRo [2408.05321].

On GEN1, the reported gains are scale-matched against event-based competitors. At nano scale, ReYOLOv8n reaches \(46.3\) mAP with \(4.7\)M parameters and \(9.2\) ms runtime, compared with RVT-T at \(44.1\) mAP and \(4.4\)M parameters [2408.05321]. At small scale, ReYOLOv8s reaches \(48.3\) mAP with \(8.4\)M parameters; the paper highlights a \(+2.8\%\) mAP gain over HMNet-L1 while using about \(27\%\) fewer parameters [2408.05321]. At medium scale, ReYOLOv8m reaches \(49.4\) mAP with \(18.1\)M parameters, exceeding SAST-CB at \(48.2\) and DTSDNet-M at \(47.7\) mAP [2408.05321].

The qualitative interpretation offered by the study is that recurrence stabilizes detection under sparse or noisy events, while VTEI keeps the end-to-end pipeline latency compatible with real-time use [2408.05321]. Failure cases remain concentrated in extremely small distant pedestrians and cluttered backgrounds whose event patterns resemble noise [2408.05321]. Those limitations are consistent with the fact that temporal integration can compensate for sparsity only up to the point where the object generates too little signal.

## 5. ReYOLOv8-style refinement for frame-based autonomous driving

A separate line of work applies the ReYOLOv8 idea to frame-based detection for autonomous driving rather than to event streams. In this formulation, the baseline is mainstream Ultralytics YOLOv8, but all positions using C2f are replaced internally by C2f_RFAConv blocks, Triplet Attention modules are inserted around key backbone and neck feature maps, and a new P2 detection head is added for small-object detection [2407.09530]. The paper keeps the detection head and the training loss of YOLOv8 basically intact, so the redesign is concentrated in feature extraction [2407.09530].

The key operation is Receptive Field Attention Convolution. Instead of applying the same kernel \(K\) at every spatial position, the method introduces attention weights \(A_p(q)\) so that the effective kernel becomes
\[
K_p(q) = A_p(q) \cdot K(q).
\]
The paper describes the resulting mechanism qualitatively as “Kernel1 = \(A_1 \cdot K\), Kernel2 = \(A_2 \cdot K\), Kernel3 = \(A_3 \cdot K\)” and so on, so that each receptive field sees a different kernel adapted to local content [2407.09530]. The C2f_RFAConv block preserves the global split–multi-block–concat–fuse layout of C2f while replacing the Bottleneck internals with RFAConv-based computation [2407.09530].

Triplet Attention is used in the form described in “Rotate to Attend: Convolutional Triplet Attention Module,” with three branches that model channel–height–width interactions through permutation, Z-pooling, \(k\times k\) convolution, sigmoid gating, and aggregation [2407.09530]. In the reported configuration, these modules are inserted into the improved YOLOv8 network structure after C2f_RFAConv is already in place [2407.09530]. The additional P2 head extends the detection scale hierarchy upward to higher-resolution features, with the explicit aim of strengthening small-object detection [2407.09530].

The reported gains are numerical and stepwise. On the COCO dataset referred to in the paper, the original YOLOv8 baseline obtains mAP@0.5 \(= 0.326\) and mAP@0.5:0.95 \(= 0.187\) [2407.09530]. After C2f_RFAConv only, the reported all-classes mAP@0.5 is about \(0.378\) [2407.09530]. After C2f_RFAConv, Triplet Attention, and P2 are combined, the final model reaches mAP@0.5 \(= 0.385\) and mAP@0.5:0.95 \(= 0.217\), corresponding to absolute gains of \(0.059\) and \(0.030\) over the baseline [2407.09530]. The same source reports per-class PR improvements, for example pedestrian from \(0.335\) to \(0.441\) and van from \(0.751\) to \(0.811\) [2407.09530].

This version of ReYOLOv8 does not provide explicit FLOPs or FPS, but it states that RFAConv and Triplet Attention are designed to be lightweight and that the model continues to satisfy the real-time requirement for autonomous driving [2407.09530]. The paper also notes that experiments are done on COCO and that generalization to autonomous-driving benchmarks such as KITTI, BDD100K, nuScenes, or Waymo is not explicitly shown [2407.09530].

## 6. Related redesign patterns, deployment logic, and limitations

The broader YOLOv8 literature shows that ReYOLOv8 sits within a wider pattern of task-aligned redesign. In distracted-driving classification, P-YOLOv8 adapts YOLOv8n-cls through transfer learning and reports \(99.46\%\) accuracy with \(1{,}451{,}098\) parameters and a \(2.84\) MB model size on the State Farm dataset [2410.15602]. That model is not called ReYOLOv8, but it demonstrates the same design principle: begin from a small YOLOv8 backbone, specialize the head to the target label space, and exploit pretraining to preserve accuracy under a strict parameter budget [2410.15602]. In autonomous-driving multi-task learning, A-YOLOM keeps a YOLOv8-style backbone and detection head, adds separate segmentation necks with an Adaptive Concatenation Module, and uses lightweight segmentation heads; the nano variant has \(4.43\)M parameters and runs at \(39.9\) FPS while jointly performing detection, drivable-area segmentation, and lane-line segmentation [2310.01641]. In UAV localization, YoloTag uses a lightweight YOLOv8 detector in a closed-loop system with EPnP and a Butterworth filter, and reports \(55\) FPS on a Quadro P2200 [2409.02334]. Collectively, these works place ReYOLOv8 in a larger engineering tradition of YOLOv8 specialization for edge or real-time systems.

Several design regularities recur across these derivatives. First, the YOLOv8 backbone–neck–head decomposition is typically preserved, even when individual blocks are replaced or auxiliary heads are added [2408.15857]. Second, the smallest viable configuration is often preferred for embedded deployment, with task-specific heads or recurrent blocks added only where the target application justifies the overhead [2410.15602]. Third, real-time use cases repeatedly push modifications into the feature extractor and leave the high-level training and deployment stack recognizable as YOLOv8 [2407.09530].

The main limitation of the term itself is semantic rather than algorithmic: there is no single standardized ReYOLOv8 architecture in the literature. One established usage denotes the recurrent event-based detector on GEN1 and PEDRo, while another denotes a refined frame-based detector built from C2f_RFAConv, Triplet Attention, and a P2 head [2408.05321]. A second limitation is evidential scope. The event-based formulation is validated on two event datasets, and the frame-based autonomous-driving refinement is evaluated on COCO-style traffic classes rather than on a broad suite of driving benchmarks [2407.09530]. The frame-based paper also does not report detailed FLOPs or FPS, and the event-based paper notes the latency cost of ConvLSTM relative to the fastest non-recurrent baselines [2408.05321].

A plausible implication is that ReYOLOv8 should be treated as a research pattern rather than a fixed model name. In that pattern, YOLOv8 supplies the scalable real-time scaffold, while domain-specific pressures—event sparsity, small-object traffic detection, multi-task scene understanding, or extreme parameter limits—determine whether the redesign is recurrent, attention-based, receptive-field-aware, or head-specialized.

Source: https://www.emergentmind.com/topics/reyolov8