- The paper introduces Evita, a unified backbone for dense RGB-event parsing that integrates geometric parallax rectification, harmonic spectral resonance, and transient global routing.
- It employs a symmetric hierarchical architecture with cross-modal co-learning modules to align asynchronous event data with RGB images while reducing computational overhead.
- Empirical evaluations demonstrate state-of-the-art performance with superior mIoU, robust reconstruction, and real-time deployment readiness under degraded visual conditions.
Unified Harmonic-Geometric Representation Learning for Dense RGB-Event Parsing
Motivation and Context
The fusion of asynchronous event streams with dense RGB imagery achieves resilient perception under degraded visual conditions prevalent in autonomous driving and robotics. Traditional multimodal frameworks employ dual encoders, which exacerbate parameter overhead and computational latency while inadequately addressing geometric parallax induced by asynchronous sensor delays and cross-spectral aliasing from orthogonal physical spectra. The representational gap between intensity grids (RGB) and kinematic spikes (events) remains an unresolved barrier for unified architectures. Evita is designed as the first backbone purpose-built for dense RGB-event parsing. Its architecture embeds modality synergy directly in each encoder block via three core operations: Geometric Parallax Rectification, Harmonic Spectral Resonance, and Transient Global Routing.
Architecture and Methodology
Evita adopts a symmetric, intertwined hierarchical backbone, integrating cross-modal co-learning modules at every stage. The pipeline initiates with parallel modality stems projecting RGB images and stochastically-encoded event streams into feature spaces. At each hierarchical layer, geometric and spectral harmonics are jointly optimized for robust cross-modal synthesis.

Figure 1: Structural overview of Evita's intertwined backbone, Evita block, parallel routing, and mechanism integration for geometric and spectral alignment.
Geometric Parallax Rectification
Evita's parallax rectification module leverages a cross-modal deformable convolution to align structural boundaries between modalities. Adaptively predicted offset magnitudes and deform masks anchor misaligned event edges to authentic spatial coordinates without requiring explicit depth calibration.

Figure 2: Visualization of geometric parallax rectification, anchoring kinematic edges to true physical boundaries.
Harmonic Spectral Resonance
Spectral resonance executes cross-modal texture transfer strictly in the complex frequency domain. The amplitude-gated harmonics derived from event features are injected into the RGB frequency profile, followed by polar fusion and inverse transform. By decoupling amplitude and phase, the mechanism mitigates cross-modal aliasing without spatial noise artifacts.
Transient Global Routing
Evita establishes an asymmetric cross-modal attention mechanism through the Transient Global Routing layer. Event features form dynamic queries, spatially aligned with RGB keys and values, supplemented by an additive spatial bias computed from event priors. This structure achieves effective routing of photometric context based on motion-centric anchors.

Figure 3: Transient Global Routing dynamically focuses attention on motion boundaries, suppressing static noise.
Pretraining and Dataset Design
Evita's pretraining paradigm employs N-ImageNetV2, a large-scale multimodal dataset constructed to ensure strict spatial alignment between RGB and event pairs using semantic-guided registration and iterative auto-alignment. To avoid overfitting and maximize generalization, training involves stochastic spatial perturbation, cycling between aligned and misaligned states to compel the parallax rectification module to learn robust deformation fields.

Figure 4: Ablation analysis shows optimal stochastic misalignment probability peaking at $0.6$, promoting spatial robustness.
Empirical Evaluation
Evita is rigorously compared to seventeen multimodal semantic segmentation frameworks across DELIVER, DDD17, and DSEC datasets. Evita-L establishes new state-of-the-art performanceโon DELIVER, it achieves 59.57% mIoU, surpassing CMNeXt-B4 by 0.36% accuracy with half the computational FLOPs and only 38% of its parameter count.

Figure 5: Evita achieves new SOTA 59.57% mIoU and optimal accuracy/compute equilibrium on DELIVER.
Qualitative results confirm superior robustness in challenging conditions, reconstructing fine-grained structures and maintaining semantic consistency under low light and occlusion, outperforming legacy architectures.

Figure 6: Evita demonstrates robust reconstruction and semantic consistency in adverse weather and high-occlusion scenarios.
Ablation studies demonstrate individual and combined benefits of geometric and spectral modules; coupling both boosts DDD17 mIoU by +2.17% versus baseline. Integrating event-derived spatial bias via additive formulation yields peak accuracy and computational efficiency in routing.
Evita's cross-modal resilience to spatial misalignment is quantitatively superior: the mIoU penalty under $48$-pixel offset is limited to โ0.18%, compared to โ2.43% and โ1.71% in CMX and CMNeXt, respectively.
Multi-scale Feature Evolution and Modality Universality
By propagating symbiotic learning across hierarchical depths, Evita converges high-level semantic representations, transforming modality heterogeneity in shallow layers to unified feature patterns at deeper stages.

Figure 7: Multi-scale representation evolution: symbiotic learning yields unified semantic features in deep layers.
Evita's harmonic-geometric paradigm generalizes to unseen modalities (thermal, LiDAR), consistently improving semantic segmentation accuracy on MFNet and KITTI-360, confirming the structural modules' universality.
Inference Latency and Practical Deployment
Evita achieves a superior accuracy-latency tradeoff. On DSEC, Evita-L delivers 76.80% mIoU in 45.6 msโalmost half the inference time of MambaSeg (85.1 ms) for higher accuracy. Evita-S achieves real-time deployment readiness with only 22.4 ms latency at comparable accuracy.
Implications and Future Developments
Evita demonstrates that embedding explicit geometric and spectral mechanisms within unified backbones systematically overcomes representational and computational bottlenecks in dense RGB-event parsing. Its resilience to spatial misalignment and generalizability to arbitrary event formats suggest that harmonically-aligned architectures present a scalable foundation for robust multimodal perception. The universal benefit to alternative modalities underscores the value of spectral and geometric priors in cross-modal fusion, opening avenues for agnostic, large-scale multimodal foundation models. Synthesis of pseudo multi-spectral datasets and scaling symbiotic co-learning to more modalities could further extend Evita's applicability to complex sensor ecosystems in real-world autonomy.
Conclusion
Evita introduces a principled, unified harmonic-geometric backbone for dense RGB-event parsing, embedding spatial alignment, frequency-domain fusion, and asymmetric routing modules into every encoder layer. Supported by strictly aligned N-ImageNetV2 and stochastic pretraining, Evita achieves superior structural robustness, cross-modal resilience, accuracy-latency equilibrium, and universality across modalities. Its architectural paradigm defines a scalable framework for real-time multimodal perception, poised for extension to broader sensory domains and application scenarios.