Papers
Topics
Authors
Recent
Search
2000 character limit reached

RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding

Published 19 May 2026 in cs.CV and cs.AI | (2605.19329v1)

Abstract: Conventional vision-LLMs (VLMs) struggle to interpret scenes captured under adverse conditions (e.g., low light, high dynamic range, or fast motion) because standard RGB images degrade in such environments. Event cameras provide a complementary modality: they asynchronously record per-pixel brightness changes with high temporal resolution and wide dynamic range, preserving motion cues where frames fail. We propose RE-VLM, the first dual-stream vision-LLM that jointly leverages RGB images and event streams for robust scene understanding across both normal and challenging conditions. RE-VLM employs parallel RGB and event encoders together with a progressive training strategy that aligns heterogeneous visual features with language. To address the scarcity of RGB-Event-Text supervision, we further propose a graph-driven pipeline that converts synchronized RGB-Event streams into verifiable scene graphs, from which we synthesize captions and question-answer (QA) pairs. To develop and evaluate RE-VLM, we construct two datasets: PEOD-Chat, targeting illumination-challenged scenes, and RGBE-Chat, covering diverse scenarios. On captioning and VQA benchmarks, RE-VLM consistently outperforms state-of-the-art RGB-only and event-only models with comparable parameter counts, with particularly large gains under challenging conditions. These results demonstrate the effectiveness of event-augmented VLMs in achieving robust vision-language understanding across a wide range of real-world environments. Code and datasets are available at https://github.com/bupt-ai-cz/RE-VLM.

Summary

  • The paper introduces RE-VLM, a dual-stream vision-language model that combines RGB appearance with event-based motion and structure, achieving 0.63 and 0.75 VQA accuracy on PEOD-Chat and RGBE-Chat.
  • The paper uses a graph-driven RGB-Event-Text pipeline and degradation-aware fusion to create more reliable supervision, reducing human-audited question-answer corrections from 54.2% to 18.1%.
  • The paper shows that joint RGB-event input consistently outperforms single-modality and larger baseline models, with the greatest gains in adverse conditions such as low light, overexposure, glare, and motion blur.

RE-VLM addresses a concrete failure mode of modern vision-LLMs: degradation under adverse imaging conditions. Standard RGB-based VLMs such as Qwen2.5-VL, InternVL, and DeepSeek-VL produce unreliable descriptions in low-light, overexposed, or fast-motion scenes because the RGB signal itself is corrupted. Event cameras, which asynchronously record per-pixel log-intensity changes with microsecond latency and high dynamic range, preserve motion and structure where frames fail—but they carry no color or texture information. The paper proposes RE-VLM, a dual-stream VLM that fuses both modalities, together with a graph-driven data pipeline that manufactures the RGB-Event-Text supervision such a model requires (2605.19329).

Graph-driven RGB-Event-Text data generation

The central data bottleneck is that no large-scale corpus of synchronized RGB frames, event streams, and text exists. The authors' pipeline converts paired streams into a verifiable intermediate representation rather than relying on an RGB-trained VLM to hallucinate captions under degradation—an approach they argue fails precisely when RGB is corrupted.

The pipeline proceeds in three steps. First, an event window of N×33N \times 33 ms (with N=4N{=}4) centered on each RGB keyframe is reconstructed into grayscale frames via NER-Net, forming a video-like event tensor that is captioned with a structured subject–motion–place–relation schema and parsed by an LLM into an event graph emphasizing motion, temporal ordering, and topology. Second, an RGB graph is built from the synchronized keyframe, encoding appearance attributes (color, texture, shape), geometry, and explicit degradation labels (low light, overexposure, glare, motion blur) attached to nodes and edges. Third, a degradation-aware fusion arbitrates at the field level: motion and structural facts are anchored to the event graph, while color, lighting, and text facts come from the RGB graph only when it is not severely degraded; degraded RGB conclusions are retained only as low-confidence candidates, and geometry fields follow RGB when the two graphs disagree. Captions and up to three VQA items are then synthesized from the fused graph.

A human audit on 855 PEOD samples quantifies the reliability gain: the RGB-only generation baseline (EventGPT) required correction on 54.2% of QA items, versus 18.1% for the proposed pipeline. This is the strongest single piece of evidence in the paper that the graph-mediated, degradation-aware route produces materially more trustworthy supervision than direct RGB-based caption generation in adverse scenes. The resulting datasets—PEOD-Chat (11k samples from PEOD, targeting illumination challenges) and RGBE-Chat (113.7k samples spanning RGBE-ImageNet, DSEC, DDD17, RGBE-SEG, MVSEC, and M3ED)—serve as both training corpora and benchmarks. Note that RGBE-ImageNet event data are synthesized rather than natively captured, which the authors acknowledge via the generation method cited.

Architecture and training

RE-VLM is built on Qwen2.5-VL-3B and adds a ViT-based event encoder, modality adapters, and a training-only Spatio-Temporal Alignment Module (STAM). Event slices (Nw=3N_w{=}3) are encoded per-slice, processed with multi-scale depthwise 1D temporal convolutions and SE-style temporal weighting to emphasize salient motion intervals, then projected into the LLM space alongside RGB tokens. At inference the LLM decodes from [P;Ti;Te][P; T_i; T_e], and the model supports RGB-only, event-only, or joint input.

STAM computes per-modality self-attention affinity matrices over a shared spatio-temporal lattice, derives token saliency as graph degree, fuses the two saliency maps, and applies a relation loss—the weighted spatial inner product between the fused importance map and the per-pixel RGB–event feature discrepancy. This penalizes cross-modal mismatch more strongly in salient regions, with total loss L=LLLM+λLCA-WTDL = L_{\text{LLM}} + \lambda L_{\text{CA-WTD}} at λ=0.1\lambda = 0.1. STAM is discarded at inference, so it adds no deployment cost.

Training follows a three-stage curriculum: (1) event–language alignment on 1,300K pairs with the LLM and RGB branch frozen; (2) event–RGB alignment with STAM on 600K pairs, again with the LLM frozen; (3) LoRA-based instruction tuning on 120K samples with both visual branches frozen. This ordering preserves the pretrained RGB-language alignment while injecting event semantics incrementally.

Experimental results

Evaluation uses GPT-3.5-Turbo as an LLM judge (Video-ChatGPT protocol) on zero-shot captioning (CI, DO, CU on a 0–5 Likert scale) and VQA (average score and attribute-level accuracy). Test sets comprise 1,750 PEOD-Chat samples (60% adverse, 40% normal) and 2,047 RGBE-Chat samples, with strict sequence-level splits.

Model Params PEOD-Chat Cap. Ave PEOD-Chat Acc RGBE-Chat Cap. Ave RGBE-Chat Acc
Qwen2.5-VL (RGB) 3B 3.47 0.52 3.80 0.66
InternVL2 (RGB) 4B 3.36 0.49 3.70 0.68
DeepSeek2-VL (RGB) 7B 3.37 0.50 3.49 0.52
EventGPT (event) 7B 3.04 0.40 3.10 0.39
RE-VLM (RGB+Event) 4B 3.82 0.63 4.20 0.75

RE-VLM at 4B parameters outperforms baselines with up to 2×2\times its parameter count on every metric on both benchmarks. The margin is largest on illumination-challenged PEOD-Chat, where the fine-tuned RGB-only Qwen2.5-VL variant reaches only 0.55 VQA accuracy versus RE-VLM's 0.63, and EventGPT reaches 0.40. The qualitative examples reinforce the complementarity claim: under severe overexposure, the RGB baseline misses a city bus entirely while the event-only model cannot determine its color; RE-VLM recovers both.

The ablations support the design choices. Single-modality inference remains functional but is consistently inferior to joint input (e.g., RGB-only 0.57 vs. joint 0.63 VQA accuracy on PEOD-Chat), confirming genuine fusion rather than RGB dominance. Replacing STAM with naive concatenation degrades PEOD-Chat accuracy from 0.63 to 0.61, though on RGBE-Chat the STAM gain is marginal (0.75 vs. 0.74)—the alignment module's benefit is concentrated in adverse conditions. Cross-validation with the open-source Qwen3-Omni-30B judge reproduces the ranking (RE-VLM 0.62 VQA accuracy vs. 0.48 for Qwen2.5-VL and 0.40 for EventGPT), addressing concerns about reliance on a closed-source judge.

Limitations and open questions

Several caveats are explicit in the paper. The data pipeline's degradation arbitration is rule-based and LLM-mediated; its correctness rests on the assumption that degradation can be reliably diagnosed from RGB-graph labels, which is not independently validated. The human audit covers 855 PEOD samples only, leaving the error profile of RGBE-Chat—where much of the event data is synthesized—unquantified. Evaluation relies on LLM-as-a-judge scoring; although cross-checked with a second judge, no human evaluation of model outputs is reported. The ablation shows STAM's contribution on general scenes is small, leaving open whether the alignment loss is justified outside adverse conditions. Finally, the comparison renders event streams into event images for all baselines "to ensure fair comparison," which may disadvantage baselines designed for asynchronous event input; whether native asynchronous processing would narrow the gap is unexamined.

Conclusion

RE-VLM demonstrates that pairing RGB and event streams in a single VLM yields consistent gains in captioning and VQA over RGB-only and event-only baselines of comparable or larger size, with the largest improvements under low-light, HDR, and fast-motion conditions. Its graph-driven data pipeline offers a reusable recipe for manufacturing verifiable multimodal supervision where paired data is scarce, and its human audit provides quantitative evidence of the pipeline's reliability advantage. The work leaves open the questions of scaling the pipeline beyond rule-based arbitration, validating supervision quality at scale on synthesized event data, and processing events asynchronously rather than as rendered frames.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.