EvGNN: Event-Driven GNN Accelerator
- EvGNN is an event-driven GNN accelerator for edge vision that leverages a directed dynamic graph formulation and event-queue storage to process sparse, asynchronous events.
- It employs a spatiotemporally decoupled prism search to efficiently construct local causal subgraphs, reducing neighborhood processing to only the necessary 1-hop regions.
- The accelerator uses layer-parallel execution with INT8 quantization to achieve competitive accuracy while drastically lowering latency and memory footprint on FPGA hardware.
EvGNN most directly denotes the hardware-software co-designed system introduced in "EvGNN: An Event-driven Graph Neural Network Accelerator for Edge Vision" (Yang et al., 2024). It is presented as the first dedicated hardware accelerator for event-driven graph neural networks in edge vision, targeting the sparse, asynchronous output of event-based cameras without converting the stream into dense frames. The system combines a directed dynamic graph formulation, event-queue-based neighborhood construction, and a layer-parallel execution scheme to achieve low-footprint, ultra-low-latency, and high-accuracy edge vision on FPGA hardware. In adjacent graph learning literature, similar abbreviations are also used for different concepts, including E(n)-equivariant GNNs, Eigen-GNN, and evolving graph convolutional models; this naming overlap makes contextual disambiguation necessary (Satorras et al., 2021).
1. Definition and nomenclature
In the 2024 edge-vision literature, EvGNN refers to an event-driven GNN accelerator designed for event-based cameras and deployed on a Xilinx KV260 Ultrascale+ MPSoC platform (Yang et al., 2024). Its central objective is to enable real-time, microsecond-resolution inference at the edge by avoiding frame conversion and by exploiting the sparsity and causality of event streams. The design is explicitly framed as a co-design of event representation, graph construction, and GNN execution for FPGA implementation.
The term is not globally unique across graph-learning research. Closely related names include EGNN for E(n)-Equivariant Graph Neural Networks (Satorras et al., 2021), Eigen-GNN as a graph structure preserving plug-in for shallow GNNs (Zhang et al., 2020), and EGCN as an evolving graph convolutional neural network that learns graph Laplacians during training (Li et al., 2017). This suggests that "EvGNN" should be interpreted by domain context rather than by acronym alone.
| Name in literature | Primary meaning | Representative paper |
|---|---|---|
| EvGNN | Event-driven GNN accelerator for edge vision | (Yang et al., 2024) |
| EGNN | E(n)-Equivariant Graph Neural Networks | (Satorras et al., 2021) |
| Eigen-GNN | Graph structure preserving plug-in via eigenspace augmentation | (Zhang et al., 2020) |
A plausible implication is that the event-driven EvGNN sits at the intersection of neuromorphic sensing, sparse dynamic graphs, and hardware-aware GNN deployment, whereas the similarly named models address symmetry, spectral structure preservation, or graph evolution in software-centric settings.
2. Event-based vision as a graph problem
EvGNN is motivated by event cameras, which emit asynchronous events only when a pixel’s intensity changes beyond a threshold. Each event is represented as
where is pixel location, is timestamp, and is polarity (Yang et al., 2024). Because events are sparse, local, and microsecond-resolved, the paper argues that dense CNN-style processing is inefficient for low-latency edge applications.
The graph formulation stores each event as a node with features . For a node , its 1-hop neighborhood is , and standard message passing is written as
For directed graphs, only messages along edges from past nodes to future nodes are considered (Yang et al., 2024). In EvGNN, this causal restriction is not merely a modeling assumption; it is the basis of the accelerator’s latency reduction.
The paper contrasts this with workflows that first construct a full event graph and then run a GNN, noting that such pipelines drive latency to millisecond scale. Even dynamic event-driven GNNs remain difficult to accelerate because graph neighborhoods are irregular, data-dependent, and costly to search and store. EvGNN addresses these bottlenecks by collapsing per-event processing to the local causal subgraph of the newly arrived event.
3. Directed dynamic graphs, edge-free storage, and prism search
EvGNN relies on three central ideas. The first is a directed dynamic graph formulation inspired by HUGNet, in which edges point only from older events to the newly arrived event (Yang et al., 2024). Two consequences are emphasized. First, only the new node’s features need to be updated for each event arrival. Second, only the 1-hop neighborhood of the new event must be processed, even for a multi-layer GNN, because past nodes’ features are fixed. The paper states that this collapses the processing range from the $1$-to--hop neighborhoods required in approaches like AEGNN to just the 1-hop subgraph.
A second consequence is edge-free storage. Since edges are transient and only connect the current event to past neighbors, EvGNN stores only nodes in event queues, with one queue per pixel location. Each queue stores event index 0, timestamp 1, and polarity 2 (Yang et al., 2024). The paper identifies this as a major memory-saving design choice.
The second central idea is the spatiotemporally decoupled prism search for neighbor identification. Previous work used a scaled 3 radius
4
whereas EvGNN separates the search into spatial and temporal conditions: 5 or, with 6 spatial distance,
7
This produces a prism-shaped search region that is hardware-friendly because the spatial and temporal checks can be done separately and pipelined (Yang et al., 2024).
The paper reports a validation study on directed AEGNN comparing four search regions: hemi-sphere with 8, semi-octahedron with 9, cylinder with 0, and prism with 1. Their reported accuracies are 2, 3, 4, and 5, respectively, and the prism is selected for hardware efficiency (Yang et al., 2024). This suggests that EvGNN accepts a small geometric approximation in neighborhood construction in exchange for a lower implementation footprint.
4. Layer-parallel GNN execution and accelerator microarchitecture
The third central idea is layer-parallel execution. EvGNN exploits the causality of directed event graphs to compute a new event’s features for all GNN layers in parallel rather than serially (Yang et al., 2024). The underlying argument is that, for the newly arrived event, each layer depends only on fixed features of past neighbors, not on newly recomputed neighbor states from deeper layers. The paper describes this as mathematically equivalent for directed event graphs while substantially reducing latency.
The simplified graph convolution used in EvGNN is a PointNet-style operator: 6 Spatial information is retained by concatenating 7 and 8 to node features before linear transformation and max aggregation (Yang et al., 2024). Compared with a SplineConv-based baseline, the paper reports 9 validation accuracy with 30.4k parameters for SplineConv and 0 with 4.8k parameters for EvGNN’s simplified convolution; a GCN baseline is listed at 1 with 4.7k parameters (Yang et al., 2024).
The network is further simplified by removing residual connections and pooling layers and by applying batchnorm folding and INT8 quantization. According to the paper, this reduces parameter memory footprint from 121.6 kB for the SplineConv baseline to 6.6 kB in the INT8 version while maintaining 2 validation accuracy (Yang et al., 2024).
At the hardware level, the accelerator is partitioned into a Graph Construction Module, a Graph Convolution Module, and Graph Readout and FC Modules, together with control/configuration and AXI communication blocks (Yang et al., 2024). The graph construction pipeline receives a new event 3, performs spatial queue selection within 4, transfers selected queues to a local event buffer, performs temporal filtering with 5, and stores valid neighbors in a neighbor buffer. The graph convolution block then retrieves previously computed features for all layers from external DRAM via AXI MM, broadcasts augmented neighbor features to all GNN layers in parallel, performs MatVec-based linear transforms with local BRAM-stored weights, aggregates by per-channel max, applies a BAQ stage comprising bias, ReLU, and INT8 quantization, and writes outputs back to DRAM (Yang et al., 2024).
The graph readout divides the 6 input space into an 7 grid and performs max pooling within 8-pixel patches before the FC prediction head (Yang et al., 2024). The FC stage reuses the MatVec unit. This reuse indicates a deliberate reduction in hardware specialization overhead.
5. Empirical evaluation and reported performance
The software stack is implemented and tested with PyTorch Geometric, trained on NVIDIA RTX A6000 GPUs, and evaluated on the N-CARS dataset (Yang et al., 2024). The training set contains 15,422 event stream samples, with an 85% training / 15% validation split using bootstrapping, batch size 64, learning rate 0.002, Adam optimizer, and 100 training epochs. Hardware benchmarking is conducted on a Xilinx KV260 development board, with the accelerator in programmable logic and the ARM subsystem handling benchmarking and I/O; the test set contains 8,607 event stream samples (Yang et al., 2024).
On N-CARS, the reported test accuracies are 9 for software and 0 for hardware (Yang et al., 2024). The paper lists comparison figures for H-First at 1, HOTS at 2, HATS at 3, YOLE at 4, AsyNet at 5, NVS-S at 6, EvS-S at 7, and AEGNN at 8 using open-source code (Yang et al., 2024). The authors also note that some comparisons are not perfectly apples-to-apples because earlier methods did not always report the exact train/validation/test splits.
The headline latency result is an average latency per event of 9 (Yang et al., 2024). The reported resource usage on KV260 is 30,908 / 117,120 LUTs, 24,083 / 234,240 FFs, 228 / 1,248 DSPs, 0.45 kB / 0.44 MB LUTRAM, 85.2 kB / 0.63 MB BRAM, and 1.68 MB / 2.25 MB UltraRAM, with total on-chip memory footprint reported as 1.76 MB (Yang et al., 2024). The paper emphasizes that most on-chip memory is consumed by the event queues in the graph construction module.
These results position EvGNN as a low-latency accelerator rather than a highest-accuracy event-vision model. A plausible implication is that its main contribution is the attainment of microsecond-scale event-native inference under edge deployment constraints, while preserving accuracy competitive with dynamic event-driven GNN baselines.
6. Design trade-offs, limitations, and relation to surrounding GNN research
The paper explicitly characterizes several trade-offs. Directed graphs reduce search and storage costs but rely on causality; prism search is simpler to implement than a scaled 3D radius search but approximates neighborhood geometry; simplified PointNetConv plus INT8 quantization reduces parameter and compute cost with only a small accuracy penalty; and layer-parallel execution lowers latency but requires storing intermediate features for each layer and fetching them efficiently, which adds memory traffic (Yang et al., 2024). The design is also tuned for low-latency edge inference rather than maximum throughput on large batched workloads.
Within the broader GNN landscape, EvGNN is distinct from symmetry-preserving geometric models such as EGNN, which updates node coordinates and features while preserving translation, rotation, reflection, and permutation equivariance through relative distances and vector coordinate updates (Satorras et al., 2021). It is also distinct from spectral structure-preserving models such as Eigen-GNN, which augments the initial feature basis with top-0 eigenvectors of a graph structure matrix to improve shallow GNN performance on structure-driven tasks (Zhang et al., 2020). Likewise, it differs from evolving-graph models such as EGCN, where the graph Laplacian itself is updated during supervised training via learned metric learning and residual Laplacian refinement (Li et al., 2017).
This contrast clarifies the scope of EvGNN. It is not primarily a new graph neural architecture in the algorithmic sense of equivariant message passing, spectral augmentation, or adaptive Laplacian learning. Rather, it is a hardware-conscious event-driven execution framework whose novelty lies in the coupling of directed dynamic event graphs, edge-free event-queue storage, spatiotemporally decoupled prism neighbor search, and layer-parallel multi-layer GNN execution (Yang et al., 2024). In that sense, EvGNN is best understood as a system-level contribution to sparse edge vision that uses GNN principles as an event-native computational substrate.