---
title: Event-aided Direct Sparse Odometry (EDS)
url: https://www.emergentmind.com/topics/event-aided-direct-sparse-odometry-eds
type: topic
---

# Event-aided Direct Sparse Odometry (EDS)

Event-aided Direct Sparse Odometry (EDS) is a class of visual odometry (VO) algorithms that directly fuses asynchronous event streams from event cameras with standard image frames (and optionally depth) to estimate 6-DoF camera motion and reconstruct semi-dense 3D maps. By leveraging the high temporal resolution, high dynamic range, and blur resilience of event cameras, EDS provides robust, low-latency odometry even under rapid motion, sparse frames, or challenging illumination. The technique uniquely integrates a direct probabilistic formulation of per-pixel brightness increments predicted via sparse 3D structure and jointly optimized with image-based photometric constraints, enabling accurate and efficient state estimation in scenarios where conventional frame-based VO and SLAM struggle [2204.07640], [2305.08962].

## 1. Event Generation and Signal Model

EDS exploits the event generation principle of event cameras: each pixel asynchronously triggers an event $e_k = (u_k, t_k, p_k)$ whenever the local logarithmic intensity $L(u, t)$ changes by a predefined contrast threshold $C$:

\[
\Delta L(u_k, t_k) = L(u_k, t_k) - L(u_k, t_k - \Delta t_k) = p_k C,\quad p_k \in \{+1, -1\}.
\]

A first-order Taylor expansion under the optical flow assumption relates observed brightness increments to pixel velocities induced by rigid body motion:

\[
\Delta L(u) \approx -\nabla L(u) \cdot v(u) \Delta t,
\]

with image-plane velocity $v(u)$ parameterized by camera body velocity $(V, \omega)$ via:

\[
v(u) = J(u, Z) \begin{pmatrix} V \\ \omega \end{pmatrix},
\]
where $J(u, Z)$ encodes projection geometry and depth $Z$. The observed polarity-weighted event increments are accumulated with Gaussian temporal weighting to minimize motion blur:

\[
\Delta L_{\mathrm{meas}}(u) = \sum_{k=1}^{N_e} w_k\, p_k\, C\, \delta(u - u_k).
\]

A probabilistic generative model describes the likelihood of each event given the predicted increment:

\[
p\left(e_k \mid \Delta L(u_k)\right) = \Phi\left(\frac{p_k\Delta L(u_k) - C}{\sigma} \right),
\]
where $\Phi$ denotes the standard normal CDF and $\sigma$ captures sensor/event noise [2204.07640].

## 2. Direct Probabilistic Motion Formulation

For each active pixel (with high gradient and sufficient events), EDS defines a brightness increment residual:

\[
r_i(T, V) = \frac{\Delta \hat{L}(u_i; T, V)}{\|\Delta \hat{L}\|_2} - \frac{\Delta L_{\mathrm{meas}}(u_i)}{\|\Delta L_{\mathrm{meas}}\|_2},
\]

where $\Delta \hat{L}$ is the model-predicted brightness increment under a given motion hypothesis. The weighted least squares objective

\[
E(T, V) = \sum_i w_i\, \| r_i(T, V) \|^2
\]

is minimized over SE(3) increments via Gauss-Newton or Levenberg–Marquardt, with robust (Huber) per-pixel weights to attenuate outliers. Parameters are iteratively updated in a small-angle, Lie-algebra parametrization for numerical stability [2204.07640].

## 3. Sparse 3D Structure Selection and Parameterization

EDS implements a semi-dense mapping strategy: each keyframe selects a sparse set of high-gradient pixels, dividing the frame into tiles and retaining those with the highest Sobel gradient magnitude (typically top 10–15% per tile). The 3D structure is parameterized via inverse depth $\rho = 1/Z$, initialized by reprojecting depths from already-mapped keyframes and interpolated using inverse-depth kd-tree search. This approach provides a computationally efficient yet geometrically informative set of points for direct photometric and event-based optimization [2204.07640].

## 4. Global Optimization: Photometric Bundle Adjustment

For map and trajectory refinement, EDS employs a sliding-window photometric bundle adjustment (BA) over all active keyframes. The optimization jointly refines all poses $\{T_j\}$ and all inverse-depths $\{\rho_i\}$ by minimizing robust semi-dense photometric error:

\[
\min_{\{T_j, \rho_i\}} \sum_{i \in \mathcal{F}} \sum_{u \in \mathcal{P}_i} \sum_{j \in \mathrm{vis}(i,u)} \rho\left( L_j(\pi(T_j, X_i)) - \Delta L_{i \rightarrow j}(u) \right),
\]

where $X_i$ is the 3D point at pixel $u$ in keyframe $i$, $\pi$ is the projection, $L_j$ is the intensity in keyframe $j$, and $\Delta L_{i \rightarrow j}(u)$ is the reprojected event-induced increment. Huber costs mitigate the influence of degenerate correspondences or outliers. This BA typically operates on windows of $\sim$7 keyframes and is implemented using automatic differentiation in Ceres [2204.07640].

## 5. Algorithmic Workflow

The end-to-end EDS pipeline consists of the following loop:

1. **Initialization**: Coarse DSO-like bootstrapping on initial frames.
2. **Frontend tracking (events)**: For each incoming frame, collect event packets; accumulate $\Delta L_{\mathrm{meas}}$; perform incremental motion tracking via direct optimization.
3. **Keyframe management**: Trigger new keyframes when tracked coverage falls or rotation threshold is exceeded; select new sparse points and initialize depths.
4. **Backend mapping**: Perform sliding-window semi-dense photometric BA; update the map.
5. **Event batching**: Events are processed in overlapping packets (e.g., 20k events, 50% overlap) to optimize signal-to-noise ratio while minimizing blur [2204.07640].

Performance is maintained at $\approx$60 Hz in the frontend and $\approx$20 Hz for frames; sparse frames are sufficient due to the event stream bridging intervals (“blind time”).

## 6. Extension to RGB-D Data and Adaptive Event Surfaces

A variant EDS pipeline for robotics fuses RGB-D frames with events. An adaptive time surface (ATS) addresses TS “whiteout”/“blackout” by deploying pixel-wise, motion-adaptive decay rates:

\[
\tau(\mathbf{x}) = \max\left(\tau_u - \frac{1}{n} \sum_{i=1}^n (t - t_{\mathrm{last}}^{(i)}),\, \tau_\ell\right),
\]
\[
\mathcal{A}(\mathbf{x}, t) = 255 \exp\left(-\frac{t - t_{\mathrm{last}}(\mathbf{x})}{\tau(\mathbf{x})}\right).
\]
Pixel selection from the ATS then prioritizes spatially well-distributed, high-contrast, high-gradient regions. The full EDS objective jointly aligns RGB-D patch photometric errors and event-ATS patch errors. The final energy is

\[
\min_{\mathbf{T}_i, \{\rho\}} E_{\mathrm{rgb}}(\mathbf{T}_i, \{\rho\}) + \alpha E_{\mathrm{evt}}(\mathbf{T}_i, \{\rho\}) + \beta R(\{\rho\}),
\]

integrating both modalities with regularization [2305.08962].

## 7. Benchmark Performance and Application Scenarios

On indoor DAVIS-based benchmarks, monocular EDS achieves 1–2 cm RMS translational error and 1–2° RMS rotational error, outperforming monocular event-only VO (EVO, USLAM) and matching or slightly exceeding DSO and ORB-SLAM under normal frame rates. When frame rates are reduced from 20 Hz to 5 Hz, frame-only methods degrade or lose track, whereas EDS remains robust due to continuous event tracking. In robotics applications, EDS with ATS demonstrates ATE below 2 cm and competitive relative pose error (RPE) across high-dynamics tasks (e.g., bounding/backflipping quadruped robots, angular rates up to 510 °/s) on datasets where classical methods diverge [2204.07640], [2305.08962].

A summary of results is as follows:

| Dataset/Scenario             | EDS Translational Error / ATE | Best Baseline         |
|------------------------------|-------------------------------|-----------------------|
| Indoor DAVIS (bin/boxes/etc) | 1–2 cm rms / 1–2° rot error   | DSO/ORB-SLAM (similar)|
| MVSEC flying (robot)         | 1.2–1.75 cm ATE               | DEVO 2.4–7.1 cm       |
| Mini-Cheetah (bounding)      | 0.42 cm ATE                   | Baselines diverged    |

EDS thus enables low-power, high-dynamic-range, and robust odometry for AR/VR, nano-UAVs, legged robots, and other environments where traditional frame-based VO is challenged by lighting, speed, or frame rate constraints [2204.07640], [2305.08962].

Source: https://www.emergentmind.com/topics/event-aided-direct-sparse-odometry-eds