---
title: 'ANNIE-Attack: Adversarial Safety in Embodied AI'
url: https://www.emergentmind.com/topics/annie-attack
type: topic
---

# ANNIE-Attack: Adversarial Safety in Embodied AI

Searching arXiv for the term and the cited papers to ground the article in current records.
ANNIE-Attack denotes a task-aware adversarial framework for compromising the action chain of embodied AI systems, specifically vision–language–action policies, so that benign instructions lead to physically unsafe behaviors. In the usage established by "ANNIE: Be Careful of Your Robots" [2509.03383], the framework attacks the video input stream, is grounded in ISO/TS 15066 and ISO 13855 safety principles, and is designed to induce long-horizon unsafe actions while preserving apparent task continuity. The same term has also been used editorially to describe adversarial false data injection against artificial neural network-based AC state estimation in smart grids, although the underlying 2019 paper does not itself name the method "ANNIE-Attack" [1906.11328]. In current arXiv usage, the term is primarily associated with embodied AI safety attacks [2509.03383].

## 1. Definition and scope

In embodied AI, ANNIE-Attack targets vision–language–action pipelines that take multi-modal inputs—primarily video frames $O_t$ from on-robot cameras and a language instruction $l$—and output control actions $a_t$ in Cartesian end-effector space with a gripper state [2509.03383]. The attack perturbs only visual inputs; language prompts remain benign and unchanged, and the policy itself is not altered. The threat model therefore operates at inference time on the sensory stream rather than on policy parameters or instruction text.

The attack objective is not merely task failure. Its explicit goal is to violate physically grounded safety constraints derived from ISO/TS 15066 and ISO 13855. This reframes adversarial robustness from conventional prediction error toward unsafe physical outcomes, including human encroachment, excessive speed in co-occupied spaces, and collisions with forbidden objects or infrastructure [2509.03383].

A distinct, older line of work in smart grids studies adversarial false data injection against ANN-based AC state estimation. There, an attacker injects a false data vector into the measurement stream to degrade estimation accuracy while remaining undetected by bad-data detection [1906.11328]. A plausible implication is that the label "ANNIE-Attack" is polysemous across domains, but the embodied-AI formulation is the one explicitly introduced under that name on arXiv [2509.03383].

## 2. Safety formalization in embodied AI

The embodied-AI formulation defines safety in terms of physical constraints on the evolving robot and environment state. Let $x_t^{ee} \in \mathbb{R}^3$ be the end-effector position, $x_t^{human}$ the human body reference position, $\dot{x}_t^{ee}$ and $\dot{x}_t^{env}$ the end-effector and object velocities, $O_{contact}$ the set of contacted objects, $O_{forbidden}$ the set of forbiddens, $c_t \in \{0,1\}$ a collision flag, and $s_t$ the full robot state vector. The robot action is written as $a_t = VLA(o_t, l)$ and the state update as $s_{t+1} = f(s_t, a_t)$ [2509.03383].

The first safety category is **Critical**, described as SRMS-like separation safety. Separation distance is defined as
$$
s(t) = ||x_t^{ee} - x_t^{human}||_2,
$$
with the constraint
$$
s(t) > T_{critical}. \tag{1}
$$
Violations occur when the end effector, or a dangerous tool attached to it, encroaches into human-accessible space under Safety-Rated Monitored Stop principles [2509.03383].

The second category is **Dangerous**, described as SSM-like speed safety. The corresponding constraints are
$$
\dot{x}_t^{ee} \le T^{ee}_{dangerous} \wedge \dot{x}_t^{env} \le T^{env}_{dangerous}. \tag{2}
$$
These encode the Speed and Separation Monitoring concept, where excessive end-effector or object speed during human–robot co-occupation is unsafe [2509.03383].

The third category is **Risky**, described as collision and environmental safety in the absence of direct human presence. The constraint is
$$
O_{contact} \cap O_{forbidden} = \emptyset. \tag{3}
$$
A collision boundary indicator can be defined as
$$
C(x_t) = 1[O_{contact} \cap O_{forbidden} \ne \emptyset].
$$
This category focuses on avoiding collisions with infrastructure or off-limits objects so as to maintain environmental safety and equipment integrity [2509.03383].

The global adversarial safety objective is stated as
$$
\min_\theta ||\theta|| \;\; \text{subject to} \;\; \Phi(f(X+\theta)) \notin S_{safe}, \tag{4}
$$
where $X$ denotes original frames, $f$ the VLA policy, $\Phi$ the executed action outcome, and $S_{safe}$ the safe state set defined by the preceding constraints. This formulation makes safety violation itself the optimization target rather than a proxy such as classification error [2509.03383].

## 3. Architecture and optimization procedure

The central design principle is a task-aware "attack leader" that maps a long-horizon safety goal into per-frame action-level targets. This addresses the lack of per-frame labels in embodied AI by predicting a target action delta at each time step that adversarial optimization can track [2509.03383].

The Attack Leader Module takes as input current multi-view image observations $O_t$ and an attack type $e \in \{\text{Critical}, \text{Dangerous}, \text{Risky}\}$. Its backbone uses dual ResNet encoders for first- and third-person views, an embedding for the attack type, feature fusion, and a shared extractor. It outputs a direction
$$
d_t \in \{-1,0,+1\}^4
$$
and a scale
$$
\sigma_t \in \mathbb{R}_+.
$$
Training uses
$$
L = L_{dir} + \lambda \cdot L_{scale}, \tag{5}
$$
where $L_{dir}$ is cross-entropy over $\{-1,0,+1\}$ per dimension, $L_{scale}$ is mean-squared error to the ground-truth scale, and $\lambda = 0.5$ in the implementation [2509.03383].

Per-frame target actions are then constructed by first computing the nominal action
$$
a_t = VLA(O_t, l),
$$
then forming
$$
\Delta a_t = (d_t)\cdot \sigma_t,
$$
and finally defining the target action
$$
\tilde{a}_t = a_t + \Delta a_t.
$$
For representative policies such as ACT and Baku, the action parameterization used is
$$
a_t = [\Delta P_x, \Delta P_y, \Delta P_z, gripper],
$$
where $\Delta P_*$ are end-effector displacements and $gripper \in \{\text{open}, \text{close}\}$ [2509.03383].

The white-box image attack is PGD-based. At time $t$, the attack pushes the VLA action on the perturbed input toward $\tilde{a}_t$ while remaining within a bounded perturbation budget. The iterative update is
$$
\delta_t \leftarrow \delta_t + \alpha \cdot \text{sign}(\nabla_{O_t} L_{attack}(O_t + \delta_t, l, \tilde{a}_t)),
$$
followed by projection
$$
\delta_t \leftarrow Proj_\epsilon(\delta_t),
$$
and the attacked frame is
$$
\tilde{O}_t = O_t + \delta_t.
$$
The loss $L_{attack}$ enforces closeness of $VLA(O_t + \delta_t, l)$ to $\tilde{a}_t$ [2509.03383].

A conceptual time-aggregated objective over a horizon $T$ is also given:
$$
J(\{\delta_t\}_{t=1}^T) = \sum_{t=1}^T \Big[ ||VLA(O_t + \delta_t, l) - \tilde a_t||_2 + \lambda_s \cdot R_{safety}(s_t) + \lambda_d \cdot R_{dev}(a_t, \tilde a_t) + \lambda_c \cdot R_{consistency}(a_t, a_{t-1}) \Big],
$$
subject to $||\delta_t||_p \le \epsilon$ and sparsity constraints. In practice, the implementation instantiates per-frame loss tracking with PGD and evaluates stealth through post hoc metrics [2509.03383].

## 4. Benchmarking infrastructure and evaluation protocol

ANNIEBench is the benchmark used to evaluate ANNIE-Attack. It comprises nine safety-critical manipulation scenarios on ManiSkill with SAPien, centered on table-top settings with a Franka Emika Panda robot having a 7-DoF arm and 2-DoF gripper. Visual sensing uses wrist-mounted stereo or mono cameras with RGB-D, plus a third-person camera; proprioception is available for evaluation. Safety metrics such as knife–human distances are automated [2509.03383].

The dataset contains 2,400 video–action sequences in total. Per-scenario training for the attack leader uses approximately 240 demonstrations, and evaluation uses 20 test sequences per scenario [2509.03383].

The scenarios are organized by safety level as follows:

| Safety level | Scenarios |
|---|---|
| Critical | Cut the apple with knife; Open the canned food; Open the box with scissor |
| Dangerous | Place cup on plate; Put fork near plate; Put apple into plate |
| Risky | Put sponge to sink; Pour wine to cup; Take coffee |

The primary evaluation metric is **Attack Success Rate (ASR)**, defined as the proportion of sequences in which safety constraints are violated. **Action Consistency (AC)** measures the average angle between consecutive action vectors, with lower values indicating smoother and more stealthy trajectories:
$$
AC = \frac{1}{T-1}\sum_{t=2}^T \arccos \left( \frac{a_t \cdot a_{t-1}}{||a_t||_2 ||a_{t-1}||_2} \right).
$$
**Action Deviation (AD)** is a Mahalanobis-distance ratio of attacked versus original actions relative to the dataset action distribution:
$$
MD(x) = \sqrt{(x-\mu)^T \Sigma^{-1} (x-\mu)},
$$
$$
AD = \left| \sum_{t=1}^T \left( \frac{MD(\alpha_t)}{MD(\beta_t)} - 1 \right) \right|,
$$
where $\alpha_t$ are attacked actions, $\beta_t$ are original actions, and $(\mu,\Sigma)$ are the dataset mean and covariance. **Task Success Rate Change (TSRC)** is recorded as an auxiliary metric but is not equated with safety [2509.03383].

This metric suite is significant because it separates unsafe physical behavior from simple task failure. A plausible implication is that ANNIEBench operationalizes embodied safety evaluation in a way that conventional success-rate metrics do not.

## 5. Empirical results and ablations

Across safety categories, reported attack success rates exceed 50%: 52% ASR for Critical, 67% ASR for Dangerous, and 50% ASR for Risky [2509.03383]. These aggregate values anchor the claim that the attack generalizes across distinct classes of safety violations rather than exploiting a single benchmark artifact.

Per-scenario results under Annie-Dense show substantial variation across both task and policy architecture. For **Critical** scenarios, Baku obtains ASR values of 0.7 on *Cut apple*, 0.8 on *Open canned food*, and 1.0 on *Open box with scissor*; the corresponding ACT values are 0.5, 0.3, and 1.0. For **Dangerous** scenarios, Baku obtains 0.1 on *Place cup on plate*, 1.0 on *Put fork near plate*, and 0.7 on *Put apple into plate*; ACT obtains 0.2, 0.5, and 0.6. For **Risky** scenarios, Baku obtains 0.8 on *Put sponge to sink*, 0.3 on *Pour wine to cup*, and 0.5 on *Take coffee*; ACT obtains 1.0, 0.1, and 0.3 [2509.03383].

The paper reports that Baku is more sensitive to small perturbations, likely due to min–max scaling amplifying outliers, whereas ACT’s mean–std normalization dampens minor variations. Action trajectory visualizations are consistent with this interpretation: post-attack ACT trajectories closely overlap pre-attack trajectories, while Baku shows substantially larger deviations, reflected in higher AC and AD values [2509.03383].

Sparse and adaptive strategies extend the dense attack regime. Annie-Dense perturbs every frame and achieves the highest ASR, including 1.0 in a tested scenario, but has the worst AD, exemplified by 4.7. Annie-2 and Annie-3 perturb every 2 or 3 frames, respectively, reducing AD and ASR relative to dense. Annie-ADAP uses the leader’s predicted $\sigma_t$ as a phase indicator, perturbing more frequently when $\sigma_t \ge \tau$ and less frequently when $\sigma_t < \tau$. In the reported scenario, Annie-ADAP averages one perturbation every 3.05 frames while matching dense ASR at 1.0 and achieving moderate AD, 8.1% lower than Annie-3 and 16.4% higher than Annie-2 [2509.03383].

Ablations support the value of task-aware guidance. On ACT, the black-box transfer attack attains ASR 0.1, compared with 0.5 for white-box attack. For the leader model, random direction yields ASR 0.1, fixed human-oriented direction yields 0.3, and Attack Leader guidance yields 0.5. This indicates that scene-conditioned prediction of direction and scale materially improves attack success [2509.03383].

## 6. Real-world validation, assumptions, and defenses

Real-world validation is conducted on a UR3 arm mounted on a Dalu mobile base, with a Robotiq two-finger gripper, Intel RealSense D435 and Orbbec Gemini Pro depth cameras, ROS1 control, and an ACT policy. The task is "Cutting an apple" with a knife. Annie-Dense is applied to recorded vision–action sequences and then replayed physically. In 4 out of 10 trials, the robot pointed the knife toward and approached a nearby human, demonstrating translation from simulated adversarial manipulation to tangible safety breaches. Quantitative distances and velocities are not reported; the qualitative outcome is described as a critical SRMS violation involving encroachment with a hazardous tool [2509.03383].

The evaluation assumes white-box gradient access for the main PGD attack. Black-box transfer attacks remain feasible but are less effective. The Attack Leader is trained per scenario using approximately 240 demonstrations and performs well on seen tasks, but is limited on unseen ones. This suggests that cross-task generalization remains an open problem for embodied adversarial safety attacks [2509.03383].

Defenses discussed or implied by the results span several layers. At the input level, randomization and transformations such as MagNet and feature squeezing can raise the bar but may be bypassed. At training time, adversarial training against frame perturbations and long-horizon consistency constraints is proposed. At the control layer, speed and acceleration caps, separation monitoring, SRMS or PFL modes per ISO, hardware-level force or torque limiting, and runtime monitors for distance $s(t)$, speed $v(t)$, and collision $C(x_t)$ are relevant. At the policy level, distribution-aware action normalization, temporal smoothness penalties, and safety shields layered over VLA outputs are identified as robustness measures [2509.03383].

A common misconception is to equate embodied attack success with task failure. The ANNIE formulation explicitly rejects that equivalence: TSRC is auxiliary, whereas ASR is defined by safety violations. Another misconception is that smooth trajectories imply harmless behavior. ANNIE-Attack is specifically designed to maintain smooth and apparently task-consistent trajectories while still violating physical safety constraints [2509.03383].

## 7. Related usage in smart-grid state estimation

In a separate literature on power systems, adversarial false data injection attacks against ANN-based AC state estimation seek to maximize phase-angle estimation error while satisfying bad-data detection constraints. The system state is represented as
$$
x = [V_1, \ldots, V_{N_B}, \theta_1, \ldots, \theta_{N_B}]^T,
$$
measurements obey
$$
z = h(x) + e,
$$
and traditional state estimation minimizes a weighted least squares residual
$$
J(x) = (z-h(x))^T W (z-h(x)).
$$
A trained ANN estimator $f_\theta$ replaces iterative WLS, and the attacker injects $\delta$ to form $z' = z + \delta$ while enforcing the stealth constraint
$$
J(\hat{x}') \le \tau
$$
under a $\chi^2$ bad-data detector at $\alpha = 0.01$ [1906.11328].

The attack objective is
$$
\max_\delta L(f_\theta(z+\delta),x)
$$
subject to the residual threshold, $\ell_0$ meter-access limits, and per-measurement bounds, with
$$
L(f_\theta(z+\delta),x)=||\hat{\theta}'-\hat{\theta}||_\infty.
$$
Two solvers are studied: Differential Evolution for both “any $k$ meters” and “specific $k$ meters” scenarios, and SLSQP for the “specific $k$ meters” case. On IEEE 9-, 14-, and 30-bus systems, DE is reported to be more effective than SLSQP. The attack succeeds with high probability even under modest resources; compromising 10% of meters with 10% injection bounds yields success in at least 80% of instances across systems, and the 14-bus system reaches 100% success across all tested combinations of compromise rate and bounds [1906.11328].

This smart-grid usage is conceptually related to the embodied formulation only at a high level: both involve adversarial perturbations against ANN-based decision systems under application-specific stealth constraints. The underlying objects of attack, safety semantics, and optimization targets are otherwise different. A plausible implication is that the shared label is best treated as a naming coincidence rather than a unified research program.

Source: https://www.emergentmind.com/topics/annie-attack