---
title: 'ReActEval: Hierarchical Drone Control'
url: https://www.emergentmind.com/topics/reacteval
type: topic
---

# ReActEval: Hierarchical Drone Control

ReActEval is a reasoning methodology for worker agents in a hierarchical multi-agent framework for autonomous drone-based visual inspection in indoor industrial settings. It is introduced as the low-level execution loop beneath a head agent that performs high-level planning and task allocation. The defining feature of ReActEval is an explicit post-action evaluation stage: given a per-drone natural-language plan and an expected outcome, a worker iterates Reason $\rightarrow$ Act $\rightarrow$ Evaluate, updating thread-local history until `end_flag` indicates task completion or `max_iters` is reached. In the reported system, this design is used for tasks ranging from simple navigation to locating and reading a pressure gauge, and its empirical behavior is strongly dependent on model capability and task complexity [2510.00259].

## 1. Definition and architectural placement

ReActEval operates strictly at the worker-agent level inside a larger hierarchical agentic framework. The head agent receives the user request, maintains session-level memory across requests, decomposes the request into per-drone subtasks, and then invokes `ReActEval(task)` for each assigned worker. Each worker controls a single drone and executes only its assigned subtask in a thread-local loop; after task completion, the worker’s thread history is reset rather than accumulated as persistent session context [2510.00259].

The interface between the head agent and the worker is natural-language but structured. For each drone, the head agent outputs a dictionary containing a step-by-step `plan`, an `expected_outcome`, an `end_flag`, and a global `response_to_user`. The paper’s example assigns one drone the plan `"1. Takeoff.\n2. Move to (5, 0, 1)."` with the expected outcome `"Drone 1 is located at (5, 0, 1)."`. This representation is central: ReActEval does not invent the overall mission. It consumes a plan already produced by the head agent and then reasons about how to execute that plan step by step [2510.00259].

This placement matters conceptually. ReActEval is not the whole multi-agent framework, and it is not a high-level planner. It is the worker-agent execution policy that supplies low-level autonomy, self-correction, and completion checks once a subtask has already been assigned. A frequent misunderstanding is to treat it as a generic ReAct replacement for all layers of an agent stack; the reported system is more specific than that.

## 2. Control loop and formal structure

The worker-side algorithm is given explicitly. `Initialize(task)` creates the initial history and task context. The loop then repeats while `\neg task.complete \land iteration < max_iters`:

1. `reasoning \gets Reason(task, history)`
2. `action \gets Act(reasoning)`
3. `evaluation \gets Evaluate(task, history, action)`
4. `task.complete \gets evaluation.end_flag`
5. `history \gets UpdateHistory(reasoning, action, evaluation)`

The return value is the final thread history [2510.00259].

Operationally, the paper describes this as a “three-step ‘Reason-Act-Evaluate’ process,” but the broader system is a full plan/reason/act/evaluate cycle because the plan is supplied by the head agent before the worker loop starts. The completion condition is not inferred implicitly from free-form reasoning; it is represented explicitly through `evaluation.end_flag`, which controls termination. This makes the evaluation stage not merely diagnostic but causal in the control flow.

Relative to the paper’s baselines, the difference is precise. The ReAct baseline removes the separate post-action evaluation phase and places `end_flag` inside Reason; the Act baseline removes both Reason and Evaluate, directly producing actions from history. The novelty claimed for ReActEval is therefore not just additional chain-of-thought, but an explicit post-action validation step grounded in the updated drone state and task history [2510.00259].

A useful comparison is with Focused ReAct for question answering, which addresses context drift and action loops through reiteration of the original question and early stop on duplicate actions, but does not introduce a distinct post-action evaluation phase [2410.10779]. ReActEval’s modification is structurally different: it inserts evaluation as a first-class stage in the worker loop.

## 3. Reason, action formulation, and state-grounded evaluation

The Reason stage conditions on the current drone state, the head agent’s `plan`, the `expected_outcome`, and the worker’s thread history. The paper specifies that the drone state includes position coordinates, heading, and other relevant parameters, while the thread history may include previous actions, previous evaluations, and prior `next_steps_notes`. The output is a structured dictionary with fields corresponding to `reasoning` and `intended_action`. An example given is a navigation rationale that decomposes motion to a target point into axis-aligned steps and selects the largest-distance move first [2510.00259].

The Act stage translates that intended action into an executable tool call. The available capabilities include drone functions such as `Takeoff`, `Land`, `Move`, `Rotate`, `Move gimbal`, and `Capture image`, as well as model-side functions such as `Analyze image` and `Analyze gauges`. The system is explicitly tool-agnostic and is described as being able to invoke VLMs, YOLO, or other custom tools. The action prompt requires a JSON object representing a function call, for example:
```json
{"function_call": "move", "parameters": { "direction": "forward", "distance": 10}}
```
The instruction is strict: the worker “MUST use a function call” [2510.00259].

The Evaluate stage is the distinctive component. It consumes the overall plan, the expected outcome, the recently executed action, and thread history; the appendix refines this to Overall Plan, Thread History, and Drone State After Action. Its output contains `evaluation_summary`, `end_flag`, and `next_steps_notes`. The instructions are to assess whether the most recent action was successful in progressing the plan, use the drone state to confirm the action’s outcome, set `end_flag` to true if and only if all steps in the plan are finished, and provide guidance for the next reasoning step. This makes evaluation both a feedback channel and a constrained replanning interface. The paper repeatedly characterizes it as a “feedback and/or replanning stage” [2510.00259].

The state representation is explicit enough to support this grounding. In simulation, the system tracks 3D Cartesian coordinates, heading, gimbal angle, and last executed command. The update rules are deterministic: `takeoff` sets flight status to True and altitude to 1.0 m; `land` sets flight status to False and altitude to 0.0 m; `move` updates coordinates using trigonometric functions based on current heading plus the specified direction and distance; `rotate` updates heading by adding the rotation angle; and `move_gimbal` updates gimbal angle within 0–90 degrees. The next evaluation and reasoning steps therefore operate on a concretely updated state rather than only on textual traces [2510.00259].

## 4. Experimental setting and task regime

The reported evaluation is conducted in a simulated indoor industrial inspection environment with two worker agents, each controlling one drone. Drone 1 starts at $(0,0,0)$ and drone 2 at $(0,2,0)$. Although the framework is described as supporting arbitrary numbers of drones, the experiments use exactly two for simplicity. The task set is organized into three complexity levels: easy, medium, and hard [2510.00259].

Easy tasks are one- or two-step commands such as taking off both drones, moving drone 1 forward 2 m, taking a picture, rotating 180 degrees, or reporting responsibilities or state. Medium tasks require multi-step explicit sequences, such as flying both drones in a square of side 3 m, executing different trajectories for the two drones, or moving one drone, taking a picture, and describing it. Hard tasks are open-ended inspection tasks, including using both drones to capture images of each corner of a 10 m $\times$ 2 m room, navigating drone 2 to a pressure gauge at $(4m, 18m, 6m)$ and returning its status, and having one drone describe an object from the left side while the other describes it from the right side [2510.00259].

The compared models are GPT-4.1 Nano, GPT-4.1, o4-mini, and o3. ReActEval is evaluated against two baselines, ReAct and Act. The scoring protocol is sequential and manually specified. For easy and medium tasks, one point is awarded per correctly executed function call, but only until the first incorrect call; subsequent steps receive no credit. The easy tasks total 14 actions and the medium tasks total 36. Hard tasks are scored by manually decomposed subtasks because multiple valid action sequences may exist; for example, imaging four corners is worth four points irrespective of the exact path. Since simulation assumes correctly specified function calls always succeed, failures are interpreted as reasoning or tool-selection failures rather than execution noise [2510.00259].

Under this protocol, ReActEval’s reported scores are: GPT-4.1 Nano, Easy $14/14$, Medium $13/36$, Hard $2/13$, Overall $0.460$; GPT-4.1, Easy $13/14$, Medium $34/36$, Hard $4/13$, Overall $0.810$; o4-mini, Easy $14/14$, Medium $34/36$, Hard $6/13$, Overall $0.857$; and o3, Easy $13/14$, Medium $34/36$, Hard $10/13$, Overall $0.905$ [2510.00259].

## 5. Comparative performance and failure modes

The main empirical conclusion is conditional rather than universal. ReActEval is not uniformly best; its benefit increases with model capability and task complexity. With GPT-4.1 Nano, ReActEval is worse than ReAct on medium tasks ($13/36$ versus $18/36$) and worse than Act on medium tasks ($13/36$ versus $21/36$). With stronger models, the ranking reverses. For GPT-4.1, ReActEval exceeds ReAct on medium ($34/36$ versus $30/36$) and hard ($4/13$ versus $2/13$) tasks. For o4-mini, it exceeds ReAct on medium ($34/36$ versus $29/36$) and hard ($6/13$ versus $4/13$). For o3, it exceeds ReAct on medium ($34/36$ versus $32/36$) and hard ($10/13$ versus $6/13$), and also exceeds Act on hard tasks ($10/13$ versus $5/13$) [2510.00259].

The paper therefore emphasizes a performance reversal with model capability. The extra reasoning burden introduced by explicit evaluation can help stronger models structure execution and correct errors, but it can create additional failure opportunities for weaker models. This is visible in the provided qualitative trace for the medium task “Drone 1, fly forward 4m, take a picture and describe what you see.” GPT-4.1 Nano incorrectly reasons that moving forward 4 m should lead to $(4,0,0)$ rather than $(0,4,0)$. After the simulated move correctly places the drone at $(0,4,1)$, the evaluation stage reasons against the wrong target and recommends moving right 4 m, reinforcing rather than correcting the spatial error. By contrast, o4-mini executes the intended sequence successfully: takeoff, move forward 4 m to $(0,4,1)$, capture image, analyze image, and land [2510.00259].

The identified failure modes are threefold: incorrect function calls, early stopping, and head-agent failure. The paper reports that ReActEval reduces incorrect and unnecessarily repeated function calls relative to ReAct and Act, and argues that the evaluation step “helps prevent errors like executing functions in wrong order or failing to recover from mistakes.” However, early stopping remains common across all methods. The pressure-gauge task provides a representative example: a drone may take off, move up 5 m, forward 16 m, right 4 m, and capture an image, yet stop before actually analyzing the gauge and returning its status. Head-agent failures are comparatively rare; only four tasks across all experiments are attributed to the head agent, indicating that the measurable effect of ReActEval is concentrated at the worker-execution level [2510.00259].

Execution-time overhead is present but not dominant relative to model size. On medium tasks, reported times are 5.79 s for GPT-4.1 Nano, 7.22 s for GPT-4.1, 21.38 s for o4-mini, and 30.60 s for o3 under ReActEval, compared with 5.97 s, 7.45 s, 21.61 s, and 40.08 s for ReAct, and 5.78 s, 7.32 s, 19.83 s, and 27.90 s for Act. Since simulation treats actions as instantaneous, these measurements exclude physical execution latency [2510.00259].

## 6. Limitations, interpretation, and research directions

Several misconceptions are directly addressed by the results. ReActEval is not a universally superior alternative to simpler execution loops; the paper explicitly shows regimes in which it underperforms. It is also not an external safety verifier. The claimed alignment benefit comes from post-action mission-progress evaluation grounded in state and expected outcome, not from a separate collision checker, explicit safety module, or formal verifier [2510.00259].

The current evidence is limited by the evaluation regime. The main experiments are in simulation, so the method is not validated under sensor noise, communication delays, actuation uncertainty, or real-time safety constraints. The paper notes that preliminary real-world tests suggest substantially greater difficulty, especially because models struggle to translate high-level goals into precise low-level command sequences. Poor spatial reasoning remains a core weakness, early stopping persists across methods, and scalability is demonstrated only for two drones even though the architecture is claimed to support arbitrary numbers [2510.00259].

These limitations shape the paper’s future directions. The proposed extensions include hybrid systems that combine LLM-based high-level planning with traditional low-level controllers, hybrid-capability agents in which powerful models handle Reason and Evaluate while smaller models handle Act, adaptive systems that choose between Act and ReActEval based on task complexity, fine-tuning smaller models on drone-control data, and broader tool suites for safety-critical and dynamic environments. A plausible implication is that ReActEval is best understood not as a final solution for physical autonomy, but as a specific control-loop modification that exposes where explicit evaluation helps and where model-level grounding remains insufficient.

In that sense, ReActEval occupies a narrow but technically clear position in the ReAct family. Standard ReAct-style execution alternates reasoning and acting; Focused ReAct for question answering adds reiteration and early stop to preserve focus and avoid loops [2410.10779]. ReActEval instead makes evaluation a persistent, state-grounded stage in the worker loop, with `end_flag` and `next_steps_notes` as explicit control variables. Its significance lies less in a new foundation model or reward function than in a concrete claim about physical-agent execution: after each action, the agent should explicitly check whether the updated world state still advances the assigned plan and whether the task is semantically complete [2510.00259].

Source: https://www.emergentmind.com/topics/reacteval