Papers
Topics
Authors
Recent
Search
2000 character limit reached

Instrumental Reliance in AI Workflows

Updated 7 July 2026
  • Instrumental reliance is a behavioral measure of dependence on AI, quantifying how tasks and cognitive effort are delegated between humans and systems.
  • It is operationalized via frameworks like decision-theoretic reliance, offloading scores, and forced behavioral choices to capture calibration and accuracy.
  • Proper calibration of reliance is crucial; over- or underreliance can impair decision quality or miss performance gains in various real-world contexts.

Instrumental reliance denotes a behaviorally specified form of dependence on an instrument—most often an AI system—within a task. Recent work treats it not as abstract trust but as the way cognitive effort, decision authority, or workflow steps are distributed between human and system. In AI-advised decision making, it can be formalized as the probability of adopting the AI recommendation when human and AI disagree; in workflow analysis, it can be quantified as the fraction of cognitive effort offloaded to a tool through a counterfactual human-only workflow; in educational writing studies, it can denote bounded use for mechanical tasks without delegation of ideational content (Guo et al., 2024, Padmakumar et al., 28 May 2026, Hossain, 27 Jun 2026). Across these formulations, the central problem is calibration: overreliance can worsen decision quality, underreliance can forgo gains, and appropriate reliance depends on task goals, user capabilities, and interaction context (Hunter et al., 2024, Ferino et al., 12 Apr 2026, Eckhardt et al., 2024).

1. Conceptual scope and definitional boundaries

A consistent theme in the literature is the distinction between reliance and trust. Trust is treated as attitudinal, whereas reliance is treated as behavioral: what a person actually does when AI advice or AI-generated output is present. The survey literature places this inside a sociotechnical system comprising a technical component, a social component, their interaction, and the surrounding environment, and argues that reliance emerges from the fit between user, advice, task, and setting rather than from model properties alone (Eckhardt et al., 2024).

Within that broad frame, different papers narrow the construct for specific purposes. The decision-theoretic literature defines reliance at the moment of choice between human and AI recommendations, explicitly separating reliance from belief formation and from the user’s ability to differentiate signals (Guo et al., 2024). Workflow-centered work instead treats reliance as displaced human effort, asking how many workflow steps the tool removed from the human path rather than whether a specific suggestion was accepted (Padmakumar et al., 28 May 2026). Organizational risk-management work adopts a deliberately narrower definition of over-reliance: a user is over-reliant when they attempt to follow AI-generated advice for a problem they would solve more effectively on their own (Hunter et al., 2024). In software engineering, appropriate reliance is framed jointly with control, spanning a range from self-reliance to full reliance on AI and from self-control to losing control over AI (Ferino et al., 12 Apr 2026).

These definitions are not interchangeable. Some are centered on advisory choice, some on workflow restructuring, and some on task-bounded support. A plausible implication is that “instrumental reliance” is best understood as a family of operational constructs tied together by a common concern with how AI changes human action, judgment, and effort allocation, rather than as a single universally fixed metric.

2. Formalizations and operational criteria

Several recent frameworks define instrumental reliance in mathematically explicit terms.

Formulation Operational definition Emphasis
Decision-theoretic reliance Pr[a=AIAIH]\Pr[a = AI \mid AI \neq H] Adoption of AI when human and AI disagree
Offloading score mnm\frac{m-n}{m} Fraction of counterfactual steps saved by AI
Pred-RC reliance rate P(di=AIhistory,xi,ci)P(d_i = \mathrm{AI} \mid \text{history}, x_i, c_i) Delegation to AI under optional cues
REL-A.I. reliance rate Proportion of “Use Response” choices In situ reliance on linguistic confidence

In the decision-theoretic framework, the original decision problem is converted into a derived binary-adoption task with action space a{0=human,1=AI}a \in \{0=\text{human},1=\text{AI}\}. Reliance is defined only when the human and AI recommendations differ, since agreement removes the adoption choice. Appropriate reliance is then the reliance level of a fully rational decision-maker in the same task, and behavioral under-reliance or over-reliance is defined by whether the observed reliance level falls below or above that rational benchmark (Guo et al., 2024).

The offloading-score framework models an observed workflow as a sequence W={w1,,wn}W=\{w_1,\dots,w_n\} and replaces each AI-assisted step with a human-only counterfactual sub-sequence wi={wi,1,,wi,ki}w'_i=\{w'_{i,1},\dots,w'_{i,k_i}\}, yielding a counterfactual workflow W={w1,,wm}W'=\{w'_1,\dots,w'_m\}. The score is then mnm\frac{m-n}{m}, with $0$ meaning no steps were saved by AI and $1$ meaning the AI replaced essentially all of the counterfactual workflow. This formulation is explicitly process-oriented: it measures displaced effort rather than output adoption alone (Padmakumar et al., 28 May 2026).

Pred-RC formalizes reliance as the probability that a human assigns the current task to the AI, with and without a reliance calibration cue. For task mnm\frac{m-n}{m}0, the model estimates mnm\frac{m-n}{m}1 and compares the predicted mismatch between reliance and AI success probability under cue-present and cue-absent conditions. Cue provision is then made selective rather than continuous (Fukuchi et al., 2023).

REL-A.I. operationalizes reliance through a forced behavioral choice. Participants see a question and an epistemic marker such as “I’m certain it’s …” or “Maybe it’s …,” but not the answer itself, and choose either “Use Response” or “Look up myself later.” The resulting reliance rate is the proportion of trials on which participants choose to rely on the agent under a given expression and context (Zhou et al., 2024).

3. Measurement regimes and validity evidence

The literature does not use a single measurement standard. The survey on AI reliance identifies agreement percentage, switch percentage, weight of advice, self-reports, accuracy proxies, delegation behavior, manual override, eye gaze, and log-based measures as competing operationalizations. Agreement percentage is easy to compute but can conflate coincidental agreement with causal influence; switch percentage is closer to causal reliance in two-stage designs but requires an initial human judgment; weight of advice provides a continuous influence measure in estimation tasks, using

mnm\frac{m-n}{m}2

The same survey argues that single-stage designs tend to increase recall of reliance instances but reduce precision, whereas two-stage designs increase precision but may miss some genuine reliance (Eckhardt et al., 2024).

Among recent measures, offloading score has unusually extensive validity analysis. For content validity, more than mnm\frac{m-n}{m}3 of sampled counterfactual steps were rated plausible at least mnm\frac{m-n}{m}4 on a Likert scale, and synthetic counterfactual workflows were significantly closer to human workflows from the same task than to workflows from other tasks, with mnm\frac{m-n}{m}5 by Wilcoxon and mnm\frac{m-n}{m}6 by permutation test. For construct validity, removing AI-assisted steps in mnm\frac{m-n}{m}7, mnm\frac{m-n}{m}8, and mnm\frac{m-n}{m}9 perturbation settings reduced mean offloading score by P(di=AIhistory,xi,ci)P(d_i = \mathrm{AI} \mid \text{history}, x_i, c_i)0, P(di=AIhistory,xi,ci)P(d_i = \mathrm{AI} \mid \text{history}, x_i, c_i)1, and P(di=AIhistory,xi,ci)P(d_i = \mathrm{AI} \mid \text{history}, x_i, c_i)2, respectively, while reruns with different reasoning settings, a different model, or paraphrased workflow steps produced little change. The same paper reports that an LLM-as-judge matched human annotations at about P(di=AIhistory,xi,ci)P(d_i = \mathrm{AI} \mid \text{history}, x_i, c_i)3 for cognitive process labels and P(di=AIhistory,xi,ci)P(d_i = \mathrm{AI} \mid \text{history}, x_i, c_i)4 for output-use labels (Padmakumar et al., 28 May 2026).

Criterion validity is especially prominent in the controlled study underlying offloading score. In a sample of P(di=AIhistory,xi,ci)P(d_i = \mathrm{AI} \mid \text{history}, x_i, c_i)5 experienced freelance developers completing web-app coding tasks, participants under a short P(di=AIhistory,xi,ci)P(d_i = \mathrm{AI} \mid \text{history}, x_i, c_i)6-hour deadline had an offloading score of P(di=AIhistory,xi,ci)P(d_i = \mathrm{AI} \mid \text{history}, x_i, c_i)7, versus P(di=AIhistory,xi,ci)P(d_i = \mathrm{AI} \mid \text{history}, x_i, c_i)8 under a long P(di=AIhistory,xi,ci)P(d_i = \mathrm{AI} \mid \text{history}, x_i, c_i)9-hour deadline, a a{0=human,1=AI}a \in \{0=\text{human},1=\text{AI}\}0 increase with a{0=human,1=AI}a \in \{0=\text{human},1=\text{AI}\}1. By contrast, baseline measures did not reliably distinguish conditions: fraction of AI code retained was a{0=human,1=AI}a \in \{0=\text{human},1=\text{AI}\}2 versus a{0=human,1=AI}a \in \{0=\text{human},1=\text{AI}\}3 with a{0=human,1=AI}a \in \{0=\text{human},1=\text{AI}\}4, and self-reported cognitive load was a{0=human,1=AI}a \in \{0=\text{human},1=\text{AI}\}5 versus a{0=human,1=AI}a \in \{0=\text{human},1=\text{AI}\}6 with a{0=human,1=AI}a \in \{0=\text{human},1=\text{AI}\}7. The offloading score also had the strongest correlation with the condition label, a{0=human,1=AI}a \in \{0=\text{human},1=\text{AI}\}8 (Padmakumar et al., 28 May 2026).

Interaction-centered evaluation provides a different kind of measurement evidence. REL-A.I. uses a self-incentivized scoring rule: participants receive a{0=human,1=AI}a \in \{0=\text{human},1=\text{AI}\}9 if they rely and the system is correct, W={w1,,wn}W=\{w_1,\dots,w_n\}0 if they rely and the system is wrong, and W={w1,,wn}W=\{w_1,\dots,w_n\}1 if they choose to look it up themselves. This makes correct reliance instrumentally rewarded and incorrect reliance instrumentally penalized within the task itself (Zhou et al., 2024).

4. Determinants and interactional dynamics

A central empirical result is that reliance is highly context-sensitive. Under time pressure, offloading score detects more reliance and also reveals how that reliance changes behavior. In the short-deadline condition, users showed more execution-oriented interactions (W={w1,,wn}W=\{w_1,\dots,w_n\}2 versus W={w1,,wn}W=\{w_1,\dots,w_n\}3) and more direct reuse of AI outputs (W={w1,,wn}W=\{w_1,\dots,w_n\}4 versus W={w1,,wn}W=\{w_1,\dots,w_n\}5), while long-deadline users rejected outputs more often (W={w1,,wn}W=\{w_1,\dots,w_n\}6 versus W={w1,,wn}W=\{w_1,\dots,w_n\}7) and used the model more for planning (W={w1,,wn}W=\{w_1,\dots,w_n\}8 versus W={w1,,wn}W=\{w_1,\dots,w_n\}9). The paper interprets these as behavioral indicators of deeper delegation and less iterative scrutiny (Padmakumar et al., 28 May 2026).

REL-A.I. shows that reliance is also shaped by interactional context, prior exposure, and social cues. The same moderate-certainty expression was relied upon wi={wi,1,,wi,ki}w'_i=\{w'_{i,1},\dots,w'_{i,k_i}\}0 in a generally confident system and wi={wi,1,,wi,ki}w'_i=\{w'_{i,1},\dots,w'_{i,k_i}\}1 in a generally unconfident system. The paper further reports that people rely wi={wi,1,,wi,ki}w'_i=\{w'_{i,1},\dots,w'_{i,k_i}\}2 more on LLMs when responding to questions involving calculations and wi={wi,1,,wi,ki}w'_i=\{w'_{i,1},\dots,w'_{i,k_i}\}3 more on LLMs that are perceived as more competent. Warm greetings such as “I’m happy to help!” increased perceived warmth, and aggregated across systems, reliance rose from wi={wi,1,,wi,ki}w'_i=\{w'_{i,1},\dots,w'_{i,k_i}\}4 to wi={wi,1,,wi,ki}w'_i=\{w'_{i,1},\dots,w'_{i,k_i}\}5 as perceived warmth increased, with Pearson’s wi={wi,1,,wi,ki}w'_i=\{w'_{i,1},\dots,w'_{i,k_i}\}6 (Zhou et al., 2024).

Selective cueing work treats reliance as calibratable rather than fixed. Pred-RC predicts how a reliance calibration cue would change the probability of delegating a task to AI, then provides the cue only when it is expected to improve the match between reliance and AI success probability. In evaluation, ANCOVA showed significant effects of number of cues, wi={wi,1,,wi,ki}w'_i=\{w'_{i,1},\dots,w'_{i,k_i}\}7, RCC selection method, wi={wi,1,,wi,ki}w'_i=\{w'_{i,1},\dots,w'_{i,k_i}\}8, and their interaction, wi={wi,1,,wi,ki}w'_i=\{w'_{i,1},\dots,w'_{i,k_i}\}9. In the random condition, fewer cues led to worse F-scores, whereas in the Pred-RC condition F-score remained relatively stable even when cues were reduced (Fukuchi et al., 2023).

The survey literature generalizes these findings into four determinant domains: environment, interaction, social component, and technical component. Task type, setting, and use case belong to the environment; decision-making approach and reliance measure belong to interaction; user training and performance feedback belong to the social component; and AI implementation and transparency mechanisms belong to the technical component. The same survey emphasizes that most studies still treat reliance as a static average even though repeated interaction, feedback, and learning can change reliance over time (Eckhardt et al., 2024).

5. Typologies and domain-specific forms

In academic writing, instrumental reliance has been operationalized as one subtype within a broader four-factor LLM Reliance Scale. The construct blueprint characterizes it as “Bounded use for mechanical tasks (outlining, sentence revision, summarizing) without delegation of ideational content,” and the scale description states that the Instrumental subscale assessed “bounded, task-directed AI use for specific mechanical functions, such as outlining, sentence revision, and summarizing, without delegation of ideational content.” The representative item is: “I use generative AI tools to rephrase complex sentences in my drafts.” In a sample of W={w1,,wm}W'=\{w'_1,\dots,w'_m\}0 undergraduates, Instrumental reliance was the dominant profile for W={w1,,wm}W'=\{w'_1,\dots,w'_m\}1 students, or W={w1,,wm}W'=\{w'_1,\dots,w'_m\}2, with descriptive mean W={w1,,wm}W'=\{w'_1,\dots,w'_m\}3, W={w1,,wm}W'=\{w'_1,\dots,w'_m\}4; the subscale showed Cronbach’s W={w1,,wm}W'=\{w'_1,\dots,w'_m\}5, McDonald’s W={w1,,wm}W'=\{w'_1,\dots,w'_m\}6, and W={w1,,wm}W'=\{w'_1,\dots,w'_m\}7. The paper also notes a limitation: Instrumental and Dialogic items co-loaded on Factor 1, indicating substantial empirical overlap at the indicator level (Hossain, 27 Jun 2026).

In software engineering, instrumental reliance is framed through the distinction between AI as an instrument and AI as a substitute for judgment. The proposed reliance-control framework includes five levels of control—Self-Control over AI, Taking Control over AI, Balanced Control over AI, Handing over the Control to AI, and Losing Control over AI—and five levels of reliance—Self-Reliance, Reliance on Colleagues, Appropriate Reliance, Overreliance on AI, and Full Automation. Examples of instrumental use include boilerplate generation, test generation in TDD, meeting summarization, information retrieval, initial prototype generation, code translation to another language, and selection-based or comment-guided enhancements. By contrast, large-block code generation without scrutiny, unsupervised agent mode, and continued debugging after loss of understanding are treated as signs of overreliance and loss of control (Ferino et al., 12 Apr 2026).

In healthcare and other safety-critical domains, the reliance drill is proposed as an organizational procedure for detecting over-reliance. A drill deliberately reduces the efficacy of a real-world AI system so that the AI is worse than the human baseline for the tested problem. Users pass if they notice and reject the mistake and fail if they accept or follow the faulty advice. The paper treats this as analogous to phishing simulations and argues that organizations may use results to trigger warnings, training, or workflow redesign. At the same time, it identifies risks of collateral harm, under-reliance, stress, morale effects, and the trade-off between ecological validity and risk minimization (Hunter et al., 2024).

These domain-specific formulations show that instrumental reliance is not always synonymous with strong deference. In writing, it can mean bounded mechanical support while retaining authorial control; in software engineering, it can mean capability extension under balanced control; in healthcare, it is often evaluated by whether oversight remains real rather than symbolic.

6. Appropriate reliance, controversy, and unresolved questions

A major controversy concerns what counts as “appropriate reliance.” The conventional formula—follow AI when it is correct and reject it when it is wrong—has been criticized as lacking formal statistical grounding and as conflating distinct phenomena. The decision-theoretic framework separates overall reliance rate from the user’s ability to discriminate which instances favor AI, and decomposes the gap between rational and behavioral performance into reliance loss and discrimination loss. On this view, high acceptance of AI advice is not automatically over-reliance; what matters is whether the reliance rate is payoff-optimal and whether it is applied to the right instances (Guo et al., 2024).

Workflow-based work reaches a related conclusion from a different direction. Offloading score is explicitly not interpreted as good or bad in isolation. When paired with code understanding, higher offloading score tends to predict lower understanding, and the paper uses W={w1,,wm}W'=\{w'_1,\dots,w'_m\}8 as a “good enough” understanding cutoff while discussing a corresponding offloading threshold. However, the same study also identifies a cluster of users with moderate-to-high offloading score and high understanding; these users reported using the tool to learn unfamiliar tasks, explore multiple lines of thought, and augment their capabilities. This suggests that identical levels of offloading may constitute overreliance in one context and productive, appropriate reliance in another (Padmakumar et al., 28 May 2026).

The writing literature adds a measurement critique. Strategic users—those who engaged AI most deliberately—scored lowest on standard outcome measures, and the paper argues that this reflects a limitation of current instruments, which index AI’s contribution rather than writing quality and therefore penalize students who show the greatest independent thinking (Hossain, 27 Jun 2026). A comparable caution appears in the offloading-score paper, which argues that usage-based measures such as fraction of AI code retained, number of AI interactions, or time spent with the tool can indicate that AI was used but do not reveal how the workflow changed or how much work was actually delegated (Padmakumar et al., 28 May 2026).

Several open problems remain. The survey literature argues that the field still lacks consensus on operational definition, often suffers from limited external validity because it relies on crowdworkers, Wizard-of-Oz setups, isolated laboratory tasks, or low-stakes settings, and frequently disregards temporal change in reliance over repeated interaction. It proposes a morphological box organized around nine subconcepts—task, setting, use case, decision-making approach, reliance measure, user training, performance feedback, AI system implementation, and transparency mechanism—as a way to structure future work. It also identifies generative AI and multi-user settings as especially important future directions, since existing agreement- or switch-based measures do not transfer cleanly to co-writing, drafting, or group decision settings (Eckhardt et al., 2024).

Taken together, these debates indicate that instrumental reliance is now studied less as a binary property of “using AI” and more as a multidimensional question about delegation, control, effort displacement, behavioral calibration, and outcome alignment. The common technical agenda is to make those dimensions observable and separable enough to support diagnosis, comparison, and intervention across real workflows.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Instrumental Reliance.