Robust Human-AI Complementarity under Uncertainty
Abstract: Machine learning models are often intended to augment rather than replace human decision makers, by providing information that is complementary to human judgement. Yet, in practice, human decision makers routinely fail to realize such complementary gains, even when models provide useful signal. In this work, we study how asymmetric information about the quality of information available to a human decision maker vs. an AI impacts the ability of a decision maker to extract complementary value from AI predictions. We show that a key factor is the error correlation structure between human and AI predictions. In particular, when the AI's prediction errors are \textit{negatively correlated} with those of the human, the decision maker can construct robust strategies which guarantee improvements in expected utility. We empirically investigate whether these conditions for complementarity arise in practice, using real-world forecasting benchmarks.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Explain it Like I'm 14
What is this paper about?
This paper looks at how people and AI can work together to make better decisions. The big idea is “complementarity”: the AI should add useful information that the human doesn’t already have. But in the real world, people often don’t get the extra benefit they expect from AI help. The authors ask why that happens and what conditions are needed so that teaming up with AI reliably helps, even when we aren’t sure how good the AI really is.
What questions did the researchers ask?
They focus on a few simple, practical questions:
- When does adding AI advice actually help a human decision-maker, especially if we don’t fully know how accurate the AI is?
- How do the kinds of mistakes humans and AIs make together affect whether teaming up is helpful or not?
- Do current AI models (like LLMs) make different mistakes than humans in real-world prediction tasks?
- Can we prompt AIs in special ways to make their mistakes more “complementary” to humans?
How did they study it?
Think of decision-making like trying to guess a hidden number (the “truth”). The human has a clue, the AI has another clue, and both clues are noisy (imperfect).
Two key ideas:
- Error correlation: Imagine you and a friend are answering quiz questions. If you both tend to miss the same questions, your mistakes are positively correlated. If you tend to get opposite questions wrong (you miss what your friend gets right, and vice versa), your mistakes are negatively correlated.
- Uncertainty about AI quality: In real life, we often don’t know exactly how accurate the AI is for the specific task right now (models change, tasks are new, data is limited). The authors ask for strategies that still help even under this uncertainty—this is called being “robust.”
Their approach had three parts:
- Theory: They built simple math models to study when combining human and AI clues can be guaranteed to reduce errors. They examined both:
- Predicting a number (like a forecast score), and
- Making a yes/no choice (like “should we run this experiment or not?”). A central trick is to look at the AI’s “residual”—the part of the AI’s prediction that isn’t predictable from the human’s prediction—so you only use the AI where it adds new information.
- Simulations: They ran controlled experiments on fake data to see how performance changes as error correlation (how similar the mistakes are) moves from negative to positive.
- Real-world tests: They compared human and AI predictions on:
- ForecastBench (a set of forecasting questions), and
- Social science experiments (TESS studies) where both humans and LLMs predicted outcomes. They also tried different prompting styles to push the AI to think differently (for example, “be contrarian” or “focus on overlooked factors”).
Glossary (in everyday terms):
- Error correlation: Do human and AI tend to be wrong in the same way at the same time (positive), or in opposite ways (negative)?
- Residual: What’s left of the AI’s prediction after removing the part that the human already knows—i.e., the AI’s new, unique information.
- Robust: A strategy that still helps even if you don’t know exact details about the AI’s accuracy.
What did they find?
Main theoretical findings:
- Negative correlation in errors is the sweet spot. If the AI tends to be wrong in different ways than the human (negative error correlation), you can build simple, robust rules that guarantee the human+AI team does better than the human alone—even when you’re unsure how good the AI is.
- Positive correlation makes complementarity hard. If human and AI make similar mistakes, it’s much tougher to ensure reliable gains. In that case, using the AI helps only in narrower situations—basically when the AI is clearly better and the human is clearly noisier in a specific way. Otherwise, you drift toward “automation” (letting the AI take over) rather than true teamwork.
- These results hold for both number-prediction tasks and yes/no decision tasks, and they extend beyond simple math assumptions to more general settings.
Simulation results:
- When errors are negatively correlated, using the AI’s residual information brings clear improvements.
- When errors are positively correlated, improvements shrink and only appear if the human’s information is much noisier than the AI’s.
Real-world results:
- On two real forecasting datasets, human and AI errors were positively correlated (they tended to make similar mistakes).
- Prompting tricks (like asking the AI to be contrarian or to focus on overlooked factors) sometimes reduced the correlation but rarely flipped it to negative. In other words, simple prompting alone usually wasn’t enough to create true complementarity.
Why this matters:
- Accuracy isn’t the whole story. It’s not just “is the AI good?” It’s “does the AI make different mistakes than the human?” That difference is crucial for teamwork to reliably help.
What’s the impact?
This work suggests a shift in how we design and evaluate AI for collaboration:
- Don’t only measure accuracy—also measure error correlation with humans. For team success, you want low or negative correlation (different mistakes).
- Train AIs to add complementary information. Future AI training might include goals that explicitly push models to cover human blind spots, not just to be accurate on average.
- Be cautious when adding AI to human workflows. If human and AI tend to fail in the same ways, the “team” may not improve much, and sometimes won’t improve at all, especially under uncertainty.
- New evaluation checklists: Before deploying AI as a teammate, check whether it provides unique, reliable signals that humans don’t already have.
In short, to build great human-AI teams, we need AIs that don’t just mirror human thinking, but truly complement it—especially by making different kinds of mistakes so that, together, the team sees more of the truth.
Knowledge Gaps
Below is a consolidated list of concrete knowledge gaps, limitations, and open questions that remain after this paper. Each item highlights a specific avenue where additional theory, methodology, or evidence is needed for actionable progress.
- Uncertainty modeling beyond a single covariance entry: The analysis assumes only Cov(DAI, θ) is uncertain while all other covariances (e.g., Var, Cov between human and AI signals) are known. How do the results change when multiple or all entries are uncertain, including Var(DAI), Cov(φH, φAI), or Var(φH|θ)?
- Eliciting/estimating credal sets in practice: The robust rules require bounds like δ = inf Cov(θ, DAI) and γ = sup Cov(φH, φAI|θ). How can practitioners elicit, estimate, or validate these bounds with limited or no ground truth, and what sample sizes are needed?
- Sensitivity to misspecification: How fragile are the guarantees if the uncertainty set U is misspecified (e.g., true Cov(θ, DAI) lies below δ, or Cov(φH, φAI|θ) exceeds γ)? Provide sensitivity analyses or robustification techniques.
- Sequential and interactive settings: The model treats signals as jointly observed and independent of presentation order. How do conclusions change when AI advice influences human updates (and vice versa), or when advice is deferred/triaged over time?
- Behavioral factors and bounded rationality: Results assume a rational decision maker follows the robust rule. What happens when humans under/over-weight the AI or misinterpret residualized advice? Quantify performance when implementation deviates from the optimal rule.
- Necessity vs. sufficiency for binary decisions: For MSE in the Gaussian case, the paper gives tight conditions; for binary investment, conditions are largely sufficient. Are there matching necessity results for binary (or other) decision tasks?
- Generality across loss functions and tasks: Extend theory to asymmetric costs, 0–1 loss, ranking/selection, resource allocation, and multi-threshold decisions. Do negative-error-correlation conditions remain primary, or do new structures emerge?
- Multivariate states and outcomes: The theory focuses on scalar θ. How do robust complementarity conditions extend to vector-valued states (e.g., multi-label prediction, structured outputs) and multi-objective utilities?
- Heteroskedasticity and nonstationarity: The core analysis assumes homoskedastic Gaussian (or monotone) noise. What if error variance depends on instance features or drifts over time as models/users evolve?
- Online adaptation under model drift: Practical settings involve changing models and data. Design and analyze online procedures that track error correlation and update robust rules with guarantees under distribution shift.
- Estimating residuals without ground truth: The residual r requires regressing out the component of DAI predictable from φH. How accurately can this be estimated from prediction-only logs (no labels), and how does estimation error affect guarantees?
- Sample complexity for detecting (negative) error correlation: Provide finite-sample bounds for reliably inferring the sign and magnitude of Cov(errors) and deciding whether robust complementarity is feasible.
- Group-specific complementarity and fairness: Error correlations may differ across subgroups. How to assess and optimize complementarity while ensuring equitable improvements and avoiding subgroup harms?
- Team composition with multiple humans/models: Extend conditions and algorithms to selecting/combining multiple forecasters and multiple AIs to enforce complementary error structures at the team level.
- Training objectives to induce complementarity: The paper suggests optimizing for complementarity but provides no concrete objectives. What losses or regularizers explicitly encourage negative error correlation with target human populations?
- Trade-off between standalone accuracy and complementarity: How to characterize and optimize the Pareto frontier between a model’s individual performance and its error decorrelation from humans?
- Prompting and representation-level interventions: Post-hoc prompting only modestly affected error correlations. What principled prompting, feature masking, or representation learning strategies reliably decorrelate errors without degrading calibration?
- Personalized complementarity: Conditions depend on the specific human’s error structure. How to learn user-specific models that adapt to and complement an individual’s mistakes with minimal data?
- Robustness of non-Gaussian results: The non-Gaussian extension assumes strictly increasing f, g with bounded derivatives and a monotone dependence condition E[eAI | eH] nonincreasing. How critical are these assumptions in real LLM outputs, and can they be relaxed (e.g., piecewise monotone, heavy-tailed errors)?
- Practical computation of conservative bounds in binary decisions: The dj policy relies on Plow and Phigh derived from pessimistic posterior parameters. How to compute and calibrate these bounds for real, non-Gaussian models without closed-form posteriors?
- Domain generality of empirical findings: Evidence comes from ForecastBench and TESS. Do similar positive error correlations hold in domains like healthcare, law, finance, and scientific subfields, and how do domain properties mediate complementarity?
- Causal and strategic considerations: If humans strategically adjust to AI (or vice versa), or if data collection affects future predictions (feedback loops), do the complementarity conditions still apply? Develop causal models that capture these dynamics.
- UI/UX for residualized advice: The robust policy relies on residual information. What interfaces best communicate “complementary” signals to humans, and how does explanation format affect uptake and reliability?
- Privacy and data availability constraints: Complementarity optimization may require human rationales or historical errors. How to achieve error decorrelation under privacy limits or sparse human data?
- Cost modeling and decision thresholds: Binary decision analysis uses a fixed cost c. How do varying, uncertain, or instance-dependent costs affect the design and guarantees of robust joint policies?
- Adversarial or worst-case environments: Investigate whether adversarial shifts can induce positive error correlation and how to design defenses (adversarial training or minimax objectives) that preserve complementarity guarantees.
- Benchmarking metrics for complementarity: Move beyond aggregate accuracy to standardized metrics capturing error correlation, residual informativeness, and robust utility improvements for human-AI teams.
- Guidelines for selecting among multiple AIs: Given several models with similar accuracy but different error structures relative to humans, how to choose or ensemble them to maximize robust complementarity?
- Replicability and measurement noise: Ground truth in social science (e.g., effect sizes) can be noisy. How do measurement errors in labels bias error-correlation estimates and downstream conclusions about complementarity?
Practical Applications
Immediate Applications
The following items translate the paper’s findings into concrete, deployable workflows and tools. They focus on safely realizing gains from human–AI teaming by exploiting residual (non-overlapping) AI signal and by monitoring error correlation.
- Bold: Complementarity guardrails for decision support (forecasting, policy analysis, finance)
- What: Add a “robust combine” module to existing dashboards that (i) logs human predictions, (ii) residualizes AI predictions against the human’s, and (iii) only adjusts the human’s decision when a conservative improvement is guaranteed under the uncertainty set.
- How: Implement linear residualization r = AI_pred − α·Human_pred (with α estimated on historical pairs) and use a conservative coefficient b ≥ 0 to form Human + b·r for regression tasks; for binary actions, use conservative success-probability bounds to overrule human choices only when uniformly beneficial.
- Tools/products: BI/forecasting platform plugin; Python/R SDK for residualization and robust gating; CI-style “complementarity checks.”
- Assumptions/dependencies: Availability of labeled historical pairs (human, AI, outcome) to estimate α and test error correlation; outcome costs/thresholds known; model version stability; data logging and privacy agreements.
- Bold: Pre-deployment complementarity audits (industry governance)
- What: Extend model evaluation to include error correlation with target users, not just accuracy. Report p(E_H, E_AI), complementarity margin K_γ, and conditions where robust rules improve vs revert to human-only.
- How: A/B tasks where humans and AI predict before ground truth is revealed; compute error correlation and maximal guaranteed improvement.
- Tools/products: “Complementarity scorecard” in model cards; procurement checklist requiring complementarity metrics.
- Assumptions/dependencies: Representative human cohort; task-relevant outcomes; enough samples for stable correlation estimates.
- Bold: Human-in-the-loop override policies (operations, risk management)
- What: Default to human-only when measured error correlation is positive and the complementarity margin is small; enable robust override only when residual evidence is sufficiently strong.
- How: Policy that dynamically switches between human-only, robust-combine, and (rarely) AI-suggested contrarian actions under explicit thresholds.
- Tools/products: Decision playbooks; workflow engines with gating logic.
- Assumptions/dependencies: Continuous monitoring; incident review to recalibrate thresholds.
- Bold: Forecasting team workflows (finance, public policy, energy)
- What: Use residualized AI inputs for: revenue forecasts, macro indicators, load/price forecasts, epidemiological curves.
- How: Calibrate on held-out historical events; deploy robust linear blend for MSE/Brier tasks; show “novel evidence” panels that visualize the residual separate from the AI’s raw forecast.
- Tools/products: Spreadsheets add-on; time-series toolkits with residual AI channel.
- Assumptions/dependencies: Reasonably stationary data-generating process; frequent revalidation under distribution shift.
- Bold: Conservative triage for binary decisions (clinical orders, fraud review, experiment selection)
- What: For invest/skip-type decisions (e.g., order test, escalate case, run experiment), apply the paper’s symmetric robust rule to overrule the human only when lower-bound success probability ≥ cost.
- How: Estimate conservative posterior bounds (Plow, Phigh) from historical data; implement “only overturn when guaranteed net gain” logic.
- Tools/products: Triage APIs that accept human score, AI score, and return action with rationale; EMR or case-management integration.
- Assumptions/dependencies: Known cost/benefit; labeled outcomes; human baseline calibration available.
- Bold: Training and enablement for analysts and clinicians (education, healthcare, finance)
- What: Teach residual thinking: “Ask the AI for what I missed,” interpret the residual channel vs the raw suggestion, understand when to ignore the AI.
- How: Short courses, job aids, and UI hints highlighting residual evidence and error-correlation dashboards.
- Tools/products: Microlearning modules; on-screen affordances in decision tools.
- Assumptions/dependencies: Behavioral adoption; leadership support.
- Bold: Prompt design to harvest non-overlapping evidence (software, knowledge work)
- What: Use contrarian or divergence prompts to elicit alternative factors the human might have missed and extract those as a separate residual input.
- How: Two-pass prompting: normal rationale → “provide distinct, independent considerations”; use only the incremental parts in the residual channel.
- Tools/products: Prompt templates; rationale differencing components.
- Assumptions/dependencies: Only partially effective; must be combined with residualization and auditing because prompts rarely flip correlation negative.
- Bold: Model selection for teams (procurement, MLOps)
- What: Choose among models not only by accuracy but by lowest positive error correlation with your users; prefer models that are “orthogonal” to user mistakes.
- How: Bake correlation and guaranteed-improvement metrics into model selection and ensemble design.
- Tools/products: Model registry tags for complementarity; automated bake-offs.
- Assumptions/dependencies: User-specific evaluation; model stability.
- Bold: Scientific screening with guardrails (academia, R&D)
- What: Use robust binary policy for experiment triage (e.g., social science, materials). Overrule a PI’s decision to run/not run only when residualized AI signal safely improves expected utility.
- How: Estimate thresholds from prior studies; log successes/failures to refine uncertainty sets.
- Tools/products: Lab management integration; preregistration utilities with complementarity checks.
- Assumptions/dependencies: Outcome ground truths are costly and sparse; conservative policies minimize wasted runs.
- Bold: “Fallback-to-human” safety in new domains (all sectors)
- What: Under high uncertainty about AI–truth correlation, run human-only decisions until enough labeled data accrues to enable robust complementarity.
- How: Phased rollout with sample-size gates; drift alarms that revert to human-only.
- Tools/products: Deployment playbooks; uncertainty-set estimators.
- Assumptions/dependencies: Label latency; monitoring infrastructure.
- Bold: Personalized assistance for novices (education, customer support)
- What: Because robust policies can help when human uncertainty is high, target AI residualization to novices while leaving experts mostly unassisted when correlation is positive.
- How: Segment users by historical error variance; apply different gating thresholds.
- Tools/products: Role-aware assistants; adaptive UI.
- Assumptions/dependencies: Fairness/privacy safeguards; accurate user modeling.
- Bold: Complementarity telemetry (product analytics)
- What: Continuously measure human-only vs robust-combine outcomes; track correlation, gains, and reversions over time.
- How: Ship a complementarity telemetry schema capturing human input, AI input, action taken, and ground truth.
- Tools/products: Observability dashboards; periodic “complementarity health” reports.
- Assumptions/dependencies: Careful governance of sensitive data; statistically sound attribution.
Long-Term Applications
These items require further research, development, or scaling—especially new training methods that target complementary error structures.
- Bold: Complementarity-optimized model training (foundation models across sectors)
- What: Train/fine-tune models to minimize positive error correlation with target-user populations, or explicitly encourage negative dependence in errors with respect to user residuals.
- How: Multi-objective loss combining accuracy with a complementarity regularizer (e.g., penalize Cov(E_H, E_AI)); adversarial data augmentation to diversify errors; conditioning on user profiles.
- Tools/products: “Complementarity-tuned” model families; per-organization adapters.
- Assumptions/dependencies: Access to paired human predictions and outcomes; privacy-preserving learning; transferability across tasks.
- Bold: Standards and regulation for human–AI teaming (policy, healthcare, finance)
- What: Require complementarity metrics in clinical decision support, underwriting, or public policy evaluations before deployment.
- How: Incorporate error-correlation and guaranteed-improvement thresholds into approvals, audits, and continuous oversight.
- Tools/products: Regulatory technical standards; certification programs.
- Assumptions/dependencies: Consensus on metrics; sector-specific risk thresholds.
- Bold: Adaptive teammates that learn each user’s error profile (software, education, healthcare)
- What: On-device or privacy-preserving systems that build user-specific error models and dynamically produce residualized, complementary outputs.
- How: Online learning of user residuals; context-aware residual prompts; per-user uncertainty sets.
- Tools/products: Personalized copilots; clinician- or analyst-specific assistants.
- Assumptions/dependencies: Strong privacy controls; sample efficiency; robust drift handling.
- Bold: Complementarity-aware RLHF/feedback loops (model alignment)
- What: Incorporate feedback that rewards helping the human where they are weak and avoiding redundant or jointly erroneous suggestions.
- How: Human-in-the-loop evaluations that score complementarity, not just helpfulness; bandit-style exploration for error orthogonality.
- Tools/products: New annotation protocols; alignment datasets tagged with complementarity signals.
- Assumptions/dependencies: Costly to collect; requires careful UX to avoid user confusion.
- Bold: Benchmarks for complementarity (academia, industry consortia)
- What: Public suites measuring error correlation with diverse human cohorts and tasks (forecasting, medical vignettes, legal reasoning, code review).
- How: Shared datasets with paired human/AI predictions and outcomes; leaderboards ranking accuracy and complementarity.
- Tools/products: “ComplementarityBench”; standardized metrics (p(E_H,E_AI), K_γ, guaranteed-gain curves).
- Assumptions/dependencies: Ethical data collection; representativeness.
- Bold: Algorithmic toolkits beyond Gaussian assumptions (software, research)
- What: General-purpose libraries implementing robust residualization and decision rules under non-Gaussian, monotone, or model-free conditions described in the paper’s extensions.
- How: Nonparametric conditional expectation estimators for residuals; conservative lower bounds on Cov(residual, outcome); PAC-style guarantees.
- Tools/products: Open-source packages; AutoML components that learn complementarity-safe policies.
- Assumptions/dependencies: Sample complexity; calibration under shift; explainability.
- Bold: Marketplaces for “diversity of errors” (AI ecosystems)
- What: Model catalogs labeled by error correlation with common user types; ensemble builders that pick complementarity-optimized model mixes.
- How: Discovery and auto-selection based on user error telemetry; pricing incentives for complementary models.
- Tools/products: Complementarity-aware model registries; ensemble routers.
- Assumptions/dependencies: Interoperability; vendor cooperation.
- Bold: Scientific discovery platforms co-optimizing teams (academia, pharma)
- What: End-to-end experiment design platforms that jointly optimize which hypotheses to pursue and how to combine human and AI judgments under uncertainty sets.
- How: Portfolio optimization using robust binary decision rules; budgeting under explicit cost–benefit and complementarity constraints.
- Tools/products: Lab portfolio managers; hypothesis-screening copilots.
- Assumptions/dependencies: Rich historical metadata; organizational buy-in.
- Bold: Complementarity-driven workforce design (operations, HR)
- What: Assign tasks to humans vs AI vs hybrid teams based on measured error correlation and human uncertainty, not just automation potential.
- How: Triage policies derived from the paper’s conditions: automate when AI dominates and correlation is high; hybridize when negative correlation emerges; human-only otherwise.
- Tools/products: Workforce orchestration platforms; policy simulators.
- Assumptions/dependencies: Reliable telemetry; change management.
- Bold: Incentive-aligned forecasting ecosystems (forecasting platforms, prediction markets)
- What: Reward forecasters and models not only for accuracy but also for providing complementary information (low correlation with consensus/human errors).
- How: Scoring rules that include a diversity/complementarity component; team-formation tools.
- Tools/products: Market mechanisms; platform features for complementarity scoring.
- Assumptions/dependencies: Mechanism design and gaming resistance.
- Bold: Safety cases for high-stakes autonomy (aviation, energy, healthcare)
- What: Use robust complementarity analysis to justify or reject mixed-initiative control, including explicit “revert to human-only” or “handover to automation” triggers based on real-time correlation proxies.
- How: Online estimation of dependence structure; certified guardrails.
- Tools/products: Safety cases; runtime monitors.
- Assumptions/dependencies: Real-time data; certification pathways.
- Bold: Education/tutoring that targets misconceptions (edtech)
- What: Train tutors to be complementary to student error profiles, giving feedback that is negatively dependent on common student mistakes.
- How: Learn misconception models; optimize tutor outputs for complementarity.
- Tools/products: Adaptive tutors; diagnostic assessments.
- Assumptions/dependencies: Student modeling ethics; generalization across curricula.
- Bold: Cross-domain deployment under shift (all sectors)
- What: Methods to estimate and update uncertainty sets U as tasks evolve, maintaining guarantees for robust complementarity as models or contexts change.
- How: Bayesian/credal updating over predictive posteriors; drift-aware re-estimation of bounds.
- Tools/products: Uncertainty-set managers; continuous validation pipelines.
- Assumptions/dependencies: Label availability; efficient re-estimation.
- Bold: IDEs with “residual hints” for coding (software engineering)
- What: Assistants that surface only those suggestions predicted to be complementary to a developer’s likely mistakes (based on context and history).
- How: Learn developer-specific error distributions; gate code suggestions through residual-evidence scoring.
- Tools/products: Complementarity-aware code copilots.
- Assumptions/dependencies: Privacy-preserving telemetry; editor integration.
Notes on Feasibility, Assumptions, and Dependencies
- Negative error correlation is the main enabler. Current LLMs often exhibit positive correlation with human errors; immediate gains may be modest unless residualization and strict gating are used.
- Estimating parameters requires labeled pairs of human/AI predictions and outcomes. When outcomes are scarce or delayed, deploy conservative policies and revert to human-only until enough data accrues.
- Model drift breaks estimates. Complementarity metrics must be part of continuous monitoring, with rapid fallback mechanisms.
- Non-Gaussian conditions are supported theoretically if signals are monotone and errors are negatively dependent; verifying these assumptions and computing conservative bounds remains an active research area.
- Human factors matter. Training users to interpret residual evidence and accept “no-change” outputs is essential to avoid overreliance or automation bias.
Glossary
- AI residual: The part of the AI signal left after removing what is predictable from the human signal; used to isolate complementary information. "Define the AI residual r by regressing out the portion of ¢AI predictable from ¢H"
- Asymmetric information: A situation where the human and AI have unequal knowledge about model quality or task-relevant distributions. "We study how asymmetric information about the quality of information available to a human decision maker vs an AI impacts the ability of a decision maker to extract complementary value from the AI."
- Binary decisions: Decision-making where the action space has only two options, typically invest/not invest or accept/reject. "We next study a setting in which a decision maker takes a binary action d € {0, 1} after observing signals."
- Brier score: A proper scoring rule equivalent to mean squared error for probabilistic predictions. "(where MSE, aka Brier score, is a common metric)"
- Conditional covariance: Covariance between two variables given a third, capturing their dependence after conditioning. "Under joint Gaussianity, the conditional covariance Cov(PH, PAI | 0) is almost surely constant as a function of 0,"
- Covariance matrix: A matrix capturing variances and covariances among variables in a multivariate distribution. "The core of the problem is encapsulated in the covariance matrix Z."
- Credal set: A set of plausible probability distributions reflecting uncertainty about the true model. "Our framework can therefore be viewed as placing a credal set over plausible predictive posteriors"
- Error correlation structure: The pattern of dependence between human and AI prediction errors that shapes complementarity. "We show that a key factor is the error correlation structure between human and AI predictions."
- Expected utility: The average utility of a decision rule under uncertainty, used as an optimization objective. "which guarantee improvements in expected utility."
- Gaussian measurement model: A data-generating setup where signals are linear functions of the state plus Gaussian noise. "We simulate data from the Gaussian measurement model"
- Generative process: A probabilistic model describing how observed signals are produced from latent states. "We study a generative process in which the signals are noisy measurements of the ground truth state."
- Joint normality: The assumption that variables are jointly Gaussian, implying linear conditional expectations. "conditional expectations are linear in the signals under joint normality."
- Law of total covariance: A decomposition expressing covariance via conditional covariances and covariances of conditional means. "using the law of total covariance."
- Linear-Gaussian: A model where relationships are linear and noise is Gaussian. "a stylized, linear-Gaussian version of the model"
- Mean squared error (MSE): The expected squared difference between predictions and the true value; a standard loss. "mean squared error l(d, 0) = (d - 0)2."
- Mean-zero Gaussian noise: Random noise with a normal distribution centered at zero. "where EH and EAT are mean-zero Gaussian noise."
- Negative correlation (of errors): A dependency where one agent’s errors tend to be counterbalanced by the other’s, enabling complementarity. "(2) is equivalent to EH and E AI being negatively correlated, i.e., negative correlation in the errors."
- Negative dependence: A broad condition where increases in one error tend to accompany decreases in the other. "That is, negative dependence in errors is sufficient to ensure robust complementarity under mean squared error"
- Pearson correlation coefficient: A statistic measuring linear correlation between two variables. "p(EH, EAI) denotes the Pearson correlation coefficient over the human error and the LLM error."
- Pessimistic regression lower bound: A conservative lower bound on a regression coefficient used to ensure uniform safety under uncertainty. "Define the positive-correlation margin and the correspond- ing pessimistic regression lower bound"
- Posterior success probability: The probability of meeting a threshold given observed signals, used to decide whether to invest. "the posterior success probability exceeds the cost:"
- Predictive posterior: The posterior distribution over the outcome of interest after observing signals. "the induced predictive posterior over the state given the observed human and AI signals."
- Regression coefficient: The slope parameter linking a predictor to the outcome in a linear regression. "Define the regression coefficient of r on 0:"
- Residualized machine signal: The AI signal after removing components predictable from the human signal, highlighting complementary parts. "the residualized machine signal becomes less complementary,"
- Robust estimator: An estimator designed to perform well across a range of plausible distributions or uncertainties. "MSE of the expert-only estimator and our robust estimator"
- Robust symmetric policy: A decision rule that symmetrically incorporates AI residuals and guarantees improvements across an uncertainty set. "the robust symmetric policy yields the largest gains"
- Statistical decision theory: A framework for choosing actions under uncertainty based on probabilistic models and loss functions. "a formal model based on statistical decision theory"
- Triage decisions: Procedures that allocate cases between human and AI or prioritize effort based on predicted difficulty. "triage decisions (Raghu et al., 2019; Okati et al., 2021)"
- Uncertainty set: A specified collection of distributions capturing what is unknown about model quality or dependencies. "they have an uncertainty set which formalizes their beliefs about the model's quality"
Collections
Sign up for free to add this paper to one or more collections.