- The paper’s main contribution is a novel protocol that adds an auxiliary mediation arm to disentangle direct and indirect effects in conjoint experiments.
- It employs doubly robust, cross-fitted machine learning estimators and principal ignorability to ensure valid mediation analysis in high-dimensional settings.
- Empirical analysis reveals that neglecting mediation can misrepresent causal mechanisms, thus establishing a new standard for mechanism evaluation in multidimensional experiments.
Introduction and Motivation
The paper "Disentangling Causal Mechanisms in Conjoint Experiments Using Mediation" (2607.03508) addresses a foundational limitation in the design and analysis of conjoint experiments for causal inference: the inability to formally distinguish direct and indirect (mediated) effects when multiple attributes are randomized. In standard factorial designs, the causal effect of a focal attribute T is estimated marginally over other randomized features. However, if T influences respondent beliefs about an unobserved attribute M (such as party label being inferred from candidate race), the design precludes decomposing total effects into theoretically-motivated direct and indirect components.
Methodological Framework
The authors introduce a novel experimental protocol that supplements a traditional Y(T,M)-conjoint with an auxiliary "mediation" experiment M(T), where respondents report beliefs about the mediator given randomized profiles, but are not asked the outcome Y. This enables the identification of nested counterfactuals required for mediation analysis, which cannot be constructed from Y(T) or Y(T,M) experiments alone.
The estimands of interest are the average marginal component effect (AMCE), the average marginal indirect effect (AMIE), and the average marginal direct effect (AMDE), all defined in terms of nested potential outcomes. The paper adopts the principal stratification framework for mediation, using principal ignorability (PI) as the identification condition. PI asserts that potential outcomes are independent of the mediator's mapping function Gi​={Mi​(0),Mi​(1)} conditional on observed covariates X. This is generally more plausible than sequential ignorability in high-dimensional contexts typical of conjoint designs.
The identification of mediation effects rests on:
- Experimental randomization and positivity of all features,
- Manipulation exclusion (type of experiment does not affect outcomes),
- Principal ignorability (no unmeasured confounding between principal strata and the outcome).
Estimation is realized using doubly robust, cross-fitted machine learning estimators. The approach leverages influence-function based estimation, allowing high-dimensional nuisance components (T0) to be estimated flexibly via random forests or related methods, while retaining valid inference for the low-dimensional causal estimands.
Empirical Illustration
To demonstrate the utility of the proposed framework, the authors conduct a pre-registered replication of [Kirkland & Coppock 2018], which studies candidate preferences in US mayoral elections with and without party labels. The analysis focuses on how the effects of candidate race, gender, occupation, and experience are mediated through beliefs about party.
The T1-experiment reveals that, e.g., Black candidates are disproportionately inferred to be Democrats, a mapping that is especially pronounced among Democratic respondents. This property is critical for testing whether the effect of candidate race on vote intentions operates directly or is substantively mediated via party inferences.
Figure 1: Average treatment effects of candidate attributes (e.g., race, occupation) on the probability the respondent infers the profile to be Democratic/Republican/Independent.
The mediation analysis recovers group-specific AMIE and AMDE for each attribute and party subgroup. Key findings include:
- Race as Treatment: The indirect effect of candidate race via party label is large, positive, and statistically significant for Democrats (i.e., non-white candidates benefit via inferred Democratic label), and large, negative, and significant for Republicans (non-white candidates are penalized via inferred Democratic label). The direct effect, holding inferred party constant, is much smaller.
- Political Experience: The effect of previous political experience is almost entirely direct, with negligible indirect effect via party.
- Gender and Occupation: There exist meaningful indirect effects for gender and certain occupations. For example, female candidates are perceived as more likely Democrats, influencing Democratic respondents' preferences indirectly. Likewise, occupations such as educator are associated with strong party stereotypes that mediate choice.
Figure 2: Estimated direct, indirect, and total effects for candidate attributes, heterogeneous by respondent party.
These patterns underscore that ignoring the mediation structure, as in a plain AMCE analysis, conflates qualitatively distinct mechanisms and may misrepresent the substantive theory.
Heterogeneous Effects and Model Diagnostics
Exploratory analysis regresses the estimated mediation effects on respondent-level covariates (e.g., ideology, race, education). The results show, for example, that partisan and ideological orientation are strong moderators of both direct and indirect effects, highlighting systematic heterogeneity not recoverable from standard conjoint analysis.
Figure 3: Best linear projections of mediation estimands showing heterogeneity by respondent covariates.
The protocol includes a falsification test: the total effect from non-mediation T2-arms is compared to the total effect reconstituted via the mediation formula. Discrepancies point to possible violations of PI or manipulation exclusion; a sensitivity analysis quantifies the impact of such violations on the reported mediation effects.
Figure 4: Marginal mean outcomes from baseline and mediation-informed estimators for various attribute treatments.
Figure 5: Sensitivity analysis of mediation effects under violation of principal ignorability.
Discussion and Implications
The paper advances the formal design and analysis of conjoint experiments by making explicit the informational and identification requirements for mediation analysis. The addition of an T3-experiment is both minimal and practically feasible, yet unlocks identification of mediation estimands otherwise unavailable. This allows analysts to move beyond the "eliminated effect" heuristic (contrasting T4 and T5 AMCEs) to precisely quantify and interpret mechanisms—a major advantage for substantive theory testing in social science and marketing.
Practically, the approach provides two major contributions: (1) a validated protocol for mediation in highly-multivariate randomized experiments, and (2) robust statistical procedures for estimation and inference with machine learning in this context.
Theoretically, the findings reveal that direct and indirect effects can differ radically in sign and magnitude—notably, in the structure of group-specific responses to candidate attributes. Analysts who use conjoint designs without attention to mediation may significantly misinterpret experimental results by attributing effects to the focal treatment that actually operate via respondent inferences or beliefs about omitted variables.
Directions for Future Research
This framework opens several paths for further work, including:
- Incorporating multiple sequential/parallel mediators (e.g., ideology and party),
- Developing crossover or within-unit designs to relax principal ignorability,
- Extending approaches to forced-choice tasks or complex profile dependency structures,
- Theory-driven selection of mediators tailored to substantive domains, especially in high-dimensional attribute spaces.
Conclusion
The paper provides a critical refinement in the design and analysis toolkit for causal inference in conjunction experiments. By integrating an additional mediation-focused arm into otherwise standard factorial designs, it enables the principled decomposition of direct and mediated causal effects using robust, machine-learning-based estimators, with accompanying diagnostic and sensitivity tools. The empirical results illustrate the substantive insights that justifiably arise from this approach but would have been missed under conventional methods. This work thus sets a new standard for the analysis of mechanisms in multidimensional experimental designs where attribute-induced inference is a key theoretical concern.