Identify the shared conversational ingredient responsible for the remaining treatment difference

Identify the conversational feature shared by all GPT-4o career-reflection interactions that distinguishes conversational reflection from static journaling and explains the trial’s outcome differences beyond the observed association between decision demands and career doubt.

Background

The transcript analysis found that cumulative decision demands were associated with greater post-test career doubt, but no other measured conversational behavior predicted the trial outcomes, and no day-level behavior predicted participants’ same-day states. This leaves unexplained the remainder of the difference between the conversational-agent and static-journaling conditions.

The authors note that their coding approach measures the frequency of conversational behaviors. A feature present in every agent conversation—such as responding to a socially perceived conversational counterpart—would not vary across participants and therefore could not be detected through those counts. Identifying that invariant feature is presented as the most important unresolved question.

References

Identifying that shared ingredient is the most important question this work leaves open.

Faithful Where It Can Be Checked: Auditing a Reflection Agent Against Its System Prompt in a Randomized Trial  (2609.19635 - Nepal et al., 17 Sep 2026) in Discussion, subsection “Implications for AI-guided reflection”

Whether decision pressure causes doubt is now a precise and testable question and we hope someone runs the experiment.