Establish the appropriate deference policy when the model lacks an independent belief

Determine whether an assistant should defer to a user who suggests the correct answer when the assistant has no independent belief of its own.

Background

PACT is designed to reduce deception while preserving legitimate context use. The paper’s belief-relative definition permits the model to decline to affirm a user’s claim when the model has no belief of its own, and the experiments show that PACT reduces such deference even when the user’s suggested answer is correct.

The authors explicitly leave unresolved whether this behavior is desirable. The issue is framed as a policy question rather than as a settled consequence of the unlearning method.

References

Where it has no belief of its own, PACT defers less to a user who suggests the correct answer (section 4.3); whether an assistant should defer there is a policy question our definition leaves open.

— Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets  (2609.38909 - Tang et al., 30 Sep 2026) in Section H, “Limitations”