Establish whether optimizing competence scores induces rigid reasoning

Establish whether optimizing the aggregate competence score defined from verdict stability, monotonicity, decisiveness, and Pareto viability causes language-model agents to adopt rigid rule-following with explicit exceptions rather than context-sensitive adjustments.

Background

The paper notes that stability, monotonicity, and decisiveness may be easiest to improve by maintaining a firm preference. Although the proposed framework permits models to flip their verdicts when morally relevant features change, such flexibility may increase stochastic variation and thereby reduce the aggregate competence score.

The authors therefore identify an unresolved empirical concern: training or optimizing systems against the proposed score could improve measured structural competence while producing excessively rigid behavior. The paper does not test whether this effect actually occurs and recommends interpreting the score as an evaluation-time coverage check rather than a reward signal.

References

Whether this materializes under optimization pressure is an empirical question we do not resolve here.

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment  (2609.05036 - Libert et al., 4 Sep 2026) in Section 3, Limitations