Determine whether large language models can self-correct rule violations without fine-tuning
Determine whether large language models can autonomously self-correct violations of formal rules without specific fine-tuning, as assessed across reasoning tasks.
References
Models are unlikely to know when they are violating formal rules and it is unclear whether they can self-correct~\, but with specific fine-tuning they might self-correct against harmful text~\, and that training on generated data might not be the best approach to preserve reasoning about outlier cases~.
Whether self-correction and oscillation recur outside IaC is unexamined. These are hypotheses: confirming them requires replications that vary domain and validator while holding the repair loop fixed.
We conjecture that tasks gated by private conventions will remain highly resistant to this loop. Because nothing in the task's instruction or environment reveals what the planted thresholds or internal normalisation tables are, self-critique alone is unlikely to recover them. However, this durability strictly applies to non-derivable construction-time knowledge; operator gates may remain susceptible to advanced tool-use, search, or scaled test-time compute.