Formal specification synthesis by LLMs
Determine how to reliably synthesize high-quality formal program specifications with large language models, addressing the unresolved challenge that current LLM-based specification synthesis remains difficult.
References
The limited overlap between human and LLM successes suggests complementary strengths: they succeed on distinct problem subsets rather than merely differing in capability level, indicating LLM-based specification synthesis remains a challenging open problem.
Writing that specification is the hard part: models that must write their own gain nothing over an unaided baseline, and only $62\%$ of their specifications pass our audit. The usual failure is faithfulness, a specification that constrains part of the required behavior and leaves the rest free. Specification quality still tracks the outcome, failing on $89\%$ of unresolved instances against $47\%$ of resolved ones, making faithful specification synthesis a concrete open problem.