Extreme-tail reliability beyond prompt-complexity bin 16

Characterize the reliability relationship between the prompt-side structural-complexity index and code-generation pass rates beyond composite bin 16, where the available audit-clean extension contains too few prompts to determine the shape of the extreme tail.

Background

The study evaluates whether a prompt-side structural-complexity index predicts code-generation reliability across 21 models and reports a nonmonotone pass-rate relationship. To improve support in the high-complexity region, the authors add an audit-clean extension consisting of prompts whose reference solutions pass all supplied tests.

The extension materially increases support at bins 15 and 16, but only 14 added prompts lie above bin 16. Consequently, the evidence is sufficient for comparisons through bin 16 but does not determine whether the observed reliability pattern continues, reverses, or changes beyond that point. The authors explicitly limit their claim about the extreme tail because its shape remains unresolved.

References

Only 14 retained tail prompts lie above bin 16, so the extreme-tail shape remains unresolved.

— The Complexity Kink: A Prompt-Side Structural Complexity Index for Code-Generation Reliability  (2609.19616 - Hernandez et al., 17 Sep 2026) in Section 4, subsection “Audit-Clean High-Complexity Extension”