Paraphrase robustness for workflow generation
Characterize paraphrase robustness for natural-language workflow generation beyond the templated and systematically varied prompts used in the 635-rule benchmark.
References
Because our prompts couple this injected noise with deep branch structure, we expect the harder bottleneck to remain structural rather than surface-level; still, paraphrase robustness for workflow generation is a known open problem \citep{xu2025}, and we do not claim to isolate fully free-form paraphrase.
Every cell is a single greedy (temperature-0) generation, so these intervals capture rule-sampling and scorer uncertainty but not generation stochasticity, prompt-paraphrase sensitivity, or few-shot-ordering sensitivity. The qualitative ordering is reproduced by the judge-independent Cond\% metric, but quantifying decode- and prompt-level variance is left to future work and is the main residual threat to the point estimates.