Paraphrase robustness for workflow generation

Characterize paraphrase robustness for natural-language workflow generation beyond the templated and systematically varied prompts used in the 635-rule benchmark.

Background

The evaluation benchmark uses programmatically generated natural-language templates with controlled variation in condition and action values, branch structure, surface phrasing, spelling errors, and some non-English entity values. Although these variations probe a degree of surface robustness, they do not constitute fully free-form paraphrase evaluation.

The paper therefore identifies robustness to paraphrase as an unresolved issue relevant to deploying natural-language-to-workflow-DAG systems in less controlled settings. Establishing performance under genuinely diverse paraphrases would test whether the reported structural-generation results generalize beyond templated inputs.

References

Because our prompts couple this injected noise with deep branch structure, we expect the harder bottleneck to remain structural rather than surface-level; still, paraphrase robustness for workflow generation is a known open problem \citep{xu2025}, and we do not claim to isolate fully free-form paraphrase.

Generating Workflow DAGs from Natural Language with Non-Reasoning LLMs  (2608.30250 - Iyer et al., 31 Aug 2026) in Section 2, “Benchmark” paragraph

Every cell is a single greedy (temperature-0) generation, so these intervals capture rule-sampling and scorer uncertainty but not generation stochasticity, prompt-paraphrase sensitivity, or few-shot-ordering sensitivity. The qualitative ordering is reproduced by the judge-independent Cond\% metric, but quantifying decode- and prompt-level variance is left to future work and is the main residual threat to the point estimates.

Generating Workflow DAGs from Natural Language with Non-Reasoning LLMs  (2608.30250 - Iyer et al., 31 Aug 2026) in Appendix, “Equivalence tests (TOST) for the bridging claim”