Determine the source of the heavy simulated-conversation tail

Determine whether the heavy tail of simulated user-message lengths in Snowglobe conversations is specific to the evaluated Snowglobe configuration or common to large-language-model user simulators.

Background

The paper compares conversation-length distributions for production Card Delivery conversations, Snowglobe-generated simulations, and an off-topic production control. Simulated conversations contain substantially more user-message words than both production and control conversations, including a much larger proportion exceeding 200 words.

Because the control consists of production traffic rather than LLM-generated user text, it cannot establish whether the unusually long simulated conversations arise from the particular Snowglobe configuration or reflect a broader property of LLM-based user simulators. Resolving this distinction would clarify the generality and reliability of the observed simulation mismatch.

References

Because the control is production traffic rather than LLM-generated user text, it cannot determine whether the heavy simulated tail is specific to this configuration or common to LLM user simulators.

— Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale  (2609.30137 - Alcoba et al., 24 Sep 2026) in Appendix, Section "Conversation length distributions" (Appendix P1)