Causes of LLM instability under deterministic settings
Determine the precise causes of response instability observed in leading large language models hosted via APIs, including OpenAI gpt-4o, Anthropic Claude-3.5, and Google Gemini-1.5, when repeatedly queried with identical legal questions under deterministic inference settings (temperature=0, fixed seed, and top-p or top-k set to 1.0); specifically ascertain whether the instability arises from nondeterministic floating‑point accumulation order, heterogeneous hardware or floating‑point implementations across servers, or parallelization-induced variations in execution order.
References
Why are some leading LLMs unstable, even with temperature=0 and setting the seed? It is impossible to say for sure, since the models are proprietary.
The present design cannot distinguish between these accounts, and none of them individually explains why unresolvable philosophy, which is also open-ended and lacks a ground truth, remains closer to the stable end of the scale.
No changes were intentionally introduced between the three evaluations. Consequently, the variation observed across runs cannot be attributed within this experiment to changes in the dataset, prompt structure, ground-truth labels, or configured generation parameters. The results demonstrate variability under repeated evaluation conditions; however, the experiment does not establish the underlying cause of that variability.
Second, batching pressure alone suffices to reproduce shared-endpoint-magnitude instability on one model and one stack; whether it is the cause of the shared-endpoint floor remains unidentifiable from these data, and this paper does not claim it.