Identify the source of differences between training trajectories

Determine whether numerical precision, random seed, learning-rate schedule, batch configuration, or another uncontrolled factor causes the differences between the bf16 and fp32 training trajectories.

Background

The study compares two trajectories that share the architecture and data stream but differ in numerical precision and potentially other factors. They differ in conflict-cycle depth, event-significance regime, and roster fluidity.

Because the experimental records do not establish that precision is the only difference, the paper cannot attribute the cross-trajectory discrepancies to bf16 versus fp32 computation alone. Identifying the responsible factor is required for interpreting the robustness of the reported dynamics.

References

The two trajectories differ in numerical precision (bf16 vs.fp32) and other factors; the source of their differences (conflict-cycle depth, event significance regime, roster fluidity) is unresolved.

Pre-carved Niches: The Formation Dynamics of Modular Task Partitions in Early LLM Training  (2609.01170 - Li et al., 1 Sep 2026) in Section 8, Limitations, item (iii); Appendix D, “Two-Trajectory Details and Probe Repair History”