Disentangle the sources of ArmSTEM’s adaptation gains

Determine the separate contributions of STEM content, question–answer format, and translation verification to the performance gains produced by the six-point ArmSTEM mixture swap in continued pretraining of Gemma-4-E4B for Armenian, using a format-matched comparison.

Background

The released arm-gemma-e4b recipe replaces six percentage points of ArmWeb with verified English–Armenian mathematics and science data, and this change reverses much of the knowledge loss caused by continued pretraining on Armenian news. However, the intervention changes several properties simultaneously: the added data contains STEM knowledge, follows a question–answer format, and has undergone answer-preserving translation verification.

The paper explicitly states that the current experiments cannot identify which of these properties is responsible for the observed gains. A format-matched comparison run is therefore needed to isolate the effects while holding the structure of the training examples constant.

References

Within the 6-point mixture swap we cannot yet separate the contributions of STEM content, QA format, and verification, and a format-matched comparison run is future work (STEM-CPT-full already provides a repetition-matched one, \S\ref{sec:results}).

From Zero to Hero: An Open LLM Ecosystem for Armenian  (2609.03350 - Arakelyan et al., 3 Sep 2026) in Section ‘Limitations’