Stability of recursive synthetic self-improvement
Determine whether recursively generating synthetic training data with a Large Language Model and using it to train successor models yields genuine capability gains or results in model collapse.
References
An open question: what, linguistically, is this common destination?
The question left open is where the spectrum comes from: why the same data, landing on different checkpoints, produces outcomes five-fold apart.
A key open question is whether such a process would lead to genuine capability gains or result in model collapse, a degenerative process where the model overfits to its own idiosyncrasies, leading to a gradual loss of diversity and accuracy.
The open problem is how to increase useful adaptation while controlling the propagation of errors across future interactions.