Specialization on truly private codebases
Ascertain whether repository specialization with SERA yields comparable performance improvements on truly private codebases that were not included in base-model or teacher-model training data and that models have never seen, given the lack of directly testable evaluation instances.
References
While specialization effects are well-studied in fine-tuning scaling laws and our results appear plausible, we have not verified specialization on truly private codebases that models have never seen because we have no evaluation data to test this directly.
One corpus is what we measured, and it is not what we claim. The generalization a buyer actually wants, will a self-hostable specialist hold up on my documents, in my domain?, is not answerable from any public benchmark.