Specialization on truly private codebases

Ascertain whether repository specialization with SERA yields comparable performance improvements on truly private codebases that were not included in base-model or teacher-model training data and that models have never seen, given the lack of directly testable evaluation instances.

Background

Specialization experiments target public repositories (Django, Sympy, Sphinx) using evaluation instances from SWE-bench Verified, which may overlap with model pretraining data.

The authors note potential bias and explicitly state they have not verified specialization on truly private codebases due to the absence of evaluation data.

References

While specialization effects are well-studied in fine-tuning scaling laws and our results appear plausible, we have not verified specialization on truly private codebases that models have never seen because we have no evaluation data to test this directly.

SERA: Soft-Verified Efficient Repository Agents  (2601.20789 - Shen et al., 28 Jan 2026) in Section 9 (Limitations), Private repo specialization

One corpus is what we measured, and it is not what we claim. The generalization a buyer actually wants, will a self-hostable specialist hold up on my documents, in my domain?, is not answerable from any public benchmark.

The Commercial Tax: Rent-vs-Own Blind Spots in Multi-Hop Retrieval Benchmarks  (2608.16096 - Sanchez et al., 17 Aug 2026) in Section 7, Conclusion, “What this measures, and what comes next”; Section 6, Limitations