Determine the appropriate candidate-selection budget

Determine the appropriate candidate-selection budget for selecting generative molecular crystal structure prediction models, accounting for the trade-off between small-budget per-sample success and larger-budget crystal-level coverage.

Background

Packora’s ablations and model selection primarily use a budget of 30 generated candidates. However, the scaling experiments show that model behavior differs across sampling budgets: increasing single-track capacity can improve single-sample success while reducing crystal-level coverage, whereas balanced scaling can improve the trade-off. Consequently, the model selected using a small candidate budget may not be optimal when substantially larger candidate pools are used.

The unresolved issue is therefore how to choose the candidate budget used for model selection in generative molecular crystal structure prediction. Establishing an appropriate selection budget would help ensure that the selected architecture and training configuration are aligned with the intended downstream workflow, in which multiple generated candidates are relaxed and ranked.

References

Our scaling results show a trade-off between small-budget success and larger-budget coverage, suggesting that model selection may benefit from larger candidate pools. Determining an appropriate selection budget therefore remains an important question for generative CSP.

Packora: Systematic Design for Generative Molecular Crystal Structure Prediction  (2608.26962 - Kim et al., 27 Aug 2026) in Section 6, “Limitations and discussion” (Conclusion)