Effect of WEECFP-token pretraining on the classification gap

Determine whether pretraining on WEECFP tokens closes the classification-performance gap between the WEECFP-SuRGE Blend and competing pretrained molecular-learning methods.

Background

The WEECFP-SuRGE Blend is trained from scratch and achieves strong regression performance on the TDC ADMET benchmarks, but it remains behind pretrained methods on the classification subset. The paper notes that WEECFP tokenization is compatible with masked-substructure pretraining objectives, suggesting a possible route to improve classification performance while retaining the proposed representation.

The unresolved issue is specifically whether applying pretraining to WEECFP tokens would eliminate the observed classification gap; the paper does not resolve this question experimentally or theoretically.

References

WEECFP tokenization is compatible with masked-substructure pretraining objectives; whether such pretraining closes the classification gap is open.

WEECFP-SuRGE: Wide Embedded Extended Connectivity Fingerprint with Substructure Rotary Graph-distance Encoding  (2609.04672 - Epps, 4 Sep 2026) in Section 5, Discussion, paragraph “On pretraining”