Sufficient and effective representation for downstream tasks
Determine a self-supervised learning representation that is both sufficient and effective across a variety of downstream tasks, where "sufficient" means downstream tasks can be completed by composing functions only on the learned representations rather than on the original data, and "effective" means the composed functions for downstream tasks are lightweight models.
References
Concretely, despite the empirical successes achieved by representation from SSL, there are essential research questions have yet to be resolved, i.e, What representation is sufficient and effective for variety of downstream tasks? How can such a representation learned in an efficient and scalable way?
Finally, it remains to be tested how well our learned HRV representations transfer to additional downstream tasks (e.g., workload estimation).
The training signal is applied to a single clip-level token, and although semantically organized patch representations emerge without dense supervision, their sufficiency for dense prediction tasks such as segmentation and tracking has not yet been evaluated.
We have not yet evaluated whether the same network capacity remains sufficient when the pattern representation is pretrained over a substantially larger and more heterogeneous task collection.