Principled evolution of evaluation metrics to avoid overfitting and point solutions
Devise principled procedures to systematically update and diversify evaluation metrics over time so as to prevent overfitting and discourage point solutions in robotics benchmarks and challenges.
References
Another open question is how to systematically change evaluation metrics in a principled way to avoid overfitting and point solutions.
At meta-harness level, better component selection and evaluator evolution remain open problems. A selector should choose next bounded update from failure attribution, uncertainty, expected improvement, and component interactions. Any adaptive evaluator must remain versioned and anchored by frozen reference tasks, adversarial probes, and periodic human audits so optimization doesn't reward-hack a moving target.