Determine the effect of fine-tuning on SafeAlert evaluation performance

Determine whether fine-tuning produces different results from prompt-level interventions for the generation-safety and Nigerian fintech communication-classification tasks evaluated by SafeAlert.

Background

SafeAlert evaluates commercial models under three prompt-level conditions: no system prompt, a generic safety prompt, and a Nigerian-fintech-specific remediation prompt. The study does not evaluate model fine-tuning, so it cannot establish whether adapting model parameters would improve harmful-output rates, false-positive rates, or false-negative rates beyond the effects observed from prompt-level interventions.

The unresolved comparison is relevant because the reported results show that prompt-level remediation improves some models and tasks substantially while leaving persistent failures for others. The paper explicitly identifies the relative effect of fine-tuning as a separate unanswered question.

References

The three conditions tested here are prompt-level interventions; whether fine-tuning produces different results is a separate question this evaluation does not address.

— Context-Aware Pre-Deployment Evaluation of AI Systems: A Regulatory Framework for Nigerian Fintech  (2609.24016 - Uduimoh et al., 21 Sep 2026) in Section 4, subsection “Limitations”