Characterize interactions and broader cue-strength effects

Characterize the interactions among co-occurring quality-preserving cues and the responses of multimodal large language model image-editing judges across a broader range of intervention strengths.

Background

EditJudgeBias evaluates a finite set of controlled cues at specified intervention strengths. The reported results therefore establish sensitivity for the tested cue types and doses, but do not determine how simultaneous cues interact or how judge responses change when intervention strengths vary beyond those settings.

The limitations section explicitly identifies these interactions and broader strength responses as unresolved aspects requiring further characterization. This is a distinct open problem from the benchmark’s primary question because it concerns generalization of the observed cue effects to combinations and dose ranges not evaluated in the study.

References

First, EditJudgeBias evaluates a finite set of controlled cues at specified intervention strengths; interactions among co-occurring cues and responses across a broader range of strengths remain to be characterized.

— Do MLLM Judges Judge the Edit? Auditing Bias in Image Editing Evaluation with Verified Quality Preservation  (2610.01670 - Huang et al., 1 Oct 2026) in Appendix, Section “Limitations”