Papers
Topics
Authors
Recent
Search
2000 character limit reached

OR-Judge: Continuous LLM Evaluation

Updated 8 July 2026
  • OR-Judge is a continuous evaluation method that detects and attributes quality drift in LLM products by distinguishing between system changes and judge updates.
  • It employs a dual-monitoring strategy with a cheap continuous judge and a strong periodic judge augmented by a fixed human-labeled anchor set to resolve attribution ambiguity.
  • Empirical results show high detection accuracy and lower false-alarm rates compared to traditional methods, ensuring robust, cost-effective performance.

OR-Judge is a continuous evaluation method for LLM products that addresses a specific ambiguity in LLM-as-a-judge pipelines: when a monitored score drifts, the change may come from the evaluated system or from the judge itself. It augments a main system monitor with a fixed, human-labeled anchor set that the current judge re-scores at a steady interleave, then combines the two signals through a guard-window rule that returns exactly one of {none,system,judge}\{\texttt{none}, \texttt{system}, \texttt{judge}\}. The method is designed for detection and attribution under anytime-valid error control, rather than for drift detection alone (Li, 13 Jun 2026).

1. Problem formulation and deployment setting

OR-Judge is built for continuous LLM evaluation in which a dashboard reports that “quality drifted,” but the monitored score is produced by an LLM judge that is itself mutable through a version bump or a scoring-prompt update (Li, 13 Jun 2026). In the deployment model, items arrive as a stream i=1,2,…i=1,2,\dots, each with rubric scores in [0,1]R[0,1]^R and a stratum label t∈{1,…,K}t \in \{1,\dots,K\}. The pipeline uses a cheap judge xix_i on every item and a strong judge yiy_i on only a sampled subset, under budget. The main monitor treats yy as ground truth for system drift, but that assumption fails when the strong judge changes silently.

The central ambiguity is therefore attributional rather than merely statistical. If the main monitor fires, the event is compatible with two distinct causes: the system may have regressed or improved, or the judge may have changed scoring behavior. OR-Judge treats this as a contamination problem inside the evaluation stack itself. Its goal is to determine not only whether drift occurred, but whether the observed drift should be attributed to the system or to the judge (Li, 13 Jun 2026).

This problem sits within a broader reliability agenda in LLM-as-a-Judge research. Survey treatments characterize judge reliability in terms of consistency, bias, robustness, calibration, and adaptability across assessment scenarios (Gu et al., 2024). Other work isolates format-induced inconsistency as a routing problem in which a shared internal judgment signal is mapped through fragile, format-specific terminal branches (Feldhus et al., 15 May 2026), while bias-focused studies document susceptibility to position, verbosity, authority, self-enhancement, and related effects (Ye et al., 2024). OR-Judge addresses a different but operationally central failure mode: silent drift of the judge during live monitoring.

2. Main system monitor

The main layer of OR-Judge is a cost-aware system monitor based on per-(rubric×stratum)(\text{rubric} \times \text{stratum}) prediction-powered inference (Li, 13 Jun 2026). For each cell (r,t)(r,t), it tests the one-sided null

H0sys:E[yi[r]∣i∈cell (r,t)]≥q0[r,t]−0.05,H_0^{\mathrm{sys}}:\quad \mathbb{E}[y_i[r] \mid i \in \text{cell } (r,t)] \ge q_0[r,t] - 0.05,

where i=1,2,…i=1,2,\dots0 is a held-out baseline from in-control data and the i=1,2,…i=1,2,\dots1 slack absorbs calibration noise.

The cheap judge provides side information, while only a subset of strong labels is queried. The monitoring statistic is a prediction-powered e-process built as a nonnegative supermartingale, with predictable sampling probabilities i=1,2,…i=1,2,\dots2 and a bet cap i=1,2,…i=1,2,\dots3. An alarm fires when any cell crosses a Bonferroni threshold, and validity follows from Ville’s inequality together with the underlying PPI e-process theorem (Li, 13 Jun 2026).

A structural detail emphasized in the method is stratification. The paper states that stratification is essential because localized blind-spot drift can disappear when pooled. OR-Judge therefore does not treat monitoring as a single aggregate test over all traffic; it preserves rubric- and stratum-level structure so that localized regressions remain detectable (Li, 13 Jun 2026).

3. Fixed human-labeled anchors and the judge monitor

To resolve ambiguity in the main process, OR-Judge introduces a fixed held-out anchor set

i=1,2,…i=1,2,\dots4

with frozen human labels i=1,2,…i=1,2,\dots5, chosen before monitoring begins and excluded from main calibration (Li, 13 Jun 2026). At every i=1,2,…i=1,2,\dots6, one anchor is sampled and re-scored by the current strong judge. Because the anchor items and human labels are fixed, the resulting process is insensitive to system drift.

The paper defines the rescaled anchor-gap statistic

i=1,2,…i=1,2,\dots7

with baseline i=1,2,…i=1,2,\dots8 equal to the exact mean rescaled gap of the time-0 judge over the full anchor set (Li, 13 Jun 2026). Conceptually, system drift changes incoming stream items, but anchors are frozen; therefore only the judge can move anchor gaps.

The anchor family then runs a betting e-process for each rubric i=1,2,…i=1,2,\dots9 and direction [0,1]R[0,1]^R0:

[0,1]R[0,1]^R1

with predictable [0,1]R[0,1]^R2 constrained by

[0,1]R[0,1]^R3

Wealth is accumulated multiplicatively,

[0,1]R[0,1]^R4

and the family alarms when

[0,1]R[0,1]^R5

The two directions encode whether the judge has become more lenient or more harsh relative to baseline (Li, 13 Jun 2026).

4. Guard-window race and verdict rule

OR-Judge converts the two monitor families into a single attribution rule. Let [0,1]R[0,1]^R6 be the first alarm time of the main system family, [0,1]R[0,1]^R7 the first alarm time of the anchor family, and [0,1]R[0,1]^R8 the guard window. The verdict is

[0,1]R[0,1]^R9

Thus, if neither process fires, the verdict is t∈{1,…,K}t \in \{1,\dots,K\}0; if anchors fire first, or within the guard window after the system family, the verdict is t∈{1,…,K}t \in \{1,\dots,K\}1; otherwise it is t∈{1,…,K}t \in \{1,\dots,K\}2 (Li, 13 Jun 2026).

The guard window exists because a changed judge can contaminate the main system monitor and make it fire first. The window gives anchors time to “win the race.” The design rule is asymmetric by construction: when both families fire inside the window, the verdict is judge, because once the judge is contaminated the system alarm is not trusted as a clean attribution (Li, 13 Jun 2026).

This produces a two-layer monitor: the main layer detects that something drifted, the anchor layer detects that the judge changed, and the guard-window race turns those facts into a categorical operational decision (Li, 13 Jun 2026).

5. Formal guarantees

The anchor e-process is a nonnegative supermartingale under the judge-null, using boundedness of t∈{1,…,K}t \in \{1,\dots,K\}3, directional bet caps, predictability of t∈{1,…,K}t \in \{1,\dots,K\}4, and the martingale condition t∈{1,…,K}t \in \{1,\dots,K\}5 (Li, 13 Jun 2026). Ville’s inequality then yields the time-uniform guarantee

t∈{1,…,K}t \in \{1,\dots,K\}6

This is the method’s anytime-validity guarantee, contrasted in the paper with repeated rolling t∈{1,…,K}t \in \{1,\dots,K\}7-tests.

A second guarantee is one-way identification. Under pure system drift, the anchor distribution is unchanged because anchors depend only on the fixed items t∈{1,…,K}t \in \{1,\dots,K\}8, frozen labels t∈{1,…,K}t \in \{1,\dots,K\}9, and the current judge. The paper states that under pure system drift,

xix_i0

uniformly in xix_i1, and that under no drift at all,

xix_i2

This is the method’s “one-way mirror”: the anchors can see judge drift, but system drift cannot move them (Li, 13 Jun 2026).

The attribution race is characterized exactly under judge drift. Misattribution to the system is

xix_i3

The paper proves that this probability is non-increasing in xix_i4, non-increasing in anchor budget xix_i5, non-increasing in anchor interleave rate xix_i6, and non-decreasing in main-process power. The associated design law is that the anchors must out-run the main process they guard (Li, 13 Jun 2026).

OR-Judge also proves process orthogonality. The anchor process is independent of the main-policy configuration, main budget, and acquisition schedule xix_i7, while the main process does not depend on anchor interleave because anchor calls are out-of-band and excluded from main calibration. Consequently, anchor budget xix_i8 and xix_i9 control judge attribution latency, main configuration controls system detection power, and the two families carry separate Bonferroni budgets (Li, 13 Jun 2026).

6. Empirical behavior, cost, and significance

The empirical study uses two datasets, HelpSteer2 and TL;DR summarization, and two real judge changes: a silent version bump from gemini-3.1-pro-preview to 3.5-flash, and a harsher prompt change, v2-strict (Li, 13 Jun 2026). On the lenient version bump, the anchor monitor detects judge drift cleanly: with yiy_i0 and rate yiy_i1, detections are yiy_i2 on HelpSteer2 and yiy_i3 on TL;DR, with judge-to-system misattribution equal to yiy_i4 on both datasets (Li, 13 Jun 2026). The abstract summarizes this result as detection of the silent version bump in yiy_i5 runs with zero judge-to-system misattribution (Li, 13 Jun 2026).

The stricter prompt change is the contaminated case in which the harsher judge can spuriously trigger the main system monitor. On HelpSteer2, judge-to-system spill decreases monotonically as the guard window grows, from yiy_i6 to yiy_i7 to yiy_i8 out of yiy_i9 over yy0; on TL;DR it decreases from yy1 to yy2 to yy3 (Li, 13 Jun 2026). The abstract reports that the contaminating strict-prompt change is correctly attributed on yy4 of yy5 runs at guard width yy6, and that attribution becomes perfect on TL;DR, where the stronger prompt shift makes the anchors fire faster, yielding yy7 correct judge attributions (Li, 13 Jun 2026).

The paper also benchmarks OR-Judge against classical alternatives on the same anchor stream. The industry-default rolling yy8-test false-alarms on yy9 of drift-free streams on HelpSteer2 and (rubric×stratum)(\text{rubric} \times \text{stratum})0 on TL;DR (Li, 13 Jun 2026). A calibrated Page–Hinkley detector controls false alarms better, at (rubric×stratum)(\text{rubric} \times \text{stratum})1 and (rubric×stratum)(\text{rubric} \times \text{stratum})2, but detects the subtle lenient bump only (rubric×stratum)(\text{rubric} \times \text{stratum})3 of the time on HelpSteer2 and (rubric×stratum)(\text{rubric} \times \text{stratum})4 on TL;DR (Li, 13 Jun 2026). The anchor e-process is reported to have false-alarm rate around (rubric×stratum)(\text{rubric} \times \text{stratum})5, without a calibration-stream requirement, while also improving detection and latency on the studied drifts (Li, 13 Jun 2026).

Cost is treated explicitly. Relative to strong-evaluating everything,

(rubric×stratum)(\text{rubric} \times \text{stratum})6

where (rubric×stratum)(\text{rubric} \times \text{stratum})7 and (rubric×stratum)(\text{rubric} \times \text{stratum})8 are the costs of the cheap and strong judges, (rubric×stratum)(\text{rubric} \times \text{stratum})9 is the number of processed items, and (r,t)(r,t)0 is the number of strong calls (Li, 13 Jun 2026). The abstract reports that the monitor runs at approximately (r,t)(r,t)1 of the cost of strong-judging every item, or (r,t)(r,t)2 in a cheaper-but-deafer regime (Li, 13 Jun 2026).

Within the broader LLM-as-a-Judge literature, OR-Judge addresses a deployment-level instability that prompt design, debiasing, or committee aggregation do not directly resolve. The literature documents prompt- and format-induced inconsistency (Feldhus et al., 15 May 2026), systematic judge bias that can reverse leaderboard rankings (Zhu et al., 2 Mar 2026), and broad bias taxonomies with automated robustness evaluation (Ye et al., 2024). OR-Judge contributes a distinct capability: anytime-valid attribution of observed score drift to the system or to the judge in a live evaluation pipeline (Li, 13 Jun 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to OR-Judge.