---
title: 'OR-Judge: Continuous LLM Evaluation'
url: https://www.emergentmind.com/topics/or-judge
type: topic
---

# OR-Judge: Continuous LLM Evaluation

OR-Judge is a continuous evaluation method for LLM products that addresses a specific ambiguity in LLM-as-a-judge pipelines: when a monitored score drifts, the change may come from the evaluated system or from the judge itself. It augments a main system monitor with a fixed, human-labeled anchor set that the current judge re-scores at a steady interleave, then combines the two signals through a guard-window rule that returns exactly one of $\{\texttt{none}, \texttt{system}, \texttt{judge}\}$. The method is designed for detection and attribution under anytime-valid error control, rather than for drift detection alone [2606.15474].

## 1. Problem formulation and deployment setting

OR-Judge is built for continuous LLM evaluation in which a dashboard reports that “quality drifted,” but the monitored score is produced by an LLM judge that is itself mutable through a version bump or a scoring-prompt update [2606.15474]. In the deployment model, items arrive as a stream $i=1,2,\dots$, each with rubric scores in $[0,1]^R$ and a stratum label $t \in \{1,\dots,K\}$. The pipeline uses a cheap judge $x_i$ on every item and a strong judge $y_i$ on only a sampled subset, under budget. The main monitor treats $y$ as ground truth for system drift, but that assumption fails when the strong judge changes silently.

The central ambiguity is therefore attributional rather than merely statistical. If the main monitor fires, the event is compatible with two distinct causes: the system may have regressed or improved, or the judge may have changed scoring behavior. OR-Judge treats this as a contamination problem inside the evaluation stack itself. Its goal is to determine not only whether drift occurred, but whether the observed drift should be attributed to the system or to the judge [2606.15474].

This problem sits within a broader reliability agenda in LLM-as-a-Judge research. Survey treatments characterize judge reliability in terms of consistency, bias, robustness, calibration, and adaptability across assessment scenarios [2411.15594]. Other work isolates format-induced inconsistency as a routing problem in which a shared internal judgment signal is mapped through fragile, format-specific terminal branches [2605.16023], while bias-focused studies document susceptibility to position, verbosity, authority, self-enhancement, and related effects [2410.02736]. OR-Judge addresses a different but operationally central failure mode: silent drift of the judge during live monitoring.

## 2. Main system monitor

The main layer of OR-Judge is a cost-aware system monitor based on per-$(\text{rubric} \times \text{stratum})$ prediction-powered inference [2606.15474]. For each cell $(r,t)$, it tests the one-sided null
$$
H_0^{\mathrm{sys}}:\quad \mathbb{E}[y_i[r] \mid i \in \text{cell } (r,t)] \ge q_0[r,t] - 0.05,
$$
where $q_0[r,t]$ is a held-out baseline from in-control data and the $0.05$ slack absorbs calibration noise.

The cheap judge provides side information, while only a subset of strong labels is queried. The monitoring statistic is a prediction-powered e-process built as a nonnegative supermartingale, with predictable sampling probabilities $\pi_i$ and a bet cap $\lambda_{\max}(\pi_i)$. An alarm fires when any cell crosses a Bonferroni threshold, and validity follows from Ville’s inequality together with the underlying PPI e-process theorem [2606.15474].

A structural detail emphasized in the method is stratification. The paper states that stratification is essential because localized blind-spot drift can disappear when pooled. OR-Judge therefore does not treat monitoring as a single aggregate test over all traffic; it preserves rubric- and stratum-level structure so that localized regressions remain detectable [2606.15474].

## 3. Fixed human-labeled anchors and the judge monitor

To resolve ambiguity in the main process, OR-Judge introduces a fixed held-out anchor set
$$
A=\{a_1,\dots,a_k\}
$$
with frozen human labels $h(a)\in[0,1]^R$, chosen before monitoring begins and excluded from main calibration [2606.15474]. At every $i \equiv 0 \pmod{\rho}$, one anchor is sampled and re-scored by the current strong judge. Because the anchor items and human labels are fixed, the resulting process is insensitive to system drift.

The paper defines the rescaled anchor-gap statistic
$$
g_i \;=\; \operatorname{clip}\!\Big(\tfrac{y_i(a) - h(a) + 1}{2},\, 0,\, 1\Big) \;\in\; [0,1]^R,
$$
with baseline $q\in(0,1)^R$ equal to the exact mean rescaled gap of the time-0 judge over the full anchor set [2606.15474]. Conceptually, system drift changes incoming stream items, but anchors are frozen; therefore only the judge can move anchor gaps.

The anchor family then runs a betting e-process for each rubric $r$ and direction $d \in \{\text{below},\text{above}\}$:
$$
e_i \;=\; 1 + \lambda_i\, s_d\,\big(q[r] - g_i[r]\big), \qquad s_{\text{below}}=+1,\; s_{\text{above}}=-1,
$$
with predictable $\lambda_i$ constrained by
$$
\lambda_i \in [0,\lambda_{\max}(q[r], d)].
$$
Wealth is accumulated multiplicatively,
$$
W_n^{(r,d)} = \prod_{i \le n} e_i,
$$
and the family alarms when
$$
\max_{r,d} W_n^{(r,d)} \ge \frac{2R}{\alpha}.
$$
The two directions encode whether the judge has become more lenient or more harsh relative to baseline [2606.15474].

## 4. Guard-window race and verdict rule

OR-Judge converts the two monitor families into a single attribution rule. Let $\tau_{\mathrm{sys}}$ be the first alarm time of the main system family, $\tau_{\mathrm{anc}}$ the first alarm time of the anchor family, and $W \ge 0$ the guard window. The verdict is
$$
\text{verdict} =
\begin{cases}
\texttt{none}   & \text{if neither family fires,}\\[2pt]
\texttt{judge}  & \text{if } \tau_{\mathrm{anc}} < \infty \text{ and } (\tau_{\mathrm{sys}} = \infty \text{ or } \tau_{\mathrm{anc}} \le \tau_{\mathrm{sys}} + W),\\[2pt]
\texttt{system} & \text{if } \tau_{\mathrm{sys}} < \infty \text{ and } (\tau_{\mathrm{anc}} = \infty \text{ or } \tau_{\mathrm{anc}} > \tau_{\mathrm{sys}} + W).
\end{cases}
$$
Thus, if neither process fires, the verdict is $\texttt{none}$; if anchors fire first, or within the guard window after the system family, the verdict is $\texttt{judge}$; otherwise it is $\texttt{system}$ [2606.15474].

The guard window exists because a changed judge can contaminate the main system monitor and make it fire first. The window gives anchors time to “win the race.” The design rule is asymmetric by construction: when both families fire inside the window, the verdict is judge, because once the judge is contaminated the system alarm is not trusted as a clean attribution [2606.15474].

This produces a two-layer monitor: the main layer detects that something drifted, the anchor layer detects that the judge changed, and the guard-window race turns those facts into a categorical operational decision [2606.15474].

## 5. Formal guarantees

The anchor e-process is a nonnegative supermartingale under the judge-null, using boundedness of $g_i[r]\in[0,1]$, directional bet caps, predictability of $\lambda_i$, and the martingale condition $\mathbb{E}[g_i[r]\mid \text{past}] = q[r]$ [2606.15474]. Ville’s inequality then yields the time-uniform guarantee
$$
\mathbb{P}\Big(\exists\, n : \max_{r,d} W_n^{(r,d)} \ge \frac{2R}{\alpha}\Big) \le \alpha.
$$
This is the method’s anytime-validity guarantee, contrasted in the paper with repeated rolling $z$-tests.

A second guarantee is one-way identification. Under pure system drift, the anchor distribution is unchanged because anchors depend only on the fixed items $A$, frozen labels $h$, and the current judge. The paper states that under pure system drift,
$$
\mathbb{P}(\text{verdict} = judge) \le \alpha
$$
uniformly in $W$, and that under no drift at all,
$$
\mathbb{P}(\text{verdict} \ne \texttt{none}) \le \alpha_{\mathrm{sys}} + \alpha_{\mathrm{anc}}.
$$
This is the method’s “one-way mirror”: the anchors can see judge drift, but system drift cannot move them [2606.15474].

The attribution race is characterized exactly under judge drift. Misattribution to the system is
$$
\mathbb{P}(\text{verdict} = system \mid \text{judge drift}) = \mathbb{P}(\tau_{\mathrm{sys}} + W < \tau_{\mathrm{anc}}).
$$
The paper proves that this probability is non-increasing in $W$, non-increasing in anchor budget $k$, non-increasing in anchor interleave rate $1/\rho$, and non-decreasing in main-process power. The associated design law is that the anchors must out-run the main process they guard [2606.15474].

OR-Judge also proves process orthogonality. The anchor process is independent of the main-policy configuration, main budget, and acquisition schedule $\pi$, while the main process does not depend on anchor interleave because anchor calls are out-of-band and excluded from main calibration. Consequently, anchor budget $(k,\rho)$ and $W$ control judge attribution latency, main configuration controls system detection power, and the two families carry separate Bonferroni budgets [2606.15474].

## 6. Empirical behavior, cost, and significance

The empirical study uses two datasets, HelpSteer2 and TL;DR summarization, and two real judge changes: a silent version bump from gemini-3.1-pro-preview to 3.5-flash, and a harsher prompt change, v2-strict [2606.15474]. On the lenient version bump, the anchor monitor detects judge drift cleanly: with $k=200$ and rate $1/5$, detections are $60/60$ on HelpSteer2 and $50/60$ on TL;DR, with judge-to-system misattribution equal to $0$ on both datasets [2606.15474]. The abstract summarizes this result as detection of the silent version bump in $60/60$ runs with zero judge-to-system misattribution [2606.15474].

The stricter prompt change is the contaminated case in which the harsher judge can spuriously trigger the main system monitor. On HelpSteer2, judge-to-system spill decreases monotonically as the guard window grows, from $29$ to $18$ to $10$ out of $120$ over $W=50,150,300$; on TL;DR it decreases from $5$ to $1$ to $0$ [2606.15474]. The abstract reports that the contaminating strict-prompt change is correctly attributed on $110$ of $120$ runs at guard width $300$, and that attribution becomes perfect on TL;DR, where the stronger prompt shift makes the anchors fire faster, yielding $240/240$ correct judge attributions [2606.15474].

The paper also benchmarks OR-Judge against classical alternatives on the same anchor stream. The industry-default rolling $z$-test false-alarms on $75\%$ of drift-free streams on HelpSteer2 and $67\%$ on TL;DR [2606.15474]. A calibrated Page–Hinkley detector controls false alarms better, at $0.08$ and $0.03$, but detects the subtle lenient bump only $8\%$ of the time on HelpSteer2 and $40\%$ on TL;DR [2606.15474]. The anchor e-process is reported to have false-alarm rate around $0.03$, without a calibration-stream requirement, while also improving detection and latency on the studied drifts [2606.15474].

Cost is treated explicitly. Relative to strong-evaluating everything,
$$
\text{cost-fraction} \;=\; \frac{L\, c_x + S\, c_y}{L\, c_y} \;=\; 0.125 + \frac{S}{L},
$$
where $c_x$ and $c_y$ are the costs of the cheap and strong judges, $L$ is the number of processed items, and $S$ is the number of strong calls [2606.15474]. The abstract reports that the monitor runs at approximately $0.64$ of the cost of strong-judging every item, or $0.21$ in a cheaper-but-deafer regime [2606.15474].

Within the broader LLM-as-a-Judge literature, OR-Judge addresses a deployment-level instability that prompt design, debiasing, or committee aggregation do not directly resolve. The literature documents prompt- and format-induced inconsistency [2605.16023], systematic judge bias that can reverse leaderboard rankings [2603.01865], and broad bias taxonomies with automated robustness evaluation [2410.02736]. OR-Judge contributes a distinct capability: anytime-valid attribution of observed score drift to the system or to the judge in a live evaluation pipeline [2606.15474].

Source: https://www.emergentmind.com/topics/or-judge