OR-Judge: Continuous LLM Evaluation
- OR-Judge is a continuous evaluation method that detects and attributes quality drift in LLM products by distinguishing between system changes and judge updates.
- It employs a dual-monitoring strategy with a cheap continuous judge and a strong periodic judge augmented by a fixed human-labeled anchor set to resolve attribution ambiguity.
- Empirical results show high detection accuracy and lower false-alarm rates compared to traditional methods, ensuring robust, cost-effective performance.
OR-Judge is a continuous evaluation method for LLM products that addresses a specific ambiguity in LLM-as-a-judge pipelines: when a monitored score drifts, the change may come from the evaluated system or from the judge itself. It augments a main system monitor with a fixed, human-labeled anchor set that the current judge re-scores at a steady interleave, then combines the two signals through a guard-window rule that returns exactly one of . The method is designed for detection and attribution under anytime-valid error control, rather than for drift detection alone (Li, 13 Jun 2026).
1. Problem formulation and deployment setting
OR-Judge is built for continuous LLM evaluation in which a dashboard reports that “quality drifted,” but the monitored score is produced by an LLM judge that is itself mutable through a version bump or a scoring-prompt update (Li, 13 Jun 2026). In the deployment model, items arrive as a stream , each with rubric scores in and a stratum label . The pipeline uses a cheap judge on every item and a strong judge on only a sampled subset, under budget. The main monitor treats as ground truth for system drift, but that assumption fails when the strong judge changes silently.
The central ambiguity is therefore attributional rather than merely statistical. If the main monitor fires, the event is compatible with two distinct causes: the system may have regressed or improved, or the judge may have changed scoring behavior. OR-Judge treats this as a contamination problem inside the evaluation stack itself. Its goal is to determine not only whether drift occurred, but whether the observed drift should be attributed to the system or to the judge (Li, 13 Jun 2026).
This problem sits within a broader reliability agenda in LLM-as-a-Judge research. Survey treatments characterize judge reliability in terms of consistency, bias, robustness, calibration, and adaptability across assessment scenarios (Gu et al., 2024). Other work isolates format-induced inconsistency as a routing problem in which a shared internal judgment signal is mapped through fragile, format-specific terminal branches (Feldhus et al., 15 May 2026), while bias-focused studies document susceptibility to position, verbosity, authority, self-enhancement, and related effects (Ye et al., 2024). OR-Judge addresses a different but operationally central failure mode: silent drift of the judge during live monitoring.
2. Main system monitor
The main layer of OR-Judge is a cost-aware system monitor based on per- prediction-powered inference (Li, 13 Jun 2026). For each cell , it tests the one-sided null
where 0 is a held-out baseline from in-control data and the 1 slack absorbs calibration noise.
The cheap judge provides side information, while only a subset of strong labels is queried. The monitoring statistic is a prediction-powered e-process built as a nonnegative supermartingale, with predictable sampling probabilities 2 and a bet cap 3. An alarm fires when any cell crosses a Bonferroni threshold, and validity follows from Ville’s inequality together with the underlying PPI e-process theorem (Li, 13 Jun 2026).
A structural detail emphasized in the method is stratification. The paper states that stratification is essential because localized blind-spot drift can disappear when pooled. OR-Judge therefore does not treat monitoring as a single aggregate test over all traffic; it preserves rubric- and stratum-level structure so that localized regressions remain detectable (Li, 13 Jun 2026).
3. Fixed human-labeled anchors and the judge monitor
To resolve ambiguity in the main process, OR-Judge introduces a fixed held-out anchor set
4
with frozen human labels 5, chosen before monitoring begins and excluded from main calibration (Li, 13 Jun 2026). At every 6, one anchor is sampled and re-scored by the current strong judge. Because the anchor items and human labels are fixed, the resulting process is insensitive to system drift.
The paper defines the rescaled anchor-gap statistic
7
with baseline 8 equal to the exact mean rescaled gap of the time-0 judge over the full anchor set (Li, 13 Jun 2026). Conceptually, system drift changes incoming stream items, but anchors are frozen; therefore only the judge can move anchor gaps.
The anchor family then runs a betting e-process for each rubric 9 and direction 0:
1
with predictable 2 constrained by
3
Wealth is accumulated multiplicatively,
4
and the family alarms when
5
The two directions encode whether the judge has become more lenient or more harsh relative to baseline (Li, 13 Jun 2026).
4. Guard-window race and verdict rule
OR-Judge converts the two monitor families into a single attribution rule. Let 6 be the first alarm time of the main system family, 7 the first alarm time of the anchor family, and 8 the guard window. The verdict is
9
Thus, if neither process fires, the verdict is 0; if anchors fire first, or within the guard window after the system family, the verdict is 1; otherwise it is 2 (Li, 13 Jun 2026).
The guard window exists because a changed judge can contaminate the main system monitor and make it fire first. The window gives anchors time to “win the race.” The design rule is asymmetric by construction: when both families fire inside the window, the verdict is judge, because once the judge is contaminated the system alarm is not trusted as a clean attribution (Li, 13 Jun 2026).
This produces a two-layer monitor: the main layer detects that something drifted, the anchor layer detects that the judge changed, and the guard-window race turns those facts into a categorical operational decision (Li, 13 Jun 2026).
5. Formal guarantees
The anchor e-process is a nonnegative supermartingale under the judge-null, using boundedness of 3, directional bet caps, predictability of 4, and the martingale condition 5 (Li, 13 Jun 2026). Ville’s inequality then yields the time-uniform guarantee
6
This is the method’s anytime-validity guarantee, contrasted in the paper with repeated rolling 7-tests.
A second guarantee is one-way identification. Under pure system drift, the anchor distribution is unchanged because anchors depend only on the fixed items 8, frozen labels 9, and the current judge. The paper states that under pure system drift,
0
uniformly in 1, and that under no drift at all,
2
This is the method’s “one-way mirror”: the anchors can see judge drift, but system drift cannot move them (Li, 13 Jun 2026).
The attribution race is characterized exactly under judge drift. Misattribution to the system is
3
The paper proves that this probability is non-increasing in 4, non-increasing in anchor budget 5, non-increasing in anchor interleave rate 6, and non-decreasing in main-process power. The associated design law is that the anchors must out-run the main process they guard (Li, 13 Jun 2026).
OR-Judge also proves process orthogonality. The anchor process is independent of the main-policy configuration, main budget, and acquisition schedule 7, while the main process does not depend on anchor interleave because anchor calls are out-of-band and excluded from main calibration. Consequently, anchor budget 8 and 9 control judge attribution latency, main configuration controls system detection power, and the two families carry separate Bonferroni budgets (Li, 13 Jun 2026).
6. Empirical behavior, cost, and significance
The empirical study uses two datasets, HelpSteer2 and TL;DR summarization, and two real judge changes: a silent version bump from gemini-3.1-pro-preview to 3.5-flash, and a harsher prompt change, v2-strict (Li, 13 Jun 2026). On the lenient version bump, the anchor monitor detects judge drift cleanly: with 0 and rate 1, detections are 2 on HelpSteer2 and 3 on TL;DR, with judge-to-system misattribution equal to 4 on both datasets (Li, 13 Jun 2026). The abstract summarizes this result as detection of the silent version bump in 5 runs with zero judge-to-system misattribution (Li, 13 Jun 2026).
The stricter prompt change is the contaminated case in which the harsher judge can spuriously trigger the main system monitor. On HelpSteer2, judge-to-system spill decreases monotonically as the guard window grows, from 6 to 7 to 8 out of 9 over 0; on TL;DR it decreases from 1 to 2 to 3 (Li, 13 Jun 2026). The abstract reports that the contaminating strict-prompt change is correctly attributed on 4 of 5 runs at guard width 6, and that attribution becomes perfect on TL;DR, where the stronger prompt shift makes the anchors fire faster, yielding 7 correct judge attributions (Li, 13 Jun 2026).
The paper also benchmarks OR-Judge against classical alternatives on the same anchor stream. The industry-default rolling 8-test false-alarms on 9 of drift-free streams on HelpSteer2 and 0 on TL;DR (Li, 13 Jun 2026). A calibrated Page–Hinkley detector controls false alarms better, at 1 and 2, but detects the subtle lenient bump only 3 of the time on HelpSteer2 and 4 on TL;DR (Li, 13 Jun 2026). The anchor e-process is reported to have false-alarm rate around 5, without a calibration-stream requirement, while also improving detection and latency on the studied drifts (Li, 13 Jun 2026).
Cost is treated explicitly. Relative to strong-evaluating everything,
6
where 7 and 8 are the costs of the cheap and strong judges, 9 is the number of processed items, and 0 is the number of strong calls (Li, 13 Jun 2026). The abstract reports that the monitor runs at approximately 1 of the cost of strong-judging every item, or 2 in a cheaper-but-deafer regime (Li, 13 Jun 2026).
Within the broader LLM-as-a-Judge literature, OR-Judge addresses a deployment-level instability that prompt design, debiasing, or committee aggregation do not directly resolve. The literature documents prompt- and format-induced inconsistency (Feldhus et al., 15 May 2026), systematic judge bias that can reverse leaderboard rankings (Zhu et al., 2 Mar 2026), and broad bias taxonomies with automated robustness evaluation (Ye et al., 2024). OR-Judge contributes a distinct capability: anytime-valid attribution of observed score drift to the system or to the judge in a live evaluation pipeline (Li, 13 Jun 2026).