Gauge-Sweep Audits Overview
- Gauge-Sweep Audits are audit procedures that dynamically adjust a scalar gauge across structured hypotheses to optimize performance and retain formal risk limits.
- They are applied across domains including ballot-level risk-limiting audits, DP-SGD privacy auditing, checkpoint verification, and metrology side-channel tests.
- Implementations use techniques like mixture martingales, adaptive betting, and side-channel profiling to ensure robust audit guarantees while reducing expected workload.
Gauge-Sweep Audits are audit procedures in which a “gauge” is swept, mixed, accumulated, or adapted over a structured family of hypotheses, latent states, or internal signals. In the supplied literature, the term appears most concretely in ballot-level comparison risk-limiting audits built on comparison-optimal betting, where the gauge is the martingale tuning parameter or bet; it is also used as an interpretive label for one-run privacy auditing in DP-SGD, artifact-level checkpoint triage, profiled side-channel auditing of public metrology releases, and a SHANGRLA-style mapping from discrepancies to half-average null tests. Across these settings, the common pattern is to preserve a formal audit guarantee while improving robustness or efficiency when the true generative or error process is unknown (Spertus, 2023, Agrawal et al., 10 Jun 2026, Hurtado, 2 Jul 2026, Alpay et al., 1 Jun 2026, Stark, 2019).
1. Domain-specific meanings of “gauge” and “sweep”
The term is not used uniformly across the cited work. In ballot-level comparison RLAs, the gauge is the per-draw bet in a supermartingale test, and sweeping means diversifying or adapting that bet across plausible cast-vote-record error rates. In the supplied terminology mapping for one-run privacy auditing, the gauge is the normalized canary score and the sweep is the accumulation of canary-aligned signals over training steps. In checkpoint auditing, gauges are internal artifact-level signals computed from activations and weights, and sweeping means applying them across a registry of checkpoints and combining them by standardized fusion. In metrology, gauge settings are instrument and processing parameters such as band edge, windowing, overlap, and segment count, and sweeping means varying those settings to map leakage. In the SHANGRLA connection, the gauge is the mapping that converts ballot-level discrepancies into nonnegative quantities fitting a half-average-null framework, and the sweep is the systematic enumeration and testing of the finite set of assertions needed for outcome verification (Spertus, 2023, Agrawal et al., 10 Jun 2026, Hurtado, 2 Jul 2026, Alpay et al., 1 Jun 2026, Stark, 2019).
| Domain | Gauge | Sweep |
|---|---|---|
| Comparison RLAs | Per-draw bet | Mixtures or adaptation across hypothesized CVR error rates |
| One-run DP auditing | Canary score | Summing canary-aligned signals across DP-SGD steps |
| Checkpoint auditing | Activation refusal-gap and weight-recovery energy | Sweeping complementary gauges across checkpoints and combining them |
| Metrology side-channel auditing | Instrument and processing settings; protected residual norm | Sweeping chemistry and release-map settings |
| SHANGRLA mapping | Assorter or discrepancy-to-list transform | Enumerating and testing all required assertions |
A plausible implication is that “Gauge-Sweep Audit” names a methodological schema rather than a single fixed algorithm. What remains invariant is the use of an internal scalar or low-dimensional statistic whose evolution under a null and an alternative can be controlled, optimized, or profiled.
2. Comparison-optimal betting in ballot-level comparison risk-limiting audits
In COBRA, gauge-sweep audits are risk-limiting audits that “sweep,” mix, or adapt the martingale tuning parameter—called the gauge or bet—across a range of hypothesized CVR error rates. For a plurality contest with ballots, CVRs , human interpretations , and assorter taking values in , the overstatement in favor of the reported winner is . With diluted margin
define
when understatement errors are ignored. The three categories are correct CVR 0, 1-vote overstatement 1, and 2-vote overstatement 2 with probabilities 3. The outcome is correct iff 4, so the complementary null is 5.
For sequential sampling with replacement, COBRA uses the supermartingale update
6
where 7 is predictable. Under 8, 9 is a nonnegative supermartingale, and Ville’s inequality yields
0
The audit stops and confirms the outcome when 1; the truncated reciprocal 2 is a sequentially valid 3-value.
COBRA’s central result is that a fixed bet 4 should maximize expected log-growth under the Kelly criterion,
5
The first-order condition is
6
In the special case 7,
8
This oracle-optimal 9 depends on the true CVR error rates. When 0, maximal betting 1 eventually “goes broke” with probability 2 because drawing 3 yields 4.
Gauge-sweep enters through two constructions. In diversified betting, one selects a grid of plausible 5, solves for 6 at each point, forms component martingales
7
and mixes them as
8
Because mixtures of supermartingales are supermartingales, risk control is unchanged. In adaptive betting, error-rate estimates from previously sampled ballots are plugged into the Kelly condition through truncated-shrinkage estimators
9
and the resulting 0 remains valid because it is a predictable function of the past.
The efficiency gains reported for these gauge-sweep strategies are substantial. Across simulated audits with diluted margins 1 and 2, comparison-optimal oracle bets reduced expected workload by 3 on average relative to a ballot-polling apKelly bet transplanted to comparison audits. At 4 and 5, apKelly often needed the entire population of 6 draws, whereas 7 stopped in mean 8 draws with 9th percentile 0. For practical strategies at 1 and 2, geometric mean workload ratios relative to oracle were fixed 3, adaptive 4, and diversified 5; mixtures were best on average. COBRA therefore recommends comparison-optimal gauge-sweep strategies rather than fixed a priori bets, small positive 6 to avoid stalls, and caution when sampling without replacement not to naively subtract observed errors from hypothesized population rates (Spertus, 2023).
3. SHANGRLA and assertion sweeps over half-average nulls
SHANGRLA reduces auditing for many social choice functions to testing a finite set of null hypotheses of the form “the average of this list is not greater than 7.” For a finite list 8 of nonnegative numbers, the canonical null is
9
Rejecting all such nulls at level 0 confirms the reported outcome with risk limit 1. This reduction applies to majority, super-majority, plurality, multi-winner plurality, Instant Runoff Voting, Borda count, approval voting, and STAR-Voting, among others.
For ballot polling, SHANGRLA works directly with assorters. For ballot-level comparison audits, it maps each discrepancy into a nonnegative quantity that again fits the half-average-null form. If 2 is an assorter mapping ballots or CVRs to 3, define 4, reported assorter mean 5, and reported margin 6. Let
7
Then 8 iff the underlying assertion is correct. In this sense, the gauge is the discrepancy-to-list transform 9, and the sweep is the systematic testing of every assertion required by the social choice rule.
SHANGRLA provides two sequential risk-measuring functions for sampling without replacement. The first is the Kaplan–Kolmogorov multiplicative martingale with a nonnegativity shift 0, yielding a sequential 1-value 2. The second is the integrated-polynomial martingale, yielding 3. Either can test each half-average null for either ballot polling or comparison auditing. Stratified audits are handled through a union–intersection construction over allocations 4 with 5, combining stratum-specific 6-values by Fisher’s method.
The framework also includes “phantoms to evil zombies” for missing ballot cards and missing or redacted CVRs, and ballot-style manifests for contests that do not appear on every ballot card. These devices preserve the risk limit by treating missing or misclassified items in the least favorable manner. SHANGRLA’s reported efficiency gains over previous comparison audits come from testing conditions that are both necessary and sufficient for correctness for most social choice functions, and from avoiding a conservative approximation used in earlier MACRO-style methods (Stark, 2019).
4. One-run Gaussian privacy auditing in DP-SGD
In the supplied terminology mapping for one-run privacy auditing, a Gauge-Sweep Audit is the paper’s one-run, white-box Gaussian audit. The gauge corresponds to the canary score
7
a normalized sum of canary-aligned per-step observations, and the sweep corresponds to accumulating those observations across all DP-SGD steps. For a canary 8, clipped canary gradient
9
and unit direction 0 when 1, the scalar signal is
2
where 3 is the privatized average gradient at step 4.
Under the paper’s idealized scalar model, if 5 indicates whether the canary is in the minibatch and 6 is Gaussian noise, then the absent and present worlds are modeled by
7
Accordingly,
8
while
9
and 0 is approximately Gaussian by the CLT. The Kolmogorov distance bound scales as
1
with a tail-localized bound that is negligible relative to 2 in the regime 3.
The audit can be used either as a one-sided membership test or, in the preferred formulation, as a Gaussian-pair fit. For a single threshold 4, the 5-value is
6
and the power is
7
More centrally, one fits
8
computes the hockey-stick divergence between Gaussians, and defines
9
with confidence-adjusted lower bound
00
The paper uses either Bonferroni rectangles or bootstrap ellipsoids for 01.
The empirical results reported for DP-SGD on CIFAR-10 use WideResNet-16-4, batch size 02, augmentation 03, and 04 steps, with analytic-Gaussian accounting at 05 and theoretical 06. For the convergence-bound example, 07, 08, and 09. The empirical 10 lower bound is approximately 11, compared with approximately 12 for an f-DP one-run audit and approximately 13 for a thresholding-based one-run audit. With only 14 canaries, the Gaussian audit attains bounds that the baselines need approximately 15 canaries to reach. The method requires white-box access to per-step privatized gradients, assumes approximate normality and cross-canary near-orthogonality, and extends exactly to DP-FTRL when ancestor-path noise sums are Gaussian (Agrawal et al., 10 Jun 2026).
5. Two-signal artifact-level checkpoint audits
In the checkpoint-auditing instantiation, Gauge-Sweep Audits are artifact-level audits that score the checkpoint as an object rather than scoring generated text. The paper’s premise is explicit: runtime guards cannot determine whether a checkpoint’s refusal mechanism has been stripped because “they score generations, not the artifact.” The two gauges are a reference-anchored activation refusal-gap and a weight-recovery energy of the base-to-candidate weight difference.
Let 16 be an attested reference model and 17 the candidate checkpoint. Using a fixed harmful/benign contrast set of 18 paired prompts from AdvBench, JailbreakBench, and HarmBench, a mid-stack layer band 19, unit-norm refusal directions 20 learned on the reference, and last-token residual-stream activations 21, define the band-averaged activation gap
22
and the base-anchored ratio
23
The interpretation supplied is that 24 for intact refusal and 25 as refusal is removed.
For watched matrices 26 consisting of attention 27 and MLP 28 layers in the same band, define
29
and the rank-1 spectral energy
30
The two signals are negatively correlated, with Pearson 31, and are fused by
32
Thresholds are then chosen by maximizing Youden’s 33 on a calibration set.
The evaluation uses a 34-checkpoint registry spanning Qwen, DeepSeek-distilled Qwen, Llama, and Gemma, with an evaluation set of 35 uncensored checkpoints plus 36 benign edits; 37 labeled-but-refusing checkpoints were excluded. The combined 38-sum achieves in-sample AUROC 39 and PR 40. Under leave-one-family-out transfer, held-out detection is 41, balanced accuracy is 42, false-positive rate is 43, and the detector misses only 44 of 45 abliterations. Single-signal baselines are weaker: AUROC 46 for 47 alone and 48 for 49 alone. A fitted two-dimensional logistic does not improve over the threshold-free 50-sum.
The paper emphasizes that the audit is “effective triage, not tamper-proofing.” The most severe failure is a spoofed reference: declaring the candidate as its own reference yields 51 and 52 by construction, and declaring an abliterated sibling as the base can evade both axes. A second failure mode is white-box owner training past the threshold. In a Qwen2.5-1.5B proof-of-concept at step 53, the evaded checkpoint had 54 and 55, remained guard-unsafe on all 56 held-out harmful prompts, had substring refusal 57, and retained coherence and most capability. The audit therefore presumes an attested reference and is intended as pre-deployment triage with manual review rather than automatic rejection (Hurtado, 2 Jul 2026).
6. Profiled side-channel audits of public metrology releases
In the metrology instantiation, a Gauge-Sweep Audit is a profiled statistical side-channel test on public power-spectral-density releases in which instrument and processing gauge settings are swept to characterize, bound, and avoid protected-parameter leakage. The release map exposes finite-band PSD statistics derived from
58
with parameters 59. The protected coordinate is the effective transport length 60, linked to quencher loading by
61
so sweeping 62 traces the affine line 63.
Two channel models are used. In the independent-bin case, averaged PSD bins follow a gamma channel:
64
with exact divergence
65
and exact Chernoff information
66
When overlap or short records induce correlations, the audit switches to a covariance-weighted log-spectrum channel,
67
with local Chernoff equal to one quarter of this when the covariance is equal across the pair.
Amplitude and blur are treated as nuisance directions and eliminated by projection in log-spectrum space. With 68 and projection 69 onto 70, the protected Fisher information is
71
The headline finite-band law for 72 is
73
so for a local pair separated by 74,
75
This yields the closed-form safe band
76
For a protected quencher bit along the chemical line,
77
The operational protocol is explicit: estimate per-bin effective shapes 78 from repeats, test gamma goodness of fit, estimate and possibly shrink 79, fit the screened model in the log domain, build nuisance tangents, project to obtain the protected residual norm, compute the exact or covariance-weighted Chernoff exponent, and report 80 together with 81. If any gate fails—gamma fit, unstable covariance, or structured residuals—the exponent is deferred and the release should either restrict the band or gather more repeats.
The EUV roughness case study calibrates the theory at an 82-nm half-pitch scale using correlation length 83 nm, 84 nm, Nyquist 85 nm86, edge length 87 nm, 88, and 89. The reported per-spectrum Chernoff exponents and corresponding 90 show a sharp transition as the band admits the transport knee: at 91 nm92, 93 and 94; at 95 nm96, 97 and 98; at 99 nm00, 01 and 02. The same setting also shows that compressing the release to RMS alone destroys most transport signal: at 03 and 04, full-PSD maximum-likelihood success is approximately 05 versus approximately 06 for RMS-only releases. Positive AR(1)-like correlation can reduce the correlated/diagonal exponent ratio to approximately 07 at 08, so diagonal audits over-state leakage unless 09 is used (Alpay et al., 1 Jun 2026).