Papers
Topics
Authors
Recent
Search
2000 character limit reached

Gauge-Sweep Audits Overview

Updated 6 July 2026
  • Gauge-Sweep Audits are audit procedures that dynamically adjust a scalar gauge across structured hypotheses to optimize performance and retain formal risk limits.
  • They are applied across domains including ballot-level risk-limiting audits, DP-SGD privacy auditing, checkpoint verification, and metrology side-channel tests.
  • Implementations use techniques like mixture martingales, adaptive betting, and side-channel profiling to ensure robust audit guarantees while reducing expected workload.

Gauge-Sweep Audits are audit procedures in which a “gauge” is swept, mixed, accumulated, or adapted over a structured family of hypotheses, latent states, or internal signals. In the supplied literature, the term appears most concretely in ballot-level comparison risk-limiting audits built on comparison-optimal betting, where the gauge is the martingale tuning parameter or bet; it is also used as an interpretive label for one-run privacy auditing in DP-SGD, artifact-level checkpoint triage, profiled side-channel auditing of public metrology releases, and a SHANGRLA-style mapping from discrepancies to half-average null tests. Across these settings, the common pattern is to preserve a formal audit guarantee while improving robustness or efficiency when the true generative or error process is unknown (Spertus, 2023, Agrawal et al., 10 Jun 2026, Hurtado, 2 Jul 2026, Alpay et al., 1 Jun 2026, Stark, 2019).

1. Domain-specific meanings of “gauge” and “sweep”

The term is not used uniformly across the cited work. In ballot-level comparison RLAs, the gauge is the per-draw bet in a supermartingale test, and sweeping means diversifying or adapting that bet across plausible cast-vote-record error rates. In the supplied terminology mapping for one-run privacy auditing, the gauge is the normalized canary score and the sweep is the accumulation of canary-aligned signals over training steps. In checkpoint auditing, gauges are internal artifact-level signals computed from activations and weights, and sweeping means applying them across a registry of checkpoints and combining them by standardized fusion. In metrology, gauge settings are instrument and processing parameters such as band edge, windowing, overlap, and segment count, and sweeping means varying those settings to map leakage. In the SHANGRLA connection, the gauge is the mapping that converts ballot-level discrepancies into nonnegative quantities fitting a half-average-null framework, and the sweep is the systematic enumeration and testing of the finite set of assertions needed for outcome verification (Spertus, 2023, Agrawal et al., 10 Jun 2026, Hurtado, 2 Jul 2026, Alpay et al., 1 Jun 2026, Stark, 2019).

Domain Gauge Sweep
Comparison RLAs Per-draw bet λi\lambda_i Mixtures or adaptation across hypothesized CVR error rates
One-run DP auditing Canary score STS_T Summing canary-aligned signals across DP-SGD steps
Checkpoint auditing Activation refusal-gap and weight-recovery energy Sweeping complementary gauges across checkpoints and combining them
Metrology side-channel auditing Instrument and processing settings; protected residual norm Sweeping chemistry and release-map settings
SHANGRLA mapping Assorter or discrepancy-to-list transform Enumerating and testing all required assertions

A plausible implication is that “Gauge-Sweep Audit” names a methodological schema rather than a single fixed algorithm. What remains invariant is the use of an internal scalar or low-dimensional statistic whose evolution under a null and an alternative can be controlled, optimized, or profiled.

2. Comparison-optimal betting in ballot-level comparison risk-limiting audits

In COBRA, gauge-sweep audits are risk-limiting audits that “sweep,” mix, or adapt the martingale tuning parameter—called the gauge or bet—across a range of hypothesized CVR error rates. For a plurality contest with NN ballots, CVRs cic_i, human interpretations bib_i, and assorter A()A(\cdot) taking values in [0,1][0,1], the overstatement in favor of the reported winner is ωi:=A(ci)A(bi)\omega_i := A(c_i)-A(b_i). With diluted margin

v:=2(N1iA(ci))1,v := 2\cdot\left(N^{-1}\sum_i A(c_i)\right)-1,

define

a:=(2v)1(0.5,1],xi:=1ωi2v{0,a/2,a}a := (2-v)^{-1}\in(0.5,1], \qquad x_i := \frac{1-\omega_i}{2-v}\in\{0,a/2,a\}

when understatement errors are ignored. The three categories are correct CVR STS_T0, 1-vote overstatement STS_T1, and 2-vote overstatement STS_T2 with probabilities STS_T3. The outcome is correct iff STS_T4, so the complementary null is STS_T5.

For sequential sampling with replacement, COBRA uses the supermartingale update

STS_T6

where STS_T7 is predictable. Under STS_T8, STS_T9 is a nonnegative supermartingale, and Ville’s inequality yields

NN0

The audit stops and confirms the outcome when NN1; the truncated reciprocal NN2 is a sequentially valid NN3-value.

COBRA’s central result is that a fixed bet NN4 should maximize expected log-growth under the Kelly criterion,

NN5

The first-order condition is

NN6

In the special case NN7,

NN8

This oracle-optimal NN9 depends on the true CVR error rates. When cic_i0, maximal betting cic_i1 eventually “goes broke” with probability cic_i2 because drawing cic_i3 yields cic_i4.

Gauge-sweep enters through two constructions. In diversified betting, one selects a grid of plausible cic_i5, solves for cic_i6 at each point, forms component martingales

cic_i7

and mixes them as

cic_i8

Because mixtures of supermartingales are supermartingales, risk control is unchanged. In adaptive betting, error-rate estimates from previously sampled ballots are plugged into the Kelly condition through truncated-shrinkage estimators

cic_i9

and the resulting bib_i0 remains valid because it is a predictable function of the past.

The efficiency gains reported for these gauge-sweep strategies are substantial. Across simulated audits with diluted margins bib_i1 and bib_i2, comparison-optimal oracle bets reduced expected workload by bib_i3 on average relative to a ballot-polling apKelly bet transplanted to comparison audits. At bib_i4 and bib_i5, apKelly often needed the entire population of bib_i6 draws, whereas bib_i7 stopped in mean bib_i8 draws with bib_i9th percentile A()A(\cdot)0. For practical strategies at A()A(\cdot)1 and A()A(\cdot)2, geometric mean workload ratios relative to oracle were fixed A()A(\cdot)3, adaptive A()A(\cdot)4, and diversified A()A(\cdot)5; mixtures were best on average. COBRA therefore recommends comparison-optimal gauge-sweep strategies rather than fixed a priori bets, small positive A()A(\cdot)6 to avoid stalls, and caution when sampling without replacement not to naively subtract observed errors from hypothesized population rates (Spertus, 2023).

3. SHANGRLA and assertion sweeps over half-average nulls

SHANGRLA reduces auditing for many social choice functions to testing a finite set of null hypotheses of the form “the average of this list is not greater than A()A(\cdot)7.” For a finite list A()A(\cdot)8 of nonnegative numbers, the canonical null is

A()A(\cdot)9

Rejecting all such nulls at level [0,1][0,1]0 confirms the reported outcome with risk limit [0,1][0,1]1. This reduction applies to majority, super-majority, plurality, multi-winner plurality, Instant Runoff Voting, Borda count, approval voting, and STAR-Voting, among others.

For ballot polling, SHANGRLA works directly with assorters. For ballot-level comparison audits, it maps each discrepancy into a nonnegative quantity that again fits the half-average-null form. If [0,1][0,1]2 is an assorter mapping ballots or CVRs to [0,1][0,1]3, define [0,1][0,1]4, reported assorter mean [0,1][0,1]5, and reported margin [0,1][0,1]6. Let

[0,1][0,1]7

Then [0,1][0,1]8 iff the underlying assertion is correct. In this sense, the gauge is the discrepancy-to-list transform [0,1][0,1]9, and the sweep is the systematic testing of every assertion required by the social choice rule.

SHANGRLA provides two sequential risk-measuring functions for sampling without replacement. The first is the Kaplan–Kolmogorov multiplicative martingale with a nonnegativity shift ωi:=A(ci)A(bi)\omega_i := A(c_i)-A(b_i)0, yielding a sequential ωi:=A(ci)A(bi)\omega_i := A(c_i)-A(b_i)1-value ωi:=A(ci)A(bi)\omega_i := A(c_i)-A(b_i)2. The second is the integrated-polynomial martingale, yielding ωi:=A(ci)A(bi)\omega_i := A(c_i)-A(b_i)3. Either can test each half-average null for either ballot polling or comparison auditing. Stratified audits are handled through a union–intersection construction over allocations ωi:=A(ci)A(bi)\omega_i := A(c_i)-A(b_i)4 with ωi:=A(ci)A(bi)\omega_i := A(c_i)-A(b_i)5, combining stratum-specific ωi:=A(ci)A(bi)\omega_i := A(c_i)-A(b_i)6-values by Fisher’s method.

The framework also includes “phantoms to evil zombies” for missing ballot cards and missing or redacted CVRs, and ballot-style manifests for contests that do not appear on every ballot card. These devices preserve the risk limit by treating missing or misclassified items in the least favorable manner. SHANGRLA’s reported efficiency gains over previous comparison audits come from testing conditions that are both necessary and sufficient for correctness for most social choice functions, and from avoiding a conservative approximation used in earlier MACRO-style methods (Stark, 2019).

4. One-run Gaussian privacy auditing in DP-SGD

In the supplied terminology mapping for one-run privacy auditing, a Gauge-Sweep Audit is the paper’s one-run, white-box Gaussian audit. The gauge corresponds to the canary score

ωi:=A(ci)A(bi)\omega_i := A(c_i)-A(b_i)7

a normalized sum of canary-aligned per-step observations, and the sweep corresponds to accumulating those observations across all DP-SGD steps. For a canary ωi:=A(ci)A(bi)\omega_i := A(c_i)-A(b_i)8, clipped canary gradient

ωi:=A(ci)A(bi)\omega_i := A(c_i)-A(b_i)9

and unit direction v:=2(N1iA(ci))1,v := 2\cdot\left(N^{-1}\sum_i A(c_i)\right)-1,0 when v:=2(N1iA(ci))1,v := 2\cdot\left(N^{-1}\sum_i A(c_i)\right)-1,1, the scalar signal is

v:=2(N1iA(ci))1,v := 2\cdot\left(N^{-1}\sum_i A(c_i)\right)-1,2

where v:=2(N1iA(ci))1,v := 2\cdot\left(N^{-1}\sum_i A(c_i)\right)-1,3 is the privatized average gradient at step v:=2(N1iA(ci))1,v := 2\cdot\left(N^{-1}\sum_i A(c_i)\right)-1,4.

Under the paper’s idealized scalar model, if v:=2(N1iA(ci))1,v := 2\cdot\left(N^{-1}\sum_i A(c_i)\right)-1,5 indicates whether the canary is in the minibatch and v:=2(N1iA(ci))1,v := 2\cdot\left(N^{-1}\sum_i A(c_i)\right)-1,6 is Gaussian noise, then the absent and present worlds are modeled by

v:=2(N1iA(ci))1,v := 2\cdot\left(N^{-1}\sum_i A(c_i)\right)-1,7

Accordingly,

v:=2(N1iA(ci))1,v := 2\cdot\left(N^{-1}\sum_i A(c_i)\right)-1,8

while

v:=2(N1iA(ci))1,v := 2\cdot\left(N^{-1}\sum_i A(c_i)\right)-1,9

and a:=(2v)1(0.5,1],xi:=1ωi2v{0,a/2,a}a := (2-v)^{-1}\in(0.5,1], \qquad x_i := \frac{1-\omega_i}{2-v}\in\{0,a/2,a\}0 is approximately Gaussian by the CLT. The Kolmogorov distance bound scales as

a:=(2v)1(0.5,1],xi:=1ωi2v{0,a/2,a}a := (2-v)^{-1}\in(0.5,1], \qquad x_i := \frac{1-\omega_i}{2-v}\in\{0,a/2,a\}1

with a tail-localized bound that is negligible relative to a:=(2v)1(0.5,1],xi:=1ωi2v{0,a/2,a}a := (2-v)^{-1}\in(0.5,1], \qquad x_i := \frac{1-\omega_i}{2-v}\in\{0,a/2,a\}2 in the regime a:=(2v)1(0.5,1],xi:=1ωi2v{0,a/2,a}a := (2-v)^{-1}\in(0.5,1], \qquad x_i := \frac{1-\omega_i}{2-v}\in\{0,a/2,a\}3.

The audit can be used either as a one-sided membership test or, in the preferred formulation, as a Gaussian-pair fit. For a single threshold a:=(2v)1(0.5,1],xi:=1ωi2v{0,a/2,a}a := (2-v)^{-1}\in(0.5,1], \qquad x_i := \frac{1-\omega_i}{2-v}\in\{0,a/2,a\}4, the a:=(2v)1(0.5,1],xi:=1ωi2v{0,a/2,a}a := (2-v)^{-1}\in(0.5,1], \qquad x_i := \frac{1-\omega_i}{2-v}\in\{0,a/2,a\}5-value is

a:=(2v)1(0.5,1],xi:=1ωi2v{0,a/2,a}a := (2-v)^{-1}\in(0.5,1], \qquad x_i := \frac{1-\omega_i}{2-v}\in\{0,a/2,a\}6

and the power is

a:=(2v)1(0.5,1],xi:=1ωi2v{0,a/2,a}a := (2-v)^{-1}\in(0.5,1], \qquad x_i := \frac{1-\omega_i}{2-v}\in\{0,a/2,a\}7

More centrally, one fits

a:=(2v)1(0.5,1],xi:=1ωi2v{0,a/2,a}a := (2-v)^{-1}\in(0.5,1], \qquad x_i := \frac{1-\omega_i}{2-v}\in\{0,a/2,a\}8

computes the hockey-stick divergence between Gaussians, and defines

a:=(2v)1(0.5,1],xi:=1ωi2v{0,a/2,a}a := (2-v)^{-1}\in(0.5,1], \qquad x_i := \frac{1-\omega_i}{2-v}\in\{0,a/2,a\}9

with confidence-adjusted lower bound

STS_T00

The paper uses either Bonferroni rectangles or bootstrap ellipsoids for STS_T01.

The empirical results reported for DP-SGD on CIFAR-10 use WideResNet-16-4, batch size STS_T02, augmentation STS_T03, and STS_T04 steps, with analytic-Gaussian accounting at STS_T05 and theoretical STS_T06. For the convergence-bound example, STS_T07, STS_T08, and STS_T09. The empirical STS_T10 lower bound is approximately STS_T11, compared with approximately STS_T12 for an f-DP one-run audit and approximately STS_T13 for a thresholding-based one-run audit. With only STS_T14 canaries, the Gaussian audit attains bounds that the baselines need approximately STS_T15 canaries to reach. The method requires white-box access to per-step privatized gradients, assumes approximate normality and cross-canary near-orthogonality, and extends exactly to DP-FTRL when ancestor-path noise sums are Gaussian (Agrawal et al., 10 Jun 2026).

5. Two-signal artifact-level checkpoint audits

In the checkpoint-auditing instantiation, Gauge-Sweep Audits are artifact-level audits that score the checkpoint as an object rather than scoring generated text. The paper’s premise is explicit: runtime guards cannot determine whether a checkpoint’s refusal mechanism has been stripped because “they score generations, not the artifact.” The two gauges are a reference-anchored activation refusal-gap and a weight-recovery energy of the base-to-candidate weight difference.

Let STS_T16 be an attested reference model and STS_T17 the candidate checkpoint. Using a fixed harmful/benign contrast set of STS_T18 paired prompts from AdvBench, JailbreakBench, and HarmBench, a mid-stack layer band STS_T19, unit-norm refusal directions STS_T20 learned on the reference, and last-token residual-stream activations STS_T21, define the band-averaged activation gap

STS_T22

and the base-anchored ratio

STS_T23

The interpretation supplied is that STS_T24 for intact refusal and STS_T25 as refusal is removed.

For watched matrices STS_T26 consisting of attention STS_T27 and MLP STS_T28 layers in the same band, define

STS_T29

and the rank-1 spectral energy

STS_T30

The two signals are negatively correlated, with Pearson STS_T31, and are fused by

STS_T32

Thresholds are then chosen by maximizing Youden’s STS_T33 on a calibration set.

The evaluation uses a STS_T34-checkpoint registry spanning Qwen, DeepSeek-distilled Qwen, Llama, and Gemma, with an evaluation set of STS_T35 uncensored checkpoints plus STS_T36 benign edits; STS_T37 labeled-but-refusing checkpoints were excluded. The combined STS_T38-sum achieves in-sample AUROC STS_T39 and PR STS_T40. Under leave-one-family-out transfer, held-out detection is STS_T41, balanced accuracy is STS_T42, false-positive rate is STS_T43, and the detector misses only STS_T44 of STS_T45 abliterations. Single-signal baselines are weaker: AUROC STS_T46 for STS_T47 alone and STS_T48 for STS_T49 alone. A fitted two-dimensional logistic does not improve over the threshold-free STS_T50-sum.

The paper emphasizes that the audit is “effective triage, not tamper-proofing.” The most severe failure is a spoofed reference: declaring the candidate as its own reference yields STS_T51 and STS_T52 by construction, and declaring an abliterated sibling as the base can evade both axes. A second failure mode is white-box owner training past the threshold. In a Qwen2.5-1.5B proof-of-concept at step STS_T53, the evaded checkpoint had STS_T54 and STS_T55, remained guard-unsafe on all STS_T56 held-out harmful prompts, had substring refusal STS_T57, and retained coherence and most capability. The audit therefore presumes an attested reference and is intended as pre-deployment triage with manual review rather than automatic rejection (Hurtado, 2 Jul 2026).

6. Profiled side-channel audits of public metrology releases

In the metrology instantiation, a Gauge-Sweep Audit is a profiled statistical side-channel test on public power-spectral-density releases in which instrument and processing gauge settings are swept to characterize, bound, and avoid protected-parameter leakage. The release map exposes finite-band PSD statistics derived from

STS_T58

with parameters STS_T59. The protected coordinate is the effective transport length STS_T60, linked to quencher loading by

STS_T61

so sweeping STS_T62 traces the affine line STS_T63.

Two channel models are used. In the independent-bin case, averaged PSD bins follow a gamma channel:

STS_T64

with exact divergence

STS_T65

and exact Chernoff information

STS_T66

When overlap or short records induce correlations, the audit switches to a covariance-weighted log-spectrum channel,

STS_T67

with local Chernoff equal to one quarter of this when the covariance is equal across the pair.

Amplitude and blur are treated as nuisance directions and eliminated by projection in log-spectrum space. With STS_T68 and projection STS_T69 onto STS_T70, the protected Fisher information is

STS_T71

The headline finite-band law for STS_T72 is

STS_T73

so for a local pair separated by STS_T74,

STS_T75

This yields the closed-form safe band

STS_T76

For a protected quencher bit along the chemical line,

STS_T77

The operational protocol is explicit: estimate per-bin effective shapes STS_T78 from repeats, test gamma goodness of fit, estimate and possibly shrink STS_T79, fit the screened model in the log domain, build nuisance tangents, project to obtain the protected residual norm, compute the exact or covariance-weighted Chernoff exponent, and report STS_T80 together with STS_T81. If any gate fails—gamma fit, unstable covariance, or structured residuals—the exponent is deferred and the release should either restrict the band or gather more repeats.

The EUV roughness case study calibrates the theory at an STS_T82-nm half-pitch scale using correlation length STS_T83 nm, STS_T84 nm, Nyquist STS_T85 nmSTS_T86, edge length STS_T87 nm, STS_T88, and STS_T89. The reported per-spectrum Chernoff exponents and corresponding STS_T90 show a sharp transition as the band admits the transport knee: at STS_T91 nmSTS_T92, STS_T93 and STS_T94; at STS_T95 nmSTS_T96, STS_T97 and STS_T98; at STS_T99 nmNN00, NN01 and NN02. The same setting also shows that compressing the release to RMS alone destroys most transport signal: at NN03 and NN04, full-PSD maximum-likelihood success is approximately NN05 versus approximately NN06 for RMS-only releases. Positive AR(1)-like correlation can reduce the correlated/diagonal exponent ratio to approximately NN07 at NN08, so diagonal audits over-state leakage unless NN09 is used (Alpay et al., 1 Jun 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Gauge-Sweep Audits.