Papers
Topics
Authors
Recent
Search
2000 character limit reached

Confidence Intervals for Rate Estimation with Importance Sampling in Autonomous Vehicle Evaluation

Published 4 Apr 2026 in stat.ME and stat.AP | (2604.03827v1)

Abstract: Accounting for both rare events and complex sampling presents challenges when quantifying uncertainty for rate estimation in autonomous vehicle performance evaluation. In this paper, we introduce a statistical formulation of this problem and develop a unified compound Poisson model framework for unbiased rate estimation through the Horvitz Thompson estimator. Though asymptotic theory for the model is available, the inference of confidence intervals (CIs) in the presence of rare events requires new investigation. We also advocate for a new monotonicity criterion for rate CIs--summing the rates of disjoint types of events should produce not only a higher point estimate but also higher confidence bounds than for the individual rates--that facilitates interpretability in real applications. We propose a novel exponential bootstrap (EB) method for CI construction based on a fiducial argument; it satisfies the monotonicity property, while novel extensions of some existing methods do not. Comprehensive numerical studies show that EB performs well for a wide range of settings relevant to our applications. Fast implementation of EB based on saddlepoint approximation is also developed, which may be of independent interest.

Summary

  • The paper introduces exponential bootstrap and weighted Gamma methods to construct monotonic confidence intervals for rare event rate estimation in AV testing.
  • It integrates a compound Poisson framework with importance sampling to address data sparsity and high-variance weighting in simulation-driven evaluations.
  • Empirical evaluations, enhanced by saddlepoint approximations, validate the robustness and practical utility of these methods in safety-critical AV assessments.

Confidence Intervals for Rare Event Rate Estimation under Importance Sampling in AV Testing

Problem Context and Statistical Framework

Autonomous Vehicle (AV) performance validation fundamentally depends on the accurate estimation of rare event occurrence rates—such as AV-induced traffic congestion or safety-critical failures—over extensive driving distances. The AV domain faces a pronounced needle-in-a-haystack challenge: relevant events can be extraordinarily infrequent (on the order of one per several million miles), requiring robust statistical inference even in data-sparse or heavily sampled settings.

To address these operational challenges, large-scale high-fidelity simulation is leveraged as an expedient complement to on-road testing. Importance sampling, guided by ML-based event likelihood prediction, is foundational for both (a) segment selection for simulation and (b) candidate event selection for human review. This two-stage probabilistic sampling pipeline is engineered to maximize signal yield for rare phenomena while minimizing resource expenditure.

The statistical estimand of central interest is the incident rate per million miles (IPMM) for a targeted event class, denoted θ\theta. Due to practical constraints, only a heavily sampled subset of the total driving corpus is subjected to simulation and subsequent human review.

The paper formalizes the resulting dataset with a semi-parametric compound Poisson (CP) model. The unbiased Horvitz-Thompson estimator is employed, incorporating both segment and event-level sampling probabilities, but quantifying the uncertainty via confidence intervals (CIs) remains nontrivial—owing to data sparsity, rare events, and high-variance weighting induced by importance sampling.

Unified Model: Importance Sampling and Compound Poisson Process

Candidate run segments are assigned features ViV_i. For each segment, simulation sampling occurs with probability s(Vi)s(V_i) (ML-driven), and downstream, candidate event sampling for human review occurs with probability h(Vi)h(V_i), potentially also ML-driven and conditional on simulation outcome.

Let BiB_i be a Bernoulli indicator for segments passing through both simulation and review (BiBernoulli(p(Vi))B_i \sim \text{Bernoulli}(p(V_i)) with p(Vi)=s(Vi)h(Vi)p(V_i) = s(V_i)h(V_i)), and let YiY_i be the binary outcome of interest (e.g., true AV-induced congestion) revealed by human review. The estimator is

θ^=i=1NBip(Vi)Yi,\hat\theta = \sum_{i=1}^N \frac{B_i}{p(V_i)} Y_i,

which, under mild assumptions, is a sum over a random number of i.i.d. contributions and thus follows a compound Poisson distribution.

This structure generalizes to multi-stage importance sampling, supports both discrete and continuous sampling probability regimes, and—under independence—offers mathematically tractable inference for θ\theta.

Challenges in Confidence Interval Construction

Traditional asymptotic methods (delta method, empirical likelihood, bootstrap) and existing specialized approaches for the weighted sum of Poissons (Gamma CIs, Poisson bootstrap) lack suitability in at least one respect: they (a) can yield non-monotone CIs under aggregation, violating logical interpretability, or (b) exhibit poor coverage under high-variance rare event scenarios. Existing gamma-based CIs may paradoxically produce lower confidence bounds for "aggregated" event rates that are less than those for any constituent subset—a key issue in interpretable safety validation.

Proposed Solution: Exponential Bootstrap with Monotonicity Guarantee

The paper introduces a new Exponential Bootstrap (EB) method, extending the weighted Gamma class of CIs using a fiducial argument. Its salient features include:

  • For discrete weights: it derives a novel weighted Gamma CI which, in contrast to classical Gamma CIs, provably satisfies the monotonicity property—i.e., aggregating categories cannot decrease confidence bounds.
  • For continuous importance weights: it introduces the EB method, which generates synthetic bootstrap replicates via sums of observed weights multiplied by i.i.d. exponential variables. Upper and lower CI bounds are determined by quantiles of sums including the "next weight," carefully estimated to reflect realistic sampling regimes.

Both the weighted Gamma and EB approaches ensure that CIs for the union of disjoint event types are at least as wide (no tighter on lower bound) as those for constituent types, preserving logical interpretability in applied settings. Figure 1

Figure 1: Empirical comparison of PB, GP2m, and EB2m for coverage error and CI width under uniform sampling (ViV_i0). All methods are equivalent here, mapping to Poisson CI performance.

Empirical Evaluation: Simulation and Practical Implications

A comprehensive experimental suite addresses: (1) single-stage and multi-stage importance sampling; (2) discrete vs continuous weights; (3) defensive vs greedy sampling policies (parametrized by ViV_i1 governing ViV_i2); and (4) robustness to model misspecification.

Key findings include:

  • Monotonicity: Only the weighted Gamma and EB approaches maintain strict monotonicity for all data realizations. Classical CIs for Poisson mixtures routinely violate this property.
  • Coverage: Poisson bootstrap (PB) CIs are shown to systematically under-cover, especially as the true event count becomes sparser or sample weights more variable. EB (with carefully selected "next weight" parameter) and the midpoint-Gamma extension (GP2m) maintain coverage close to the nominal level across a broad spectrum of sampling policies and budgets.
  • CI Width: In sparse data regimes, EB CIs are necessarily wider than PB due to maintained coverage—but this width gap closes rapidly as budget increases.
  • Robustness to Misspecification: The approach is robust to moderate deviations from the assumed power-law model for sampling weights, with material coverage loss only when sampling is highly adversarial or budgets are minuscule. Figure 2

    Figure 2: Coverage error and mean CI width for varying ViV_i3 (importance sampling quality). EB2m outperforms PB under rare event conditions, with particularly sharp differences at low budget ratios.

    Figure 3

    Figure 3: Histogram of sampling probabilities for candidate events from real AV simulation data revealing a heavy-tailed distribution, a challenging regime for inference.

Application to Real-World AV Test Data

In a confidential AV system evaluation, involving millions of simulations and tens of thousands of event reviews (budget ratio ViV_i4), the EB and weighted Gamma CIs delivered monotonic, appropriately wide, and interpretable CIs for congestion-related rates, reconciling rare but high-weighted event occurrences. By contrast, traditional Gamma CI extensions and PB produced anomalies violating monotonicity or under-coverage, respectively.

This is exemplified by the observation of 38 events in one category (low weight) and one event in another (high weight): only the proposed EB/weighted Gamma CIs properly reflect that the combined rate's CI lower bound exceeds those for the constituent categories, as demanded by logical aggregation.

Algorithmic Enhancements

The computational cost of EB—the need to estimate quantiles of high-dimensional convolutions of weighted exponentials—is addressed with a saddlepoint approximation, providing a highly accurate analytic solution, and enabling large-scale application in simulation-heavy AV pipelines. Figure 4

Figure 4: The empirical accuracy of the saddlepoint approximation for a three-component weighted Gamma variable, demonstrating close alignment with Monte Carlo simulation.

Theoretical and Practical Implications

The introduction of a monotonic, coverage-respecting CI methodology for importance-sampled rare event rate estimation directly advances safety-critical AV validation. Confidence intervals produced by the EB method and its weighted Gamma analog not only align with risk-averse safety engineering practices but are also readily interpretable by domain stakeholders, avoiding paradoxical results that complicate actionable decision making.

From a theoretical vantage, this work addresses a recognized gap in the inference for compound Poisson processes under high-variance, importance-sampled, and rare event regimes. The monotonicity property is non-trivial and previously unaddressed, and the fiducial-based EB construction, extended to complex multi-stage sampling, may foster further development in high-stakes rare event inference across domains.

Future Directions

Open avenues include improved data-driven or adaptive calibration of the "next weight" parameter, especially under adversarial sampling. Diagnosis and quantification of sampling quality (e.g., estimation of the ViV_i5 parameter) merit further study. The method is directly extendable to continuous outcome classes (beyond binary events) and difference-in-rates estimation; adaptation for high-frequency, high-dimensional AV metrics is promising.

Conclusion

The exponential bootstrap and weighted Gamma CIs for rare event rate estimation under multi-stage importance sampling advance the statistical toolkit for AV evaluation. Their formal monotonicity, robust empirical coverage, and computational feasibility across regimes of practical import mark them as essential tools for safety validation in the era of large-scale simulation and rare event analysis. Figure 5

Figure 5: Coverage and CI width for PB, GO2m, GP2m, EB2, and EB2m methods in two-stage importance sampling, confirming the consistency of findings from the single-stage setting.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.