Papers
Topics
Authors
Recent
Search
2000 character limit reached

Runs-Based Distribution-Free Charts

Updated 2 May 2026
  • The topic presents a nonparametric method that uses the sequential structure of runs from thresholded data to detect shifts in process location and distribution.
  • It leverages finite Markov chain imbedding and permutation arguments to derive exact control limits and performance measures like in-control ARL without relying on distributional assumptions.
  • Applications include detecting local anomalies and gradual shifts using binary statistics, making the approach robust, tuning-free, and adaptable to various SPC scenarios.

Runs-based distribution-free charts are a family of nonparametric statistical process control (SPC) tools designed to monitor shifts in the distributional location or structure of a process. These charts leverage the combinatorial structure of runs and patterns in binary sequences—typically created by thresholding continuous or discrete data—to offer exact finite-sample guarantees for performance metrics such as in-control average run length (ARL), without requiring knowledge of the underlying data distribution. Their theoretical foundations center on random permutation arguments and the finite Markov chain imbedding (FMCI) technique, enabling robust, tuning-free monitoring of both location changes and complex distributional departures (Wu, 17 Nov 2025, Wu, 2018, Kumar et al., 2017).

1. Fundamental Concepts and Definitions

Runs-based distribution-free charts operate by converting raw process data Y1,…,YnY_1, \ldots, Y_n into a sequence of binary indicators via a thresholding operation: Xi=1{Yi≥c}X_i = \mathbf{1}\{Y_i \geq c\} where cc is chosen such that Pr⁡(Y≥c)=p0\Pr(Y \geq c) = p_0 for a user-specified baseline p0p_0 (Wu, 17 Nov 2025). This binarization yields two summary statistics:

  • N1=∑i=1nXiN_1 = \sum_{i=1}^n X_i (number of successes)
  • N0=n−N1N_0 = n - N_1 (number of failures)

Key statistics underpinning runs-based charts include:

  • Number of success runs (RnR_n): the total maximal consecutive blocks of 1’s, computed as

Rn=∑i=1n1{Xi=1,Xi−1=0}R_n = \sum_{i=1}^n \mathbf{1}\{X_i = 1, X_{i-1} = 0\}

(with X0=0X_0 = 0).

  • Scan statistic (Xi=1{Yi≥c}X_i = \mathbf{1}\{Y_i \geq c\}0): for a window size Xi=1{Yi≥c}X_i = \mathbf{1}\{Y_i \geq c\}1, defined as

Xi=1{Yi≥c}X_i = \mathbf{1}\{Y_i \geq c\}2

  • Longest run statistic (Xi=1{Yi≥c}X_i = \mathbf{1}\{Y_i \geq c\}3-length of consecutive 1’s):

Xi=1{Yi≥c}X_i = \mathbf{1}\{Y_i \geq c\}4

Each of these statistics targets specific types of distributional departures: Xi=1{Yi≥c}X_i = \mathbf{1}\{Y_i \geq c\}5 is sensitive to general clustering, Xi=1{Yi≥c}X_i = \mathbf{1}\{Y_i \geq c\}6 detects local anomalies, and the longest-run is specialized for sustained shifts.

2. Distribution-Free Properties and Permutation Arguments

Conditioning on Xi=1{Yi≥c}X_i = \mathbf{1}\{Y_i \geq c\}7 (the total number of ones), Xi=1{Yi≥c}X_i = \mathbf{1}\{Y_i \geq c\}8 is uniformly distributed over all permutations of Xi=1{Yi≥c}X_i = \mathbf{1}\{Y_i \geq c\}9 ones and cc0 zeros. Thus, the conditional laws of cc1, cc2, and the longest run depend only on cc3, not on the underlying value or distribution of cc4. This yields exact finite-sample distribution-free inference under the null hypothesis (in-control process), as the type I error and in-control ARL are fixed independently of cc5 (Wu, 17 Nov 2025, Wu, 2018).

3. FMCI Techniques for Control Limit and Distribution Derivation

Construction of exact finite-sample control limits relies on the finite Markov chain imbedding (FMCI) technique. For specified runs or scan statistics and given cc6, FMCI models the pattern formation as a Markov process on an explicit state-space:

  • Runs Chart (cc7): The state space cc8 consists of the possible numbers of runs as zeros are inserted between blocks of ones.
  • Scan Statistic (cc9): The state space is constructed as the set of proper suffixes of potential pattern occurrences, plus an absorbing “pattern detected” state.

Control limits are derived as follows:

  • The lower control limit Pr⁡(Y≥c)=p0\Pr(Y \geq c) = p_00 is the largest Pr⁡(Y≥c)=p0\Pr(Y \geq c) = p_01 with Pr⁡(Y≥c)=p0\Pr(Y \geq c) = p_02.
  • The upper control limit for Pr⁡(Y≥c)=p0\Pr(Y \geq c) = p_03 is the smallest Pr⁡(Y≥c)=p0\Pr(Y \geq c) = p_04 for which Pr⁡(Y≥c)=p0\Pr(Y \geq c) = p_05. Randomization at the control boundary is used to achieve exactly the nominal false alarm rate Pr⁡(Y≥c)=p0\Pr(Y \geq c) = p_06 when discrete distributions do not permit precise thresholds (Wu, 17 Nov 2025).

4. Stepwise Implementation Procedure

The standard implementation proceeds as follows (Wu, 17 Nov 2025):

  1. Threshold Selection: Calculate Pr⁡(Y≥c)=p0\Pr(Y \geq c) = p_07 as the empirical Pr⁡(Y≥c)=p0\Pr(Y \geq c) = p_08 quantile of the Phase I sample, yielding Pr⁡(Y≥c)=p0\Pr(Y \geq c) = p_09.
  2. Binarization: Compute p0p_00, p0p_01, p0p_02.
  3. Chart Type Selection:
    • Runs Chart (R-1): Use p0p_03; signal if p0p_04.
    • Scan Chart (R-2): Select window p0p_05; signal if p0p_06.
  4. Randomization (if required): For discrete statistics, randomize on control boundary values to achieve the exact false alarm probability.
  5. Localization Diagnostics: Upon a signal, report the run or window indices and conditional p0p_07-values for post hoc analysis.

The procedure is directly supported by closed-form recurrences for transition matrices in FMCI, as detailed in the references of Fu–Lou (2003) and Fu–Lou–Wu (2012).

5. Generalizations: p0p_08-Run Patterns and Waiting Times

Runs-based charts can employ a broad class of patterns characterized as p0p_09-runs, as developed by Kumar & Upadhye (2017) (Kumar et al., 2017). Three canonical pattern families are used:

  • Setup I: At least N1=∑i=1nXiN_1 = \sum_{i=1}^n X_i0 and at most N1=∑i=1nXiN_1 = \sum_{i=1}^n X_i1 zeros, followed by at least N1=∑i=1nXiN_1 = \sum_{i=1}^n X_i2 ones.
  • Setup II: At least N1=∑i=1nXiN_1 = \sum_{i=1}^n X_i3 zeros, followed by at least N1=∑i=1nXiN_1 = \sum_{i=1}^n X_i4 and at most N1=∑i=1nXiN_1 = \sum_{i=1}^n X_i5 ones.
  • Setup III: At least N1=∑i=1nXiN_1 = \sum_{i=1}^n X_i6, at most N1=∑i=1nXiN_1 = \sum_{i=1}^n X_i7 zeros, then at least N1=∑i=1nXiN_1 = \sum_{i=1}^n X_i8, at most N1=∑i=1nXiN_1 = \sum_{i=1}^n X_i9 ones.

The waiting-time until first appearance of the chosen pattern is modeled as a geometric distribution with explicit parameter N0=n−N1N_0 = n - N_10 computable from the pattern parameters and the (unknown) success probability N0=n−N1N_0 = n - N_11. The in-control and out-of-control ARLs are then N0=n−N1N_0 = n - N_12 and N0=n−N1N_0 = n - N_13, respectively, allowing precise calibration of chart sensitivity (Kumar et al., 2017).

6. Performance Metrics and Comparative Results

The principal performance metrics are:

  • In-control Average Run Length (N0=n−N1N_0 = n - N_14): Under the null, N0=n−N1N_0 = n - N_15, exactly attainable using randomization.
  • Out-of-control Average Run Length (N0=n−N1N_0 = n - N_16): Calculable for specified alternatives (e.g., shift in N0=n−N1N_0 = n - N_17 or a location parameter).
  • Signal Probability: Type I error controlled at N0=n−N1N_0 = n - N_18, regardless of N0=n−N1N_0 = n - N_19.
  • Detection Power: Simulation studies show runs charts are effective for early or mid-sample location shifts; scan charts excel at detecting local clusters or window-matching anomalies.

Comparison to alternative nonparametric methods (Mann–Whitney, Kolmogorov–Smirnov, Empirical Likelihood Ratio charts) finds that runs-based approaches are robust to heavy tails and distributional skew, require no asymptotic theory or Gumbel approximations, and maintain exact finite-sample error control (Wu, 17 Nov 2025). Adjustment of parameters such as RnR_n0 and window size RnR_n1 tunes the OC sensitivity profile.

7. Numerical Illustrations and Tuning Recommendations

Concrete numerical examples validate the theoretical properties. For RnR_n2 and RnR_n3, applying FMCI transition matrices yields in-control probabilities for RnR_n4 at distinct values, and allows construction of control limits for fixed RnR_n5 (Wu, 17 Nov 2025). More extensive simulations under normal, exponential, and RnR_n6 distributions confirm that finite-sample ARL, false alarm, and power are insensitive to the underlying RnR_n7.

Empirical findings recommend:

  • Smaller threshold RnR_n8 and larger window RnR_n9 for enhanced detection of small shifts.
  • Intermediate parameters for optimal detection of moderate to large shifts.
  • Pattern selection in Rn=∑i=1n1{Xi=1,Xi−1=0}R_n = \sum_{i=1}^n \mathbf{1}\{X_i = 1, X_{i-1} = 0\}0-runs to balance Rn=∑i=1n1{Xi=1,Xi−1=0}R_n = \sum_{i=1}^n \mathbf{1}\{X_i = 1, X_{i-1} = 0\}1 and Rn=∑i=1n1{Xi=1,Xi−1=0}R_n = \sum_{i=1}^n \mathbf{1}\{X_i = 1, X_{i-1} = 0\}2 (Kumar et al., 2017, Wu, 2018).

References:

  • "Phase I Distribution-Free Control Charts for Individual Observations Using Runs and Patterns" (Wu, 17 Nov 2025)
  • "Distribution-free runs-based control charts" (Wu, 2018)
  • "Generalizations of Distributions Related to (Rn=∑i=1n1{Xi=1,Xi−1=0}R_n = \sum_{i=1}^n \mathbf{1}\{X_i = 1, X_{i-1} = 0\}3)-runs" (Kumar et al., 2017)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Runs-Based Distribution-Free Charts.