---
title: Runs-Based Distribution-Free Charts
url: https://www.emergentmind.com/topics/runs-based-distribution-free-charts
type: topic
---

# Runs-Based Distribution-Free Charts

Runs-based distribution-free charts are a family of nonparametric statistical process control (SPC) tools designed to monitor shifts in the distributional location or structure of a process. These charts leverage the combinatorial structure of runs and patterns in binary sequences—typically created by thresholding continuous or discrete data—to offer exact finite-sample guarantees for performance metrics such as in-control average run length (ARL), without requiring knowledge of the underlying data distribution. Their theoretical foundations center on random permutation arguments and the finite Markov chain imbedding (FMCI) technique, enabling robust, tuning-free monitoring of both location changes and complex distributional departures [2511.13672, 1801.06532, 1707.08367].

## 1. Fundamental Concepts and Definitions

Runs-based distribution-free charts operate by converting raw process data $Y_1, \ldots, Y_n$ into a sequence of binary indicators via a thresholding operation:
\[
X_i = \mathbf{1}\{Y_i \geq c\}
\]
where $c$ is chosen such that $\Pr(Y \geq c) = p_0$ for a user-specified baseline $p_0$ [2511.13672]. This binarization yields two summary statistics:
- $N_1 = \sum_{i=1}^n X_i$ (number of successes)
- $N_0 = n - N_1$ (number of failures)

Key statistics underpinning runs-based charts include:
- **Number of success runs ($R_n$):** the total maximal consecutive blocks of 1’s, computed as
  \[
  R_n = \sum_{i=1}^n \mathbf{1}\{X_i = 1, X_{i-1} = 0\}
  \]
  (with $X_0 = 0$).
- **Scan statistic ($S_n(r)$):** for a window size $r$, defined as
  \[
  S_n(r) = \max_{1 \leq t \leq n - r + 1} \sum_{i = t}^{t + r - 1} X_i
  \]
- **Longest run statistic ($\max$-length of consecutive 1’s):** 
  \[
  \max\{\text{length of consecutive 1’s in } (X_1, \ldots, X_n)\}
  \]
  
Each of these statistics targets specific types of distributional departures: $R_n$ is sensitive to general clustering, $S_n(r)$ detects local anomalies, and the longest-run is specialized for sustained shifts.

## 2. Distribution-Free Properties and Permutation Arguments

Conditioning on $N_1 = n_1$ (the total number of ones), $(X_1, \ldots, X_n)$ is uniformly distributed over all permutations of $n_1$ ones and $n_0$ zeros. Thus, the conditional laws of $R_n$, $S_n(r)$, and the longest run depend only on $(n, n_1)$, not on the underlying value or distribution of $p = \Pr(Y \geq c)$. This yields exact finite-sample distribution-free inference under the null hypothesis (in-control process), as the type I error and in-control ARL are fixed independently of $F_0$ [2511.13672, 1801.06532].

## 3. FMCI Techniques for Control Limit and Distribution Derivation

Construction of exact finite-sample control limits relies on the finite Markov chain imbedding (FMCI) technique. For specified runs or scan statistics and given $N_1 = n_1$, FMCI models the pattern formation as a Markov process on an explicit state-space:
- **Runs Chart ($R_n$):** The state space $\Omega_1$ consists of the possible numbers of runs as zeros are inserted between blocks of ones.
- **Scan Statistic ($S_n(r)$):** The state space is constructed as the set of proper suffixes of potential pattern occurrences, plus an absorbing “pattern detected” state.

Control limits are derived as follows:
- The lower control limit $R_n(\alpha)$ is the largest $L$ with $\Pr(R_n \leq L \mid N_1 = n_1) \leq \alpha$.
- The upper control limit for $S_n(r)$ is the smallest $s$ for which $\Pr(S_n(r) \geq s \mid N_1 = n_1) \leq \alpha$.
Randomization at the control boundary is used to achieve exactly the nominal false alarm rate $\alpha$ when discrete distributions do not permit precise thresholds [2511.13672].

## 4. Stepwise Implementation Procedure

The standard implementation proceeds as follows [2511.13672]:
1. **Threshold Selection:** Calculate $c$ as the empirical $(1 - p_0)$ quantile of the Phase I sample, yielding $\Pr(Y \geq c) \approx p_0$.
2. **Binarization:** Compute $X_i = \mathbf{1}\{Y_i \geq c\}$, $N_1$, $N_0$.
3. **Chart Type Selection:**
   - **Runs Chart (R-1):** Use $R_n$; signal if $R_n \leq R_n(\alpha)$.
   - **Scan Chart (R-2):** Select window $r$; signal if $S_n(r) \geq S_n(r, \alpha)$.
4. **Randomization (if required):** For discrete statistics, randomize on control boundary values to achieve the exact false alarm probability.
5. **Localization Diagnostics:** Upon a signal, report the run or window indices and conditional $p$-values for post hoc analysis.

The procedure is directly supported by closed-form recurrences for transition matrices in FMCI, as detailed in the references of Fu–Lou (2003) and Fu–Lou–Wu (2012).

## 5. Generalizations: $(k_1, k_2)$-Run Patterns and Waiting Times

Runs-based charts can employ a broad class of patterns characterized as $(k_1, k_2)$-runs, as developed by Kumar & Upadhye (2017) [1707.08367]. Three canonical pattern families are used:
- **Setup I:** At least $\ell_1$ and at most $k_1$ zeros, followed by at least $\ell_2$ ones.
- **Setup II:** At least $\ell_1$ zeros, followed by at least $\ell_2$ and at most $k_2$ ones.
- **Setup III:** At least $\ell_1$, at most $k_1$ zeros, then at least $\ell_2$, at most $k_2$ ones.

The waiting-time until first appearance of the chosen pattern is modeled as a geometric distribution with explicit parameter $\alpha$ computable from the pattern parameters and the (unknown) success probability $p$. The in-control and out-of-control ARLs are then $1/\alpha(p_0)$ and $1/\alpha(p_1)$, respectively, allowing precise calibration of chart sensitivity [1707.08367].

## 6. Performance Metrics and Comparative Results

The principal performance metrics are:
- **In-control Average Run Length ($\mathrm{ARL}_0$):** Under the null, $\mathrm{ARL}_0 = 1/\alpha$, exactly attainable using randomization.
- **Out-of-control Average Run Length ($\mathrm{ARL}_1$):** Calculable for specified alternatives (e.g., shift in $p$ or a location parameter).
- **Signal Probability:** Type I error controlled at $\alpha$, regardless of $F_0$.
- **Detection Power:** Simulation studies show runs charts are effective for early or mid-sample location shifts; scan charts excel at detecting local clusters or window-matching anomalies.

Comparison to alternative nonparametric methods (Mann–Whitney, Kolmogorov–Smirnov, Empirical Likelihood Ratio charts) finds that runs-based approaches are robust to heavy tails and distributional skew, require no asymptotic theory or Gumbel approximations, and maintain exact finite-sample error control [2511.13672]. Adjustment of parameters such as $p_0$ and window size $r$ tunes the OC sensitivity profile.

## 7. Numerical Illustrations and Tuning Recommendations

Concrete numerical examples validate the theoretical properties. For $n = 5$ and $n_1 = 3$, applying FMCI transition matrices yields in-control probabilities for $R_5$ at distinct values, and allows construction of control limits for fixed $\alpha$ [2511.13672]. More extensive simulations under normal, exponential, and $t_3$ distributions confirm that finite-sample ARL, false alarm, and power are insensitive to the underlying $F_0$.

Empirical findings recommend:
- **Smaller threshold $c$ and larger window $r$** for enhanced detection of small shifts.
- **Intermediate parameters** for optimal detection of moderate to large shifts.
- **Pattern selection** in $(k_1, k_2)$-runs to balance $\mathrm{ARL}_0$ and $\mathrm{ARL}_1$ [1707.08367, 1801.06532].

---

**References:**  
- "Phase I Distribution-Free Control Charts for Individual Observations Using Runs and Patterns" [2511.13672]  
- "Distribution-free runs-based control charts" [1801.06532]  
- "Generalizations of Distributions Related to ($k_1,k_2$)-runs" [1707.08367]

Source: https://www.emergentmind.com/topics/runs-based-distribution-free-charts