---
title: 'Ben-Util: Beyond Benford Analysis'
url: https://www.emergentmind.com/topics/ben-util
type: topic
---

# Ben-Util: Beyond Benford Analysis

Ben-Util (BeyondBenford) is an R package designed to compare the empirical digit distributions of real-world datasets with two theoretical models: Newcomb–Benford’s Law and the Blondeau Da Silva (BDS) model. Its primary utility is for researchers and practitioners evaluating whether observed digit frequencies—particularly leading or p-th digits—are better explained by the widely cited logarithmic Benford distribution or the parameterized, data-bound BDS model [1910.06104].

## 1. Decision Problem: Modeling Digit Distributions

Empirical investigations consistently demonstrate that leading digits in many natural datasets are not uniformly distributed, instead often exhibiting a logarithmic decay described by Newcomb–Benford’s Law. The BDS model extends this, formalizing how an upper-bound constraint on the data induces fluctuations around Benford frequencies. Ben-Util operationalizes this comparative inquiry by offering a workflow to assess, for any dataset of positive real numbers, whether the Benford or BDS distribution more appropriately models its observed digits, using goodness-of-fit metrics [1910.06104].

## 2. Explicit Theoretical Formulations

Let $d$ be the digit of interest (first digit $d \in \{1, \ldots, 9\}$; in general, the $p$th digit $d \in \{0,\ldots,9\}$).

**Newcomb–Benford’s Law**:

$$
P_B(d) = \log_{10}(1 + 1/d)
$$

**Blondeau Da Silva’s (BDS) Model**:  
If the integer upper bound for the data is $u_b$ and focusing on the $p$th digit,
$$
P_{BDS}(d; p, u_b) = \frac{1}{u_b - 10^{p-1} + 1} \sum_{i = d \cdot 10^{p-1}}^{u_b} \frac{1}{i}
$$
For first digit $(p=1)$:
$$
P_{BDS}(d; u_b) = \frac{1}{u_b} \sum_{i=d}^{u_b} \frac{1}{i}
$$

These formulas codify the two candidate distributions underlying Ben-Util comparisons. The Benford model is parameter-free; the BDS model depends explicitly on $u_b$ matched to the dataset’s observed maximum [1910.06104].

## 3. Core Functionality and Workflow

Ben-Util (BeyondBenford) provides a suite of R functions facilitating full digit-distribution analysis:

- **obs.numb.dig(dat, dig)**: Extracts observed frequencies for each digit (positions 0–9) at digit position `dig` from a vector or data-frame `dat`.
- **Benf.val(fig, dig)**: Computes Benford’s theoretical probability for digit `fig` at position `dig`.
- **Blon.val(fig, dig, upbound)**: Computes BDS’s probability for digit `fig` at position `dig`, with parameter `upbound`.
- **dat.distr(dat, dig=1, upbound=ceiling(max(dat)), ...)**: Plots histograms of observed digit frequencies, optionally overlaying theoretical curves and providing $\chi^2$ summaries.
  - Key arguments:  
    - `nclass`: Number of bins,
    - `theor`: Whether to plot the BDS curve,
    - `nchi`: If $>0$, merges bins for robust $\chi^2$.
- **digit.distr(dat, dig, upbound, mod="ben"|"BDS"|"benblo", ...)**: Plots comparisons between the observed, Benford, and/or BDS digit distributions.
- **chi2(dat, dig, upbound=…, mod="ben"|"BDS", pval=0|1)**: Computes Pearson’s $\chi^2$ and optionally $p$-value for goodness-of-fit to the chosen model [1910.06104].

This workflow enables detailed side-by-side quantitative and graphical comparisons between candidate distributions.

## 4. Observed Versus Expected Frequencies

Frequencies are computed as follows:

- **Observed counts**:
  $$
  O_0, \ldots, O_9 = \text{obs.numb.dig(dat, dig)}
  $$
  For first digit analyses, focus on $O_1, \ldots, O_9$.

- **Expected counts**:
  - **Benford**:
    $$
    E_i = P_B(i) \times N
    $$
  - **BDS**:
    $$
    E_i = P_{BDS}(i; dig, u_b) \times N
    $$
  with $N = \sum O_i$ indicating records with at least `dig` digits [1910.06104].

These serve as empirical-theoretical pairs input to the $\chi^2$ test.

## 5. Statistical Testing: Pearson’s $\chi^2$ Test

For bins $i=1,\ldots,k$ (typically digits 1–9 or 0–9):

$$
X^2 = \sum_{i=1}^k \frac{(O_i - E_i)^2}{E_i}
$$

Under the null hypothesis ($H_0$), $X^2 \sim \chi^2_{df}$ with $df = k-1$ if no parameters are estimated from the data. The $p$-value is

$$
p = 1 - F_{\chi^2_{df}}(X^2)
$$

In Ben-Util, bins are merged if necessary to ensure all expected frequencies $E_i \geq 5$, following the standard Pearson’s rule. This procedure provides a statistical measure of the adequacy of each model’s fit [1910.06104].

## 6. Example Analysis Pipeline

The illustrative R workflow using Ben-Util includes:

```r
library(BeyondBenford)
data(address_PierreBuffiere)        # Sample: 346 street addresses

# Analyze second-digit frequencies
obs2 = obs.numb.dig(address_PierreBuffiere, dig=2)
N2 = sum(obs2)                      # Only those with ≥2 digits

# Compute theoretical probabilities for 0…9
benf2 = sapply(0:9, function(d) Benf.val(d, dig=2))
blon2 = sapply(0:9, function(d) Blon.val(fig=d, dig=2, upbound=74))

# Convert to expected counts
E_benf2 = benf2 * N2
E_blon2 = blon2 * N2

# Print normalized table
tab = data.frame(Digit=0:9, Observed=obs2/N2, Benford=round(benf2,4), BDS=round(blon2,4))
print(tab)

# Visual comparison: histogram + both models
digit.distr(address_PierreBuffiere, dig=2, mod="benblo", upbound=74, main="2nd‐digit: Observed vs. Benford vs. BDS")

# Compute and compare χ² statistics
res_benf = chi2(address_PierreBuffiere, dig=2, upbound=74, mod="ben", pval=1)
res_bds  = chi2(address_PierreBuffiere, dig=2, upbound=74, mod="BDS", pval=1)
res_benf;  res_bds
```

Reported outputs for this example indicate strong fits for both models:  
- Benford $\chi^2=3.85$, $p=0.92$  
- BDS $\chi^2=2.48$, $p=0.98$  
The lower $\chi^2$ and higher $p$ indicate BDS provides a marginally better fit [1910.06104].

## 7. Model Selection and Interpretation

Assessment proceeds by juxtaposing $\chi^2$ statistics and $p$-values. The preferred model is the one yielding the smaller $\chi^2$ and larger $p$-value:
- If both $p \gg 0.05$, either model may be considered plausible.
- If only one model yields $p>0.05$, preference should be given to that model.
- For closely matching $p$-values, differences in the $\chi^2$ statistic may provide additional guidance [1910.06104].

A plausible implication is that, in datasets subject to explicit upper bounds or selection effects, BDS should be favored over Benford when it offers a superior empirical fit. However, with both models highly plausible, interpretative caution and context-specific considerations remain essential.

---

This framework enables rigorous, reproducible comparison of empirical digit distributions in diverse disciplines, using the operational tools provided by the Ben-Util (BeyondBenford) package [1910.06104].

Source: https://www.emergentmind.com/topics/ben-util