---
title: Discrete Chi-square Method (DCM)
url: https://www.emergentmind.com/topics/discrete-chi-square-method-dcm
type: topic
---

# Discrete Chi-square Method (DCM)

Discrete Chi-square Method (DCM) denotes, in the literature represented here, a chi-square-centered methodology on explicitly discrete structures. In one explicit usage, DCM is a brute-force, grid-based time-series procedure that minimizes a Chi-square criterion over a multidimensional tested frequency space in order to recover multiple periodic signals superimposed on an unknown polynomial trend in unevenly spaced data [2002.03890]. In a broader statistical usage, the same label can reasonably denote chi-square-based procedures for empirical measures, multinomial counts, contingency tables, and histograms, including projection-based divergence estimation, exact finite-sample calibration of Pearson’s statistic, and variance-corrected or sparsity-aware testing for discrete data [1101.4353]. The term is therefore not a single universally standardized object; rather, it names a methodological family whose common feature is the use of chi-square structure on discrete supports.

## 1. Scope and terminological usage

The expression “Discrete Chi-square Method” is used explicitly in the time-series literature, where it denotes a numerical model-selection and signal-detection framework. In that usage, DCM is contrasted with the Discrete Fourier Transform and is presented as a method for detecting many signals superimposed on an unknown trend by performing a massive number of linear least-squares fits over a tested frequency grid [2509.01540].

At the same time, several directly relevant papers do not use the term itself, yet develop procedures that are closely aligned with what can be called a discrete chi-square method. The signed-measure divergence framework of Broniatowski and Keziou does not name its procedure DCM, but it produces discrete quadratic forms and minimum chi-square fits on empirical supports and contingency tables [1101.4353]. The exact-distribution paper on uniform histograms likewise does not define DCM as a formal name, yet it replaces continuous chi-square calibration by exact discrete calibration for Pearson’s statistic on multinomial histograms [2506.23416]. This suggests that “DCM” has both a narrow, named meaning and a broader descriptive meaning.

A recurring distinction is therefore necessary. In the narrow sense, DCM is a particular regression-and-grid-search algorithm for time series. In the broader sense, DCM refers to any procedure that treats chi-square statistics as inherently discrete objects, calibrates them on their attainable support, or reformulates them through duality, projection, or variance correction on discrete models.

## 2. Named DCM in time-series analysis

In its explicit, named formulation, DCM analyzes observations
$$
y_i = y(t_i) \pm \sigma_i,\qquad i=1,\dots,n,
$$
with model
$$
g(t)=g(t,K_1,K_2,K_3)=h(t)+p(t),
$$
where \(K_1\) is the number of periodic signals, \(K_2\) is the harmonic order of each signal, and \(K_3\) is the polynomial trend order. The periodic part is
$$
h(t)=\sum_{i=1}^{K_1} h_i(t,f_i), \qquad
h_i(t,f_i)=\sum_{j=1}^{K_2} B_{i,j}\cos(2\pi j f_i t)+C_{i,j}\sin(2\pi j f_i t),
$$
and the trend is a polynomial in normalized centered time [2002.03890].

The essential structural idea is a nonlinear-linear split. The frequencies
$$
\boldsymbol{\beta}_I=[f_1,\dots,f_{K_1}]
$$
are nonlinear parameters, while the harmonic amplitudes and trend coefficients are linear parameters. Once the tested frequencies are fixed to grid values, the model becomes linear and is solved by least squares. The residuals are
$$
\epsilon_i=y_i-g_i,
$$
with either residual sum of squares
$$
R=\sum_{i=1}^n \epsilon_i^2
$$
or
$$
\chi^2=\sum_{i=1}^n \frac{\epsilon_i^2}{\sigma_i^2},
$$
and the search statistic is
$$
z(f_1,\dots,f_{K_1})=\sqrt{R/n}
\quad\text{or}\quad
z(f_1,\dots,f_{K_1})=\sqrt{\chi^2/n},
$$
depending on whether \(\sigma_i\) are unknown or known [2002.03890].

The search is discrete and combinatorial. DCM tests only ordered frequency combinations satisfying
$$
f_{\max}\ge f_1>f_2>\cdots>f_{K_1}\ge f_{\min},
$$
which removes permutation duplicates because the statistic is symmetric in the tested frequency space. A long search on a coarse grid of \(n_L\) frequencies is followed by a short search on denser local grids of \(n_S\) frequencies around the best coarse optima. The total number of tested models scales as
$$
\binom{n_f}{K_1},
$$
or, including bootstrap reanalysis,
$$
\binom{n_L}{K_1} + (1+n_B)\binom{n_S}{K_1}.
$$
The method therefore relies on brute computational force, but every tested candidate is a linear least-squares problem [2509.01540].

Model selection is performed with nested-model Fisher statistics,
$$
F_R = \left( \frac{R_1}{R_2}-1 \right) \left( \frac{n-\eta_2-1}{\eta_2-\eta_1} \right),
\qquad
F_{\chi} = \left( \frac{\chi_1^2}{\chi_2^2}-1 \right) \left( \frac{n-\eta_2-1}{\eta_2-\eta_1} \right),
$$
where
$$
\eta = K_1(2K_2+1)+K_3+1.
$$
The decision threshold is \(Q_F<\gamma_F=0.001\). A separate Predictivity-test evaluates out-of-sample performance through the same \(z\)-statistic on predicted data. The method also uses residual bootstrap,
$$
\boldsymbol{y}^*=\boldsymbol{g}+\boldsymbol{\epsilon}^*,
$$
to estimate parameter uncertainties [2509.01540].

In this named sense, DCM is not Pearson’s goodness-of-fit test. It is a frequency-grid least-squares framework whose objective function is chi-square or residual sum of squares. Its “discrete” character lies in the tested frequency grid and the exhaustive search over discrete model combinations, not in multinomial cell counts.

## 3. Projection-based discrete chi-square methods for measures, counts, and contingency tables

In the broader statistical sense, a discrete chi-square method is exemplified by the signed-measure divergence framework of Broniatowski and Keziou. They define, for \(P\in M_1\) and signed \(Q\in M\) with total mass \(1\),
$$
\chi^{2}(Q,P)=
\begin{cases}
\displaystyle \int \left(\frac{dQ-dP}{dP}\right)^2\,dP, & Q \ll P,\\[1ex]
\infty, & \text{otherwise,}
\end{cases}
$$
and extend the analysis from probability measures to signed measures so that projections onto constrained sets become linear-quadratic optimization problems without positivity constraints [1101.4353].

The central obstacle in continuous models is that the empirical measure
$$
P_n=\frac1n\sum_{i=1}^n\delta_{X_i}
$$
renders the classical plug-in divergence unusable whenever \(Q\not\ll P_n\). The paper resolves this through a dual representation. For fixed \(Q\),
$$
\chi_n^2(Q,P):=\sup_{f\in\mathcal F}\int m_f(x)\,dP_n(x),
$$
and for a model \(\Omega\),
$$
\chi_n^2(\Omega,P):=\inf_{Q\in\Omega}\sup_{f\in\mathcal F}\int m_f(x)\,dP_n(x).
$$
This operational definition is computable from empirical data without grouping or smoothing and is the mechanism by which the method remains feasible on discrete empirical supports [1101.4353].

The discrete connection becomes explicit when \(\Omega\) is defined by linear constraints. In the finite-constraint case, the empirical optimization is posed directly on
$$
\Omega\cap\Lambda_n,
$$
where \(\Lambda_n\) is the set of signed measures supported on the sample points. The resulting statistic reduces to a quadratic form,
$$
\chi_n^2=\underline{\nu}_n'S_n^{-1}\underline{\nu}_n,
$$
and under \(H_0\),
$$
n\chi_n^2\Rightarrow \chi^2(k).
$$
For contingency tables and related multinomial problems, the same framework yields minimum chi-square criteria over fitted cell probabilities subject to constraints, such as
$$
n\chi_{n,k}^{2} = \min_{Q\in\mathcal Q} \sum_{i=1}^{m+1}\sum_{j=1}^{m+1}
\frac{(nq_{i,j}-N_{i,j})^2}{N_{i,j}\,\mathbf 1_{\{N_{i,j}>0\}}}.
$$
In this setting, the procedure is a direct generalization of Pearson-type discrete chi-square fitting, but derived from divergence duality rather than from a raw plug-in formula [1101.4353].

This broader DCM interpretation has two consequences. First, it places classical cell-count chi-square procedures inside a more general convex-analytic framework. Second, it clarifies that discrete quadratic forms may arise either from genuine categorical data or from projection of a more general measure-theoretic problem onto empirical support.

## 4. Exact calibration, approximation theory, and the discrete null law

A central theme in discrete chi-square methodology is that the null distribution of Pearson’s statistic is itself discrete at finite sample size. For a histogram \(\mathbf{x}_n=(x_1,\dots,x_n)\) with \(\sum_i x_i=N\) under a discrete uniform null, Pearson’s statistic can be written as
$$
\chi_{n-1}^2=\frac{n}{N}\sum_{i=1}^n x_i^2 - N.
$$
Defining
$$
s=\sum_{i=1}^{n}x_i^2,
$$
the exact distribution of \(\chi^2\) is equivalent to the exact distribution of the integer-valued statistic \(s\). Zero-disparity Distribution Synthesis computes this law by dynamic programming through the recurrence
$$
C_{N,n}(i,M,s)=\sum_{m=0}^{M}\binom{N-M+m}{m}\,
C_{N,n}(i-1,M-m,s-m^2),
$$
which yields exact masses
$$
p_{N,n}^{(s)}=\frac{C_{N,n}(n,N,s)}{n^N}.
$$
The exact upper-tail \(p\)-value is then a discrete tail probability over attainable support points, not a continuous \(\chi^2\) cdf evaluation [2506.23416].

This exact-calibration viewpoint materially changes interpretation in sparse or low-count settings. The paper shows that the usual rule of thumb \(N/n\ge 5\) is too crude, that tail errors can remain substantial even when the bulk looks acceptable, and that treating a discrete null as continuous can distort rejection rates. Its NIST example with \(n=10\), significance threshold \(10^{-4}\), and \(N=55\) gives exact \(p\)-values that are substantially larger than the asymptotic chi-square approximation, so the continuous approximation would reject when the exact discrete test would not [2506.23416].

Approximation theory supplies complementary, rather than competing, results. Stein-based analysis yields explicit bounds for the distance between Pearson’s multinomial statistic
$$
W=\sum_{j=1}^m \frac{(U_j-np_j)^2}{np_j}
$$
and its limiting \(\chi^2_{m-1}\) law, with smooth-test-function error of order \(n^{-1}\) under \(np_j\ge 1\) and explicit dependence on \(m\) and \(p_*=\min_j p_j\) [1507.01707]. A different non-asymptotic route compares the multinomial Pearson statistic directly with its Gaussian quadratic-form analogue and proves
$$
\sup_{\mathbf{p}\in \mathscr{P}_{\tau}}
\big|P(X^2\le \ell)-P(\chi_d^2\le \ell)\big|
\le
\frac{1.26 \,\tau^3 (d + 1) (\log n)^{3/2}}{n^{1/2}},
$$
uniformly over probability vectors bounded away from the simplex boundary [2309.01882].

When calibration leads to generalized rather than ordinary chi-square distributions, further numerical machinery becomes relevant. Exact and approximate methods for the cdf, pdf, and inverse cdf of generalized chi-square laws include ray-tracing, inverse Fourier transform, and finite-tail ellipse approximations, with distinct tradeoffs in speed and tail accuracy [2404.05062]. A plausible implication is that advanced DCM variants can be limited as much by calibration numerics as by statistic construction.

## 5. Robustness, pathologies, and methodological disputes

A persistent misconception is that exact \(p\)-values automatically repair chi-square inference. They do not. In discrete goodness-of-fit problems with highly nonuniform null probabilities, Pearson’s weighting by \(1/p_k\) can cause bins with tiny model probabilities to dominate the statistic. The result can be severe loss of power even under exact computational calibration. In response, one line of work advocates separating discrepancy choice from significance computation and recommends statistics such as the unweighted root-mean-square
$$
X=\sqrt{\frac{1}{n}\sum_{k=1}^n (q_k-p_k(\hat\theta))^2}
$$
together with Monte Carlo calibration under the fitted null, precisely because exactness of significance does not rescue a poor discrepancy measure [1108.4126].

A second pathology appears in histogram fitting when overall normalization is itself a fitted parameter. For binned count data with variance tied to the expected counts, naive chi-square minimization can be biased because the denominator depends on the parameters. In the two-bin toy model,
$$
\chi^2=\frac{(N-x_1)^2}{N}+\frac{(N-x_2)^2}{N},
$$
and maximum likelihood introduces an additional normalization term absent from naive chi-square differentiation. The resulting score equations coincide with Poisson maximum likelihood, and the recommended chi-square-like fix is to omit derivatives of the inverse error matrix when solving the stationarity equations [1506.09077].

A third dispute concerns over-dispersion and sparse, high-dimensional discreteness. In temporal allele-frequency studies, the classical chi-square and CMH null models underestimate variance because they ignore Wright–Fisher drift, finite population sampling, and pool sequencing. The adapted chi-square statistic
$$
T_{\chi^2}^a (s_1^2, s_2^2)
=
\frac{(x_{11} x_{22} - x_{12} x_{21})^2}{x_{2+}^2 s_1^2 + x_{1+}^2 s_2^2}
$$
replaces the standard denominator by a variance estimator matched to the actual hierarchical null, and the corresponding adjusted CMH statistic does the same across strata [1902.08127]. In sparse multinomial goodness-of-fit, Pearson’s statistic itself can become biased because a signed linear component can pull the test statistic leftward under alternatives. Decomposing
$$
\mathcal X_n^2=S_{n1}+S_{n2}
$$
and replacing \(S_{n2}\) by \(|S_{n2}|\) or a weighted version yields more powerful tests with asymptotic null law based on \(Z_1+s|Z_2|\), not a standard chi-square limit [2112.03231].

These results collectively show that DCM is not synonymous with uncritical Pearson testing. The phrase marks a family of methods whose practical validity depends on which of four objects is made discrete and model-aware: the statistic, the null distribution, the optimization problem, or the variance model.

## 6. Applications, extensions, and related domains

The application range of discrete chi-square methodology is broad. In astronomy, the named time-series DCM has been used on eclipse timing data for XZ And, where a model \(g(t,2,1,2)\) yielded two periodicities,
$$
P_1 = 13418^{\mathrm d} \approx 37^{\mathrm y},
\qquad
P_2 = 32192^{\mathrm d} \approx 88^{\mathrm y},
$$
and was interpreted as evidence for the possible presence of a third and a fourth body [2002.03890]. In a different astronomical setting, Pearson-style chi-square is used as a pointwise similarity metric on normalized asteroid and laboratory reflectance spectra after normalization, Savitzky–Golay smoothing, and interpolation to a common wavelength grid. There the statistic functions primarily as a ranking score rather than a formal \(p\)-value, and lower \(\chi^2\) indicates closer spectral resemblance [2411.18705].

In discrete distribution testing beyond classical multinomials, the convolution statistic provides a nonparametric maximum-likelihood estimator for the distribution of a sum of independent discrete random variables from unequal sample sizes. The resulting generalized Wald statistics have asymptotic chi-square laws with degrees of freedom given by covariance rank rather than raw cell count, and simulation evidence indicates higher power than Pearson’s chi-square when classical matched-tuple constructions would discard data [2008.13657].

Chi-square methodology also extends beyond second-order discrepancy. For members of the same affine exponential family, closed-form formulas exist for Pearson and Neyman chi-square distances and for higher-order Pearson–Vajda \(\chi^k\) quantities. These enter Taylor expansions for general \(f\)-divergences,
$$
I_f(X_1:X_2)=\sum_{k=0}^{\infty}\frac{f^{(k)}(1)}{k!}\chi_P^k(X_1:X_2),
$$
and specialize, for example, to explicit Poisson-family formulas [1309.3029]. This suggests that a “chi-square method” need not be confined to classical second-order Pearson residuals; it can also serve as an analytic basis for approximating broader divergence functionals.

The common misconception that DCM is a single recipe should therefore be rejected. In the literature represented here, the term encompasses at least three distinct but connected ideas: a named grid-based time-series search procedure; projection- and duality-based minimum chi-square methods on empirical supports; and finite-sample discrete calibration or correction of Pearson-type statistics for histograms, contingency tables, sparse multinomials, and over-dispersed count data. What unifies these variants is not a single algorithm, but a shared insistence that chi-square methodology must respect the discrete structure of the problem it is applied to.

Source: https://www.emergentmind.com/topics/discrete-chi-square-method-dcm