- The paper introduces Add-MINAR, which separates row effects, column effects, cell-specific persistence, and Poisson innovations to make matrix-valued count dynamics more interpretable than bilinear MINAR models.
- The paper establishes stationarity, causality, consistency, and asymptotic normality for both projection estimation and iterative conditional least squares, with simulations showing that ICLSE generally achieves lower estimation error.
- The paper finds that Add-MINAR with ICLSE provides the strongest overall fit and forecasting performance on Chicago crime counts, using 36 parameters compared with 90 for MGINAR(1), while retaining interpretable spatial and crime-type effects.
Motivation and contribution
Matrix-valued integer-valued time series—cross-regional crime counts, multi-category sales, network traffic matrices—require models that respect both discreteness and the row–column structure of the observation matrix. The existing matrix integer-valued autoregressive (MINAR) model handles matrices directly via a matrix thinning operator, but its bilinear specification A⊛Xt−1⊛B⊤ couples row and column effects into a single multiplicative interaction, so parameters lack clear empirical meaning and one cannot separately attribute dynamics to rows, columns, or cell-specific lagged inertia. This paper proposes the additive matrix integer-valued autoregressive (Add-MINAR) model, which decomposes the conditional mean into three additive components: a row-effect matrix A, a column-effect matrix B, and an element-wise idiosyncratic lag matrix C, plus a Poisson intercept matrix D. The framework extends the continuous additive MAR model of Zhang et al. to count data.
The contributions are threefold: (i) probabilistic foundations—stationarity and causality conditions for the additive structure; (ii) two estimation procedures, projection estimation (PROJ) and iterative conditional least squares estimation (ICLSE), both shown consistent and asymptotically normal; and (iii) Monte Carlo and Chicago crime-data evidence that Add-MINAR, particularly under ICLSE, outperforms MGINAR(1), iINAR(1), and bilinear MINAR(1) benchmarks.
Model specification and identifiability
Building on Poisson thinning operators, the general Model (III) is
Xt=A⊛Xt−1+Xt−1⊛B⊤+C⊛⋅Xt−1+Et,
with element-wise innovations et,i,j∼Poisson(di,j). Two reduced variants impose only one of the zero-diagonal constraints. A key structural point is that Models (I) and (II) are nested boundary cases of Model (III): if C is column-replicated it can be absorbed into the diagonal of A, degenerating to Model (II), and symmetrically for row-replicated C.
Identifiability rests on zero-diagonal constraints, A0 and A1. Without them, the self-lagged term receives combined coefficient A2, so any mass shift A3 leaves the conditional expectation unchanged—an infinite family of observationally equivalent parameterizations. Delegating all self-lagged dynamics to A4 removes this ambiguity. Vectorization yields a standard MGINAR(1) form with global coefficient matrix
A5
whose main diagonal is governed exclusively by A6 under the constraints. The condition A7 guarantees a unique strictly stationary, ergodic solution with geometrically decaying innovation influence.
Estimation methodology
Projection estimation (PROJ) exploits the vectorized representation: a closed-form conditional least squares fit of the unconstrained MGINAR(1) is obtained, then structural parameters are recovered by projection—A8 averages off-diagonal blocks of A9, B0 averages block diagonals, and B1, B2 come from the diagonal of B3 and the intercept column.
ICLSE minimizes the constrained Frobenius objective directly, reformulating the zero-diagonal restrictions as trace equality constraints handled by Lagrange multipliers. Block coordinate descent alternates closed-form updates of B4, B5, then unconstrained updates of B6 and B7, initialized from PROJ estimates and terminated when squared Frobenius changes fall below B8 for all four matrices. Convergence is reported as rapid in practice.
Asymptotic theory
Under B9 and i.i.d. innovations with positive-definite covariance C0 (allowing contemporaneous cross-sectional dependence), the residual process C1 is shown to be matrix white noise with explicit covariance C2; under strict cross-sectional independence this simplifies to C3. A worked C4 example demonstrates that even when both C5 and C6 are singular, C7 and the stationary covariance C8 remain strictly positive definite—a property formalized in a proposition showing positive definiteness follows from C9 plus positive-definite D0, independent of the invertibility of the interactive coefficient matrices. This matters because non-singularity of D1 is exactly the regularity condition needed for asymptotic normality.
Both estimators are then shown asymptotically normal at rate D2: PROJ componentwise, with covariance matrices obtained by linear maps (D3) applied to D4; ICLSE jointly over all free parameters, with covariance expressed through the constraint-handling matrices D5 and the moment matrix D6. Diagonal entries of D7 and D8 have degenerate (zero) variance and are excluded from the limiting distributions.
Simulation evidence
Experiments use dimensions D9, sample sizes Xt=A⊛Xt−1+Xt−1⊛B⊤+C⊛⋅Xt−1+Et,0, 100 replications per configuration, and three innovation covariance settings (identity; random diagonal; Kronecker-structured dense). Three findings emerge. First, ICLSE consistently attains lower log-Frobenius errors than PROJ across all coefficient matrices and settings. Second, errors decrease monotonically with sample size, corroborating consistency, while increasing with dimension due to parameter proliferation. Third, convergence-curve experiments extending to very large samples show ICLSE dominating PROJ for nearly all sample sizes, with lower initial error, smaller converged error, and more stable trajectories; errors stabilize beyond roughly Xt=A⊛Xt−1+Xt−1⊛B⊤+C⊛⋅Xt−1+Et,1. The practical implication is that ICLSE should be preferred in moderate-to-large sample regimes, though the paper does not establish formal convergence guarantees for the iterative algorithm itself.
Empirical application to Chicago crime data
Monthly counts of three crime types (sexual offenses, vehicle theft, robbery) across three deliberately non-contiguous police districts (4, 5, 11) form a Xt=A⊛Xt−1+Xt−1⊛B⊤+C⊛⋅Xt−1+Et,2 matrix series of 271 observations (2001.01–2023.07); the first 251 months train, the last 20 test. The design lets Xt=A⊛Xt−1+Xt−1⊛B⊤+C⊛⋅Xt−1+Et,3 capture cross-crime-type contagion, Xt=A⊛Xt−1+Xt−1⊛B⊤+C⊛⋅Xt−1+Et,4 capture spatial spillover between adjacent districts (4–5) and non-adjacent associations (11 vs. 4/5), and Xt=A⊛Xt−1+Xt−1⊛B⊤+C⊛⋅Xt−1+Et,5 capture cell-specific persistence.
In-sample, Add-MINAR(1)_ICLSE achieves the best overall accuracy (Xt=A⊛Xt−1+Xt−1⊛B⊤+C⊛⋅Xt−1+Et,6, Xt=A⊛Xt−1+Xt−1⊛B⊤+C⊛⋅Xt−1+Et,7, Xt=A⊛Xt−1+Xt−1⊛B⊤+C⊛⋅Xt−1+Et,8) with 36 parameters, while the heavily parameterized MGINAR(1) (90 parameters) performs worst on every metric—direct evidence against vectorization-based modeling. Out-of-sample, Add-MINAR(1)_ICLSE again leads on Xt=A⊛Xt−1+Xt−1⊛B⊤+C⊛⋅Xt−1+Et,9 (1043.87) and et,i,j∼Poisson(di,j)0 (21.22), while Add-MINAR(1)_PROJ attains the best et,i,j∼Poisson(di,j)1 (1.6497) and et,i,j∼Poisson(di,j)2 (0.4742). Notably, MINAR(1)_ICLSE, competitive in-sample, degrades out-of-sample on et,i,j∼Poisson(di,j)3 and et,i,j∼Poisson(di,j)4, indicating weaker generalization than the additive variant. Estimated coefficients show ICLSE assigns systematically larger entries to et,i,j∼Poisson(di,j)5 and et,i,j∼Poisson(di,j)6 than PROJ, suggesting stronger explanatory weight on inter-variable relationships; the largest idiosyncratic persistences appear for vehicle theft and robbery (entries of et,i,j∼Poisson(di,j)7 around 0.38–0.60). One-step-ahead prediction plots reveal a concrete weakness: PROJ fails to track sexual offense trends in Districts 4 and 5, systematically underpredicting, whereas all models fit vehicle theft and robbery adequately.
Limitations and open questions
Several caveats bear directly on the results. The superiority claim for ICLSE rests on simulation and a single empirical dataset rather than formal optimization theory—the paper does not prove global convergence or uniqueness of the ICLSE fixed point, nor does it characterize finite-sample bias from the alternating scheme. Asymptotic normality requires positive-definite et,i,j∼Poisson(di,j)8 and non-singular et,i,j∼Poisson(di,j)9; behavior under overdispersed, zero-inflated, or heavy-tailed innovations common in crime data is not examined. The empirical analysis uses only first-order dependence and a small C0 matrix, leaving open whether the additive decomposition remains identifiable and efficient for higher-order lags, larger dimensions where parameter counts approach sample size, or higher-order Add-MINAR(C1) specifications. Finally, no inferential tools (standard errors, hypothesis tests on individual entries of C2 or C3) are applied to the crime-data estimates, despite the stated motivation of testing spatial-transmission hypotheses independently.
Conclusion
The Add-MINAR model replaces the coupled bilinear thinning structure of existing MINAR frameworks with an additive decomposition into row, column, and idiosyncratic lag effects, using zero-diagonal constraints to secure identifiability. With consistency and asymptotic normality established for both PROJ and ICLSE estimators, and with simulations and Chicago crime data showing ICLSE-based Add-MINAR achieving the best balance of fit, parsimony, and forecasting stability relative to MGINAR(1), iINAR(1), and bilinear MINAR(1), the framework offers a theoretically grounded and interpretable alternative for matrix-valued count time series with explicit row–column interaction structure.