Arithmetic Width in Computation
- Arithmetic width is a domain-specific parameter that measures the size of intermediate computations in read-once oblivious arithmetic branching programs and constant-width circuits.
- In computer arithmetic, it controls bit-width allocation, influencing dynamic range, precision, energy consumption, and overall computational performance.
- In discrete geometry, arithmetic width captures the number of distinct integer levels attained by linear functionals on lattice points, aiding in quasipolynomial analyses and structural insights.
Searching arXiv for recent and foundational uses of “arithmetic width” across algebraic complexity, computer arithmetic, and discrete geometry. {"3query3 width\"3 OR title:\3"Arithmetic Width\"","max_results":3all:\3query3,"sort_by":"relevance"} {"3query3 arithmetic width ROABP (&&&3all:\3&&&) word-length (&&&3 OR title:\3&&&) constant width arithmetic circuits (Loera et al., 5 Sep 2025) convex bodies","max_results":5,"sort_by":"submittedDate"} Arithmetic width is a domain-dependent notion of bounded intermediate complexity. In algebraic complexity, it is the width parameter PRESERVED_PLACEHOLDER_3query3^ of a read-once oblivious arithmetic branching program (ROABP) or of a constant-width arithmetic circuit, measuring the size of the largest intermediate matrix or layer. In computer arithmetic, it is the total bit-width or word-length that controls dynamic range and precision. In discrete geometry, it is the number of distinct values an integer linear functional attains on the lattice points of a convex body. Across these settings, width serves as a structural restriction that simultaneously limits expressivity and enables stronger algorithmic results, especially for polynomial identity testing, customized-precision arithmetic, and lattice-slice analysis (&&&3query3&&&, &&&3all:\3&&&, Loera et al., 5 Sep 2025).
3all:\3. Arithmetic width in algebraic computation
An Arithmetic Branching Program (ABP) is a layered directed acyclic graph computing a polynomial by summing over source–sink path-weights. In matrix form, an ABP of length PRESERVED_PLACEHOLDER_3all:\3^ and width PRESERVED_PLACEHOLDER_3 OR title:\3^ over a field is a sequence
computing
Its width is the size of the largest intermediate matrix, and the total size is the cost of representing all the univariate entries. When each layer is a univariate polynomial in a different variable, the ABP is oblivious; if each variable appears exactly once, it is a read-once oblivious ABP. In this setting, the width of an ROABP is often called the arithmetic width of the polynomial it computes (&&&3query3&&&).
This width parameter has an operational interpretation. It measures the “space” needed to build up a polynomial in a one-pass read of the variables, or, in the non-commutative viewpoint, the matrix-multiplication dimension required by the computation. Small width imposes strong structural constraints. The ROABP literature uses those constraints to obtain stronger black-box PIT algorithms than were known for general arithmetic circuits (&&&3query3&&&).
A related notion appears in layered arithmetic circuits. If non-leaf nodes are partitioned into layers , then the width is
A layered circuit is staggered if in every layer PRESERVED_PLACEHOLDER_3all:\3query3^ at most one gate is not of the form “multiply by PRESERVED_PLACEHOLDER_3all:\3all:\3” from the previous layer; equivalently, staggered width-PRESERVED_PLACEHOLDER_3all:\3 OR title:\3^ circuits are exactly straight-line programs with PRESERVED_PLACEHOLDER_3all:\33^ registers. Here too, width bounds how many partial values can be kept alive at once, restricting how much parallelism or mixing of variables the circuit can perform (&&&3 OR title:\3&&&).
3 OR title:\3. Hitting sets and derandomization for bounded-width ROABPs
The ROABP identity-testing results of Gurjar, Korwar, and Saxena sharpen the algorithmic meaning of arithmetic width. For the known-order model,
PRESERVED_PLACEHOLDER_3all:\34
if PRESERVED_PLACEHOLDER_3all:\35 or PRESERVED_PLACEHOLDER_3all:\36, one can deterministically construct in time
PRESERVED_PLACEHOLDER_3all:\37
a hitting set of size
PRESERVED_PLACEHOLDER_3all:\38
In particular, for constant PRESERVED_PLACEHOLDER_3all:\39 this is a truly polynomial-size hitting set. For the any-order model,
PRESERVED_PLACEHOLDER_3 OR title:\3query3^
one can build over any field a hitting set of size
PRESERVED_PLACEHOLDER_3 OR title:\3all:\3^
These improve the previously known bounds and mark the first polynomial-size hitting set for constant-width ROABP in the known-order case (&&&3query3&&&).
The structural core of the known-order argument is the rank of the partial-derivative matrix. For a bivariate polynomial
PRESERVED_PLACEHOLDER_3 OR title:\3 OR title:\3^
the matrix
PRESERVED_PLACEHOLDER_3 OR title:\33^
satisfies a rank-width relation: if
PRESERVED_PLACEHOLDER_3 OR title:\34
then PRESERVED_PLACEHOLDER_3 OR title:\35. This is used to prove that the substitution
PRESERVED_PLACEHOLDER_3 OR title:\36
preserves nonzeroness for width-PRESERVED_PLACEHOLDER_3 OR title:\37 ROABP polynomials. The proof then iterates a variable-halving lemma: if PRESERVED_PLACEHOLDER_3 OR title:\38 has a width-PRESERVED_PLACEHOLDER_3 OR title:\39 ROABP in the order 3query3, then substituting
3all:\3^
for 3 OR title:\3^ yields a nonzero polynomial in 3 still computable by a width-4 ROABP in the order 5. After 6 rounds one gets a univariate test of degree 7 (&&&3query3&&&).
The any-order argument proceeds differently. First, a low-degree bivariate shift forces all small ROABPs of width 8 in any order to become 9-concentrated, where every coefficient of a large-support monomial lies in the span of coefficients of support 3query3, for 3all:\3. The proof uses a basis-isolating weight assignment and the fact that shifting by a basis-isolating map yields 3 OR title:\3-concentration in any algebra 3. Second, once 4 is known to have an 5-support monomial with nonzero coefficient, black-box testing reduces to an 6-variate ROABP in any order of width 7 and individual degree 8, recovering the 9 bound (&&&3query3&&&).
3. Constant-width arithmetic circuits and lower bounds
Width is not only an algorithmic aid; it is also a source of lower bounds. For every 3query3, Arvind, Joglekar, and Srinivasan construct an explicit polynomial that can be computed by a linear-sized monotone circuit of width 3all:\3^ but has no subexponential-sized monotone circuit of width 3 OR title:\3. Their family 3 is defined recursively, is homogeneous of degree 4 on 5 variables, and has exactly
6
distinct monomials. The separation theorem states that for fixed 7, for all sufficiently large 8, any monotone width-9 circuit 3query3^ with
3all:\3^
must have size
3 OR title:\3^
On the other hand, 3 admits a width-4 monotone circuit of size 5 (&&&3 OR title:\3&&&).
This yields infinite hierarchies. For every fixed 6, the class of monotone width-7 circuits is strictly contained in monotone width-8 circuits for infinitely many input sizes, so the monotone constant-width hierarchy is infinite. Because width-9 staggered circuits can be unfolded into depth-3query3^ circuits of polynomial blow-up, the same strictness holds for the monotone constant-depth hierarchy. The argument also carries over to the noncommutative setting (&&&3 OR title:\3&&&).
The same paper develops hardness–randomness tradeoffs for identity testing constant-width commutative circuits. Two technical lemmata are central: if 3all:\3^ of total degree 3 OR title:\3^ has a staggered width-3, size-4 circuit, then each homogeneous component 5 has a staggered width-6, size 7 circuit; and if 8 has a staggered width-9, size-3query3^ circuit and 3all:\3, then under 3 OR title:\3, any annihilated series 3 with 4 admits a staggered circuit of width 5 and size 6. These ingredients support a Nisan–Wigderson-style generator and a deterministic subexponential-time PIT for width-7 circuits under suitable lower-bound assumptions (&&&3 OR title:\3&&&).
A common misconception is to identify bounded width with a mere size bound. The circuit results indicate that width is a finer-grained restriction: linear-size width-8 computations can be inaccessible to subexponential-size width-9 computations, even in monotone models. This suggests that arithmetic width controls a qualitative resource, not only a quantitative one (&&&3 OR title:\3&&&).
4. Arithmetic width as word-length in computer arithmetic
In digital number representation, arithmetic width is the total bit-width allocated to a datum. The total bit-width 3query3, sometimes called word-length, is the single most critical parameter because it must be apportioned between range and precision. In signed fixed-point 3all:\3, the total width is
3 OR title:\3^
where 3 is the integer-part word-length, including the sign bit, and 4 is the fractional-part word-length. If 5 is the 6-bit two’s-complement integer, then
7
with representable range
8
and quantization step
9
In radix-3 OR title:\3^ floating-point, the total width is partitioned as
PRESERVED_PLACEHOLDER_3all:\3query3query3^
and a real value is represented as
PRESERVED_PLACEHOLDER_3all:\3query3all:\3^
Two immediate consequences are the dynamic range
PRESERVED_PLACEHOLDER_3all:\3query3 OR title:\3^
and the unit round-off
PRESERVED_PLACEHOLDER_3all:\3query33^
These formulas make explicit that width governs both dynamic range and resolution (&&&3all:\3&&&).
At the operator level, the reported 3 OR title:\38 nm FDSOI, 3 OR title:\3query3query3^ MHz synthesis results show that the effect of width depends strongly on representation.
| Operator | Area, delay | Energy |
|---|---|---|
| 33 OR title:\3-bit float add | 653 PRESERVED_PLACEHOLDER_3all:\3query34, 3 OR title:\3.43 OR title:\3^ ns | PRESERVED_PLACEHOLDER_3all:\3query35 fJ |
| 33 OR title:\3-bit int add | 3all:\389 PRESERVED_PLACEHOLDER_3all:\3query36, 3all:\3.3query3 ns | PRESERVED_PLACEHOLDER_3all:\3query37 fJ |
| 33 OR title:\3-bit float mul | 3all:\3543 PRESERVED_PLACEHOLDER_3all:\3query38, 3 OR title:\3.3query39 ns | PRESERVED_PLACEHOLDER_3all:\3query39 fJ |
| 33 OR title:\3-bit int mul | 3 OR title:\3 OR title:\389 PRESERVED_PLACEHOLDER_3all:\3all:\3query3, 3 OR title:\3.38 ns | PRESERVED_PLACEHOLDER_3all:\3all:\3all:\3^ fJ |
The data imply distinct tradeoffs. The 33 OR title:\3-bit float adder is approximately PRESERVED_PLACEHOLDER_3all:\3all:\3 OR title:\3^ larger, approximately PRESERVED_PLACEHOLDER_3all:\3all:\33^ slower, and approximately PRESERVED_PLACEHOLDER_3all:\3all:\34 more energy than the 33 OR title:\3-bit integer adder. By contrast, the 33 OR title:\3-bit float multiplier is area-smaller but approximately PRESERVED_PLACEHOLDER_3all:\3all:\35 more energy because of exponent and shift management. In the width sweep from 8 to 3all:\36 bits, floating-point adders are always approximately PRESERVED_PLACEHOLDER_3all:\3all:\36 larger, approximately PRESERVED_PLACEHOLDER_3all:\3all:\37 slower, and PRESERVED_PLACEHOLDER_3all:\3all:\38–PRESERVED_PLACEHOLDER_3all:\3all:\3 more energy, while floating-point multipliers are smaller in area, slightly faster, but still PRESERVED_PLACEHOLDER_3all:\3 OR title:\3query3–PRESERVED_PLACEHOLDER_3all:\3 OR title:\3all:\3^ more energy (&&&3all:\3&&&).
Application-level studies refine that picture. For 8-bit 3 OR title:\3-D K-Means, the floating-point format PRESERVED_PLACEHOLDER_3all:\3 OR title:\3 OR title:\3^ is approximately PRESERVED_PLACEHOLDER_3all:\3 OR title:\33^ bigger and its distance kernel consumes approximately PRESERVED_PLACEHOLDER_3all:\3 OR title:\34 more energy than PRESERVED_PLACEHOLDER_3all:\3 OR title:\35, yet it converges in half as many iterations and yields PRESERVED_PLACEHOLDER_3all:\3 OR title:\36 lower CMSE and roughly half the misclassification rate. For 3all:\36-bit K-Means, the converse holds: fixed-point PRESERVED_PLACEHOLDER_3all:\3 OR title:\37 is both smaller, uses PRESERVED_PLACEHOLDER_3all:\3 OR title:\38 less energy overall, and reaches equal or better CMSE and ER than floatPRESERVED_PLACEHOLDER_3all:\3 OR title:\39. For FFT-3all:\36, fixed-point consistently beats floating-point for any equal total width, both in energy per operation and in MSE, because the dynamic range is modest and a 7-bit exponent in floating-point wastes bits that could have boosted the mantissa (&&&3all:\3&&&).
These results support an application-contingent view of arithmetic width. Fixed-point is almost always smaller, faster, and lower-energy per operation than floating-point when uniform quantization and explicit scaling are acceptable. Floating-point becomes advantageous when automatic scaling over a wide dynamic range materially improves convergence or prevents saturation and overflow (&&&3all:\3&&&).
5. Customized precision and non-native-width arithmetic
When the desired arithmetic width is not supported natively, recent work treats width itself as a software- or compiler-controlled design variable. The SAMD method, “Scalar Arithmetic Multiple Data,” embeds a vector architecture with custom bit-width lanes inside fixed-width scalar arithmetic. For native register width PRESERVED_PLACEHOLDER_3all:\33query3, virtual lane width PRESERVED_PLACEHOLDER_3all:\33all:\3, and PRESERVED_PLACEHOLDER_3all:\33 OR title:\3^ spacer bits per lane, one uses
PRESERVED_PLACEHOLDER_3all:\333^
lanes. Spacer bits prevent carries or overflows from bleeding between lanes, and ordinary scalar arithmetic plus bitwise operations implements lane-wise add, sub, mul, and even 3all:\3-D convolution. The paper reports a maximum speedup of approximately PRESERVED_PLACEHOLDER_3all:\334 on an Intel platform and approximately PRESERVED_PLACEHOLDER_3all:\335 on an ARM platform versus quantization to native 8-bit integers; on Cortex-A57, the temporary-spacer configuration reaches approximately PRESERVED_PLACEHOLDER_3all:\336 at 6-bit, PRESERVED_PLACEHOLDER_3all:\337 at 4-bit, PRESERVED_PLACEHOLDER_3all:\338 at 3-bit, and PRESERVED_PLACEHOLDER_3all:\339 at 3 OR title:\3-bit, while permanent spacer reaches up to PRESERVED_PLACEHOLDER_3all:\3max_results3query3^ at 3 OR title:\3-bit (&&&3 OR title:\3all:\3&&&).
The essential point is that width below 8 bits can be exploited without specialized hardware. One scalar multiply can simultaneously compute all PRESERVED_PLACEHOLDER_3all:\3max_results3all:\3^ lanes of partial products, so for small convolution kernels the overall multiply-plus-add count can drop substantially. The method therefore targets a regime in which custom widths are not merely storage formats but active algorithmic accelerants (&&&3 OR title:\3all:\3&&&).
A different non-native-width strategy appears in multi-word modular arithmetic for cryptographic kernels. In MoMA, if PRESERVED_PLACEHOLDER_3all:\3max_results3 OR title:\3^ is the machine-word size and PRESERVED_PLACEHOLDER_3all:\343 the required integer bit-width, each PRESERVED_PLACEHOLDER_3all:\344-bit integer is represented as
PRESERVED_PLACEHOLDER_3all:\345
A rewrite system recursively splits every PRESERVED_PLACEHOLDER_3all:\346-bit type into two PRESERVED_PLACEHOLDER_3all:\347-bit halves, replaces high-level add, sub, mul, and mod by multi-word sequences, prunes limbs known to be zero for non-power-of-two widths, and continues until all operations fit in the native word size. The implementation integrates this pass into SPIRAL and combines it with Barrett reduction, fused add-with-carry when PRESERVED_PLACEHOLDER_3all:\348, and rule selection between schoolbook and Karatsuba multiplication (&&&3 OR title:\33&&&).
The reported GPU results make arithmetic width an explicit performance knob. For BLAS Level-3all:\3^ operations on 3all:\3 OR title:\38–3all:\3query3 OR title:\34-bit widths, MoMA outperforms GRNS by PRESERVED_PLACEHOLDER_3all:\349–PRESERVED_PLACEHOLDER_3all:\3sort_by3query3^ and GMP by PRESERVED_PLACEHOLDER_3all:\3sort_by3all:\3–PRESERVED_PLACEHOLDER_3all:\3sort_by3 OR title:\3^ for add/sub, and by PRESERVED_PLACEHOLDER_3all:\353–PRESERVED_PLACEHOLDER_3all:\3 for mul/axpy. For NTT, 3 OR title:\356-bit performance on RTX 43query3max_results3query3^ is PRESERVED_PLACEHOLDER_3all:\355 faster than ICICLE on H3all:\3query3query3^ and within PRESERVED_PLACEHOLDER_3all:\356 of an ASIC; for 3all:\3 OR title:\38-bit NTT, MoMA on RTX 43query3max_results3query3^ outperforms an FHE ASIC by PRESERVED_PLACEHOLDER_3all:\357, and on H3all:\3query3query3^ by PRESERVED_PLACEHOLDER_3all:\358. The paper also notes that non-power-of-two widths such as 383all:\3^ and 753 bits avoid zero-padding and gain PRESERVED_PLACEHOLDER_3all:\359–PRESERVED_PLACEHOLDER_3all:\3relevance3query3^ speedups over naively padded code, while dynamically varying precision is not supported because kernels must fix PRESERVED_PLACEHOLDER_3all:\3relevance3all:\3^ at compile time (&&&3 OR title:\33&&&).
Taken together, SAMD and MoMA show that arithmetic width can be virtualized. In one case, width is packed into scalar registers; in the other, it is decomposed into limbs and lowered by compiler rewrites. Both approaches treat width as a first-class compilation parameter rather than a fixed ISA feature (&&&3 OR title:\3all:\3&&&, &&&3 OR title:\33&&&).
6. Arithmetic width of convex bodies
In discrete geometry, arithmetic width has a different and more literal meaning. For a convex body PRESERVED_PLACEHOLDER_3all:\3relevance3 OR title:\3^ and a nonzero integer vector PRESERVED_PLACEHOLDER_3all:\363, let
PRESERVED_PLACEHOLDER_3all:\364
The arithmetic width of PRESERVED_PLACEHOLDER_3all:\365 in direction PRESERVED_PLACEHOLDER_3all:\366 is
PRESERVED_PLACEHOLDER_3all:\367
the number of distinct integer values attained by PRESERVED_PLACEHOLDER_3all:\368 on lattice points of PRESERVED_PLACEHOLDER_3all:\369. Equivalently, it is the number of integer-normal hyperplanes slicing PRESERVED_PLACEHOLDER_3all:\3query3query3^ that actually contain lattice points (Loera et al., 5 Sep 2025).
This definition refines lattice width. The classical directional lattice width is
PRESERVED_PLACEHOLDER_3all:\3query3all:\3^
and the global lattice width is PRESERVED_PLACEHOLDER_3all:\3query3 OR title:\3. Since PRESERVED_PLACEHOLDER_3all:\373 ranges over an integer interval, one has
PRESERVED_PLACEHOLDER_3all:\374
Equality holds exactly when there are no gaps: every integer PRESERVED_PLACEHOLDER_3all:\375 between PRESERVED_PLACEHOLDER_3all:\376 and PRESERVED_PLACEHOLDER_3all:\377 occurs as PRESERVED_PLACEHOLDER_3all:\378 for some lattice point PRESERVED_PLACEHOLDER_3all:\379. Arithmetic width is therefore sensitive to missing slices that the classical lattice width does not detect (Loera et al., 5 Sep 2025).
The main structural theorem states that for a convex body PRESERVED_PLACEHOLDER_3all:\3(Gurjar et al., 2016) arithmetic width ROABP (Sentieys et al., 2022) word-length (0907.3780) constant width arithmetic circuits (Loera et al., 5 Sep 2025) convex bodies3query3^ whose affine span contains a lattice point, and PRESERVED_PLACEHOLDER_3all:\3(Gurjar et al., 2016) arithmetic width ROABP (Sentieys et al., 2022) word-length (0907.3780) constant width arithmetic circuits (Loera et al., 5 Sep 2025) convex bodies3all:\3, there exist constants PRESERVED_PLACEHOLDER_3all:\3(Gurjar et al., 2016) arithmetic width ROABP (Sentieys et al., 2022) word-length (0907.3780) constant width arithmetic circuits (Loera et al., 5 Sep 2025) convex bodies3 OR title:\3^ depending only on PRESERVED_PLACEHOLDER_3all:\383 such that for all integers PRESERVED_PLACEHOLDER_3all:\384,
PRESERVED_PLACEHOLDER_3all:\385
is exactly an arithmetic progression of step PRESERVED_PLACEHOLDER_3all:\386, where
PRESERVED_PLACEHOLDER_3all:\387
All gaps occur in two bounded end intervals. For rational polytopes PRESERVED_PLACEHOLDER_3all:\388, the fixed-direction function PRESERVED_PLACEHOLDER_3all:\389 is eventually a quasipolynomial of degree PRESERVED_PLACEHOLDER_3all:\3max_results3query3^ and period dividing PRESERVED_PLACEHOLDER_3all:\3max_results3all:\3, and the minimized global arithmetic width PRESERVED_PLACEHOLDER_3all:\3max_results3 OR title:\3^ is again eventually quasipolynomial of degree PRESERVED_PLACEHOLDER_3all:\393, with optimizing directions recurring periodically (Loera et al., 5 Sep 2025).
The algorithmic theory is dimension-sensitive. For fixed dimension PRESERVED_PLACEHOLDER_3all:\394 and rational input size PRESERVED_PLACEHOLDER_3all:\395, PRESERVED_PLACEHOLDER_3all:\396 can be computed in polynomial time by first using Barvinok’s algorithm to compute a short rational generating function
PRESERVED_PLACEHOLDER_3all:\397
then applying the monomial substitution PRESERVED_PLACEHOLDER_3all:\398 to obtain a univariate rational function
PRESERVED_PLACEHOLDER_3all:\399
whose distinct exponents are exactly PRESERVED_PLACEHOLDER_3 OR title:\3query3query3. Computing the global minimum PRESERVED_PLACEHOLDER_3 OR title:\3query3all:\3^ is reduced to a finite, though at worst exponential, set of candidate directions orthogonal to differences of lattice points (Loera et al., 5 Sep 2025).
This notion should not be confused with mean width in convex geometry. Mean width of a convex body PRESERVED_PLACEHOLDER_3 OR title:\3query3 OR title:\3^ is defined by
PRESERVED_PLACEHOLDER_3 OR title:\3query33^
with PRESERVED_PLACEHOLDER_3 OR title:\3query34, a spherical average of directional widths. Arithmetic width of a convex body, by contrast, counts attained integer levels on lattice points. The two notions address different questions: one is an averaged metric of continuous support geometry, the other a discrete measure of lattice occupation (&&&33all:\3&&&, Loera et al., 5 Sep 2025).