Papers
Topics
Authors
Recent
Search
2000 character limit reached

Arithmetic Width in Computation

Updated 10 July 2026
  • Arithmetic width is a domain-specific parameter that measures the size of intermediate computations in read-once oblivious arithmetic branching programs and constant-width circuits.
  • In computer arithmetic, it controls bit-width allocation, influencing dynamic range, precision, energy consumption, and overall computational performance.
  • In discrete geometry, arithmetic width captures the number of distinct integer levels attained by linear functionals on lattice points, aiding in quasipolynomial analyses and structural insights.

Searching arXiv for recent and foundational uses of “arithmetic width” across algebraic complexity, computer arithmetic, and discrete geometry. {"3query3 width\"3 OR title:\3"Arithmetic Width\"","max_results":3all:\3query3,"sort_by":"relevance"} {"3query3 arithmetic width ROABP (&&&3all:\3&&&) word-length (&&&3 OR title:\3&&&) constant width arithmetic circuits (Loera et al., 5 Sep 2025) convex bodies","max_results":5,"sort_by":"submittedDate"} Arithmetic width is a domain-dependent notion of bounded intermediate complexity. In algebraic complexity, it is the width parameter PRESERVED_PLACEHOLDER_3query3^ of a read-once oblivious arithmetic branching program (ROABP) or of a constant-width arithmetic circuit, measuring the size of the largest intermediate matrix or layer. In computer arithmetic, it is the total bit-width or word-length that controls dynamic range and precision. In discrete geometry, it is the number of distinct values an integer linear functional attains on the lattice points of a convex body. Across these settings, width serves as a structural restriction that simultaneously limits expressivity and enables stronger algorithmic results, especially for polynomial identity testing, customized-precision arithmetic, and lattice-slice analysis (&&&3query3&&&, &&&3all:\3&&&, Loera et al., 5 Sep 2025).

3all:\3. Arithmetic width in algebraic computation

An Arithmetic Branching Program (ABP) is a layered directed acyclic graph computing a polynomial by summing over source–sink path-weights. In matrix form, an ABP of length PRESERVED_PLACEHOLDER_3all:\3^ and width PRESERVED_PLACEHOLDER_3 OR title:\3^ over a field F\mathbb{F} is a sequence

D1(xi1)F[xi1]1×w,D2(xi2)F[xi2]w×w,,Dn(xin)F[xin]w×1,D_1(x_{i_1})\in\mathbb{F}[x_{i_1}]^{1\times w},\quad D_2(x_{i_2})\in\mathbb{F}[x_{i_2}]^{w\times w},\dots, D_n(x_{i_n})\in\mathbb{F}[x_{i_n}]^{w\times 1},

computing

f(x1,,xn)=D1(xi1)D2(xi2)Dn(xin)F[x1,,xn].f(x_1,\dots,x_n)=D_1(x_{i_1})D_2(x_{i_2})\cdots D_n(x_{i_n})\in\mathbb{F}[x_1,\dots,x_n].

Its width ww is the size of the largest intermediate matrix, and the total size is the cost of representing all the univariate entries. When each layer is a univariate polynomial in a different variable, the ABP is oblivious; if each variable appears exactly once, it is a read-once oblivious ABP. In this setting, the width ww of an ROABP is often called the arithmetic width of the polynomial it computes (&&&3query3&&&).

This width parameter has an operational interpretation. It measures the “space” needed to build up a polynomial in a one-pass read of the variables, or, in the non-commutative viewpoint, the matrix-multiplication dimension required by the computation. Small width imposes strong structural constraints. The ROABP literature uses those constraints to obtain stronger black-box PIT algorithms than were known for general arithmetic circuits (&&&3query3&&&).

A related notion appears in layered arithmetic circuits. If non-leaf nodes are partitioned into layers V1,V2,,VtV_1,V_2,\dots,V_t, then the width is

maxi>1Vi.\max_{i>1}|V_i|.

A layered circuit is staggered if in every layer PRESERVED_PLACEHOLDER_3all:\3query3^ at most one gate is not of the form “multiply by PRESERVED_PLACEHOLDER_3all:\3all:\3 from the previous layer; equivalently, staggered width-PRESERVED_PLACEHOLDER_3all:\3 OR title:\3^ circuits are exactly straight-line programs with PRESERVED_PLACEHOLDER_3all:\33^ registers. Here too, width bounds how many partial values can be kept alive at once, restricting how much parallelism or mixing of variables the circuit can perform (&&&3 OR title:\3&&&).

3 OR title:\3. Hitting sets and derandomization for bounded-width ROABPs

The ROABP identity-testing results of Gurjar, Korwar, and Saxena sharpen the algorithmic meaning of arithmetic width. For the known-order model,

PRESERVED_PLACEHOLDER_3all:\34

if PRESERVED_PLACEHOLDER_3all:\35 or PRESERVED_PLACEHOLDER_3all:\36, one can deterministically construct in time

PRESERVED_PLACEHOLDER_3all:\37

a hitting set of size

PRESERVED_PLACEHOLDER_3all:\38

In particular, for constant PRESERVED_PLACEHOLDER_3all:\39 this is a truly polynomial-size hitting set. For the any-order model,

PRESERVED_PLACEHOLDER_3 OR title:\3query3^

one can build over any field a hitting set of size

PRESERVED_PLACEHOLDER_3 OR title:\3all:\3^

These improve the previously known bounds and mark the first polynomial-size hitting set for constant-width ROABP in the known-order case (&&&3query3&&&).

The structural core of the known-order argument is the rank of the partial-derivative matrix. For a bivariate polynomial

PRESERVED_PLACEHOLDER_3 OR title:\3 OR title:\3^

the matrix

PRESERVED_PLACEHOLDER_3 OR title:\33^

satisfies a rank-width relation: if

PRESERVED_PLACEHOLDER_3 OR title:\34

then PRESERVED_PLACEHOLDER_3 OR title:\35. This is used to prove that the substitution

PRESERVED_PLACEHOLDER_3 OR title:\36

preserves nonzeroness for width-PRESERVED_PLACEHOLDER_3 OR title:\37 ROABP polynomials. The proof then iterates a variable-halving lemma: if PRESERVED_PLACEHOLDER_3 OR title:\38 has a width-PRESERVED_PLACEHOLDER_3 OR title:\39 ROABP in the order F\mathbb{F}3query3, then substituting

F\mathbb{F}3all:\3^

for F\mathbb{F}3 OR title:\3^ yields a nonzero polynomial in F\mathbb{F}3 still computable by a width-F\mathbb{F}4 ROABP in the order F\mathbb{F}5. After F\mathbb{F}6 rounds one gets a univariate test of degree F\mathbb{F}7 (&&&3query3&&&).

The any-order argument proceeds differently. First, a low-degree bivariate shift forces all small ROABPs of width F\mathbb{F}8 in any order to become F\mathbb{F}9-concentrated, where every coefficient of a large-support monomial lies in the span of coefficients of support D1(xi1)F[xi1]1×w,D2(xi2)F[xi2]w×w,,Dn(xin)F[xin]w×1,D_1(x_{i_1})\in\mathbb{F}[x_{i_1}]^{1\times w},\quad D_2(x_{i_2})\in\mathbb{F}[x_{i_2}]^{w\times w},\dots, D_n(x_{i_n})\in\mathbb{F}[x_{i_n}]^{w\times 1},3query3, for D1(xi1)F[xi1]1×w,D2(xi2)F[xi2]w×w,,Dn(xin)F[xin]w×1,D_1(x_{i_1})\in\mathbb{F}[x_{i_1}]^{1\times w},\quad D_2(x_{i_2})\in\mathbb{F}[x_{i_2}]^{w\times w},\dots, D_n(x_{i_n})\in\mathbb{F}[x_{i_n}]^{w\times 1},3all:\3. The proof uses a basis-isolating weight assignment and the fact that shifting by a basis-isolating map yields D1(xi1)F[xi1]1×w,D2(xi2)F[xi2]w×w,,Dn(xin)F[xin]w×1,D_1(x_{i_1})\in\mathbb{F}[x_{i_1}]^{1\times w},\quad D_2(x_{i_2})\in\mathbb{F}[x_{i_2}]^{w\times w},\dots, D_n(x_{i_n})\in\mathbb{F}[x_{i_n}]^{w\times 1},3 OR title:\3-concentration in any algebra D1(xi1)F[xi1]1×w,D2(xi2)F[xi2]w×w,,Dn(xin)F[xin]w×1,D_1(x_{i_1})\in\mathbb{F}[x_{i_1}]^{1\times w},\quad D_2(x_{i_2})\in\mathbb{F}[x_{i_2}]^{w\times w},\dots, D_n(x_{i_n})\in\mathbb{F}[x_{i_n}]^{w\times 1},3. Second, once D1(xi1)F[xi1]1×w,D2(xi2)F[xi2]w×w,,Dn(xin)F[xin]w×1,D_1(x_{i_1})\in\mathbb{F}[x_{i_1}]^{1\times w},\quad D_2(x_{i_2})\in\mathbb{F}[x_{i_2}]^{w\times w},\dots, D_n(x_{i_n})\in\mathbb{F}[x_{i_n}]^{w\times 1},4 is known to have an D1(xi1)F[xi1]1×w,D2(xi2)F[xi2]w×w,,Dn(xin)F[xin]w×1,D_1(x_{i_1})\in\mathbb{F}[x_{i_1}]^{1\times w},\quad D_2(x_{i_2})\in\mathbb{F}[x_{i_2}]^{w\times w},\dots, D_n(x_{i_n})\in\mathbb{F}[x_{i_n}]^{w\times 1},5-support monomial with nonzero coefficient, black-box testing reduces to an D1(xi1)F[xi1]1×w,D2(xi2)F[xi2]w×w,,Dn(xin)F[xin]w×1,D_1(x_{i_1})\in\mathbb{F}[x_{i_1}]^{1\times w},\quad D_2(x_{i_2})\in\mathbb{F}[x_{i_2}]^{w\times w},\dots, D_n(x_{i_n})\in\mathbb{F}[x_{i_n}]^{w\times 1},6-variate ROABP in any order of width D1(xi1)F[xi1]1×w,D2(xi2)F[xi2]w×w,,Dn(xin)F[xin]w×1,D_1(x_{i_1})\in\mathbb{F}[x_{i_1}]^{1\times w},\quad D_2(x_{i_2})\in\mathbb{F}[x_{i_2}]^{w\times w},\dots, D_n(x_{i_n})\in\mathbb{F}[x_{i_n}]^{w\times 1},7 and individual degree D1(xi1)F[xi1]1×w,D2(xi2)F[xi2]w×w,,Dn(xin)F[xin]w×1,D_1(x_{i_1})\in\mathbb{F}[x_{i_1}]^{1\times w},\quad D_2(x_{i_2})\in\mathbb{F}[x_{i_2}]^{w\times w},\dots, D_n(x_{i_n})\in\mathbb{F}[x_{i_n}]^{w\times 1},8, recovering the D1(xi1)F[xi1]1×w,D2(xi2)F[xi2]w×w,,Dn(xin)F[xin]w×1,D_1(x_{i_1})\in\mathbb{F}[x_{i_1}]^{1\times w},\quad D_2(x_{i_2})\in\mathbb{F}[x_{i_2}]^{w\times w},\dots, D_n(x_{i_n})\in\mathbb{F}[x_{i_n}]^{w\times 1},9 bound (&&&3query3&&&).

3. Constant-width arithmetic circuits and lower bounds

Width is not only an algorithmic aid; it is also a source of lower bounds. For every f(x1,,xn)=D1(xi1)D2(xi2)Dn(xin)F[x1,,xn].f(x_1,\dots,x_n)=D_1(x_{i_1})D_2(x_{i_2})\cdots D_n(x_{i_n})\in\mathbb{F}[x_1,\dots,x_n].3query3, Arvind, Joglekar, and Srinivasan construct an explicit polynomial that can be computed by a linear-sized monotone circuit of width f(x1,,xn)=D1(xi1)D2(xi2)Dn(xin)F[x1,,xn].f(x_1,\dots,x_n)=D_1(x_{i_1})D_2(x_{i_2})\cdots D_n(x_{i_n})\in\mathbb{F}[x_1,\dots,x_n].3all:\3^ but has no subexponential-sized monotone circuit of width f(x1,,xn)=D1(xi1)D2(xi2)Dn(xin)F[x1,,xn].f(x_1,\dots,x_n)=D_1(x_{i_1})D_2(x_{i_2})\cdots D_n(x_{i_n})\in\mathbb{F}[x_1,\dots,x_n].3 OR title:\3. Their family f(x1,,xn)=D1(xi1)D2(xi2)Dn(xin)F[x1,,xn].f(x_1,\dots,x_n)=D_1(x_{i_1})D_2(x_{i_2})\cdots D_n(x_{i_n})\in\mathbb{F}[x_1,\dots,x_n].3 is defined recursively, is homogeneous of degree f(x1,,xn)=D1(xi1)D2(xi2)Dn(xin)F[x1,,xn].f(x_1,\dots,x_n)=D_1(x_{i_1})D_2(x_{i_2})\cdots D_n(x_{i_n})\in\mathbb{F}[x_1,\dots,x_n].4 on f(x1,,xn)=D1(xi1)D2(xi2)Dn(xin)F[x1,,xn].f(x_1,\dots,x_n)=D_1(x_{i_1})D_2(x_{i_2})\cdots D_n(x_{i_n})\in\mathbb{F}[x_1,\dots,x_n].5 variables, and has exactly

f(x1,,xn)=D1(xi1)D2(xi2)Dn(xin)F[x1,,xn].f(x_1,\dots,x_n)=D_1(x_{i_1})D_2(x_{i_2})\cdots D_n(x_{i_n})\in\mathbb{F}[x_1,\dots,x_n].6

distinct monomials. The separation theorem states that for fixed f(x1,,xn)=D1(xi1)D2(xi2)Dn(xin)F[x1,,xn].f(x_1,\dots,x_n)=D_1(x_{i_1})D_2(x_{i_2})\cdots D_n(x_{i_n})\in\mathbb{F}[x_1,\dots,x_n].7, for all sufficiently large f(x1,,xn)=D1(xi1)D2(xi2)Dn(xin)F[x1,,xn].f(x_1,\dots,x_n)=D_1(x_{i_1})D_2(x_{i_2})\cdots D_n(x_{i_n})\in\mathbb{F}[x_1,\dots,x_n].8, any monotone width-f(x1,,xn)=D1(xi1)D2(xi2)Dn(xin)F[x1,,xn].f(x_1,\dots,x_n)=D_1(x_{i_1})D_2(x_{i_2})\cdots D_n(x_{i_n})\in\mathbb{F}[x_1,\dots,x_n].9 circuit ww3query3^ with

ww3all:\3^

must have size

ww3 OR title:\3^

On the other hand, ww3 admits a width-ww4 monotone circuit of size ww5 (&&&3 OR title:\3&&&).

This yields infinite hierarchies. For every fixed ww6, the class of monotone width-ww7 circuits is strictly contained in monotone width-ww8 circuits for infinitely many input sizes, so the monotone constant-width hierarchy is infinite. Because width-ww9 staggered circuits can be unfolded into depth-ww3query3^ circuits of polynomial blow-up, the same strictness holds for the monotone constant-depth hierarchy. The argument also carries over to the noncommutative setting (&&&3 OR title:\3&&&).

The same paper develops hardness–randomness tradeoffs for identity testing constant-width commutative circuits. Two technical lemmata are central: if ww3all:\3^ of total degree ww3 OR title:\3^ has a staggered width-ww3, size-ww4 circuit, then each homogeneous component ww5 has a staggered width-ww6, size ww7 circuit; and if ww8 has a staggered width-ww9, size-V1,V2,,VtV_1,V_2,\dots,V_t3query3^ circuit and V1,V2,,VtV_1,V_2,\dots,V_t3all:\3, then under V1,V2,,VtV_1,V_2,\dots,V_t3 OR title:\3, any annihilated series V1,V2,,VtV_1,V_2,\dots,V_t3 with V1,V2,,VtV_1,V_2,\dots,V_t4 admits a staggered circuit of width V1,V2,,VtV_1,V_2,\dots,V_t5 and size V1,V2,,VtV_1,V_2,\dots,V_t6. These ingredients support a Nisan–Wigderson-style generator and a deterministic subexponential-time PIT for width-V1,V2,,VtV_1,V_2,\dots,V_t7 circuits under suitable lower-bound assumptions (&&&3 OR title:\3&&&).

A common misconception is to identify bounded width with a mere size bound. The circuit results indicate that width is a finer-grained restriction: linear-size width-V1,V2,,VtV_1,V_2,\dots,V_t8 computations can be inaccessible to subexponential-size width-V1,V2,,VtV_1,V_2,\dots,V_t9 computations, even in monotone models. This suggests that arithmetic width controls a qualitative resource, not only a quantitative one (&&&3 OR title:\3&&&).

4. Arithmetic width as word-length in computer arithmetic

In digital number representation, arithmetic width is the total bit-width allocated to a datum. The total bit-width maxi>1Vi.\max_{i>1}|V_i|.3query3, sometimes called word-length, is the single most critical parameter because it must be apportioned between range and precision. In signed fixed-point maxi>1Vi.\max_{i>1}|V_i|.3all:\3, the total width is

maxi>1Vi.\max_{i>1}|V_i|.3 OR title:\3^

where maxi>1Vi.\max_{i>1}|V_i|.3 is the integer-part word-length, including the sign bit, and maxi>1Vi.\max_{i>1}|V_i|.4 is the fractional-part word-length. If maxi>1Vi.\max_{i>1}|V_i|.5 is the maxi>1Vi.\max_{i>1}|V_i|.6-bit two’s-complement integer, then

maxi>1Vi.\max_{i>1}|V_i|.7

with representable range

maxi>1Vi.\max_{i>1}|V_i|.8

and quantization step

maxi>1Vi.\max_{i>1}|V_i|.9

In radix-3 OR title:\3^ floating-point, the total width is partitioned as

PRESERVED_PLACEHOLDER_3all:\3query3query3^

and a real value is represented as

PRESERVED_PLACEHOLDER_3all:\3query3all:\3^

Two immediate consequences are the dynamic range

PRESERVED_PLACEHOLDER_3all:\3query3 OR title:\3^

and the unit round-off

PRESERVED_PLACEHOLDER_3all:\3query33^

These formulas make explicit that width governs both dynamic range and resolution (&&&3all:\3&&&).

At the operator level, the reported 3 OR title:\38 nm FDSOI, 3 OR title:\3query3query3^ MHz synthesis results show that the effect of width depends strongly on representation.

Operator Area, delay Energy
33 OR title:\3-bit float add 653 PRESERVED_PLACEHOLDER_3all:\3query34, 3 OR title:\3.43 OR title:\3^ ns PRESERVED_PLACEHOLDER_3all:\3query35 fJ
33 OR title:\3-bit int add 3all:\389 PRESERVED_PLACEHOLDER_3all:\3query36, 3all:\3.3query3 ns PRESERVED_PLACEHOLDER_3all:\3query37 fJ
33 OR title:\3-bit float mul 3all:\3543 PRESERVED_PLACEHOLDER_3all:\3query38, 3 OR title:\3.3query39 ns PRESERVED_PLACEHOLDER_3all:\3query39 fJ
33 OR title:\3-bit int mul 3 OR title:\3 OR title:\389 PRESERVED_PLACEHOLDER_3all:\3all:\3query3, 3 OR title:\3.38 ns PRESERVED_PLACEHOLDER_3all:\3all:\3all:\3^ fJ

The data imply distinct tradeoffs. The 33 OR title:\3-bit float adder is approximately PRESERVED_PLACEHOLDER_3all:\3all:\3 OR title:\3^ larger, approximately PRESERVED_PLACEHOLDER_3all:\3all:\33^ slower, and approximately PRESERVED_PLACEHOLDER_3all:\3all:\34 more energy than the 33 OR title:\3-bit integer adder. By contrast, the 33 OR title:\3-bit float multiplier is area-smaller but approximately PRESERVED_PLACEHOLDER_3all:\3all:\35 more energy because of exponent and shift management. In the width sweep from 8 to 3all:\36 bits, floating-point adders are always approximately PRESERVED_PLACEHOLDER_3all:\3all:\36 larger, approximately PRESERVED_PLACEHOLDER_3all:\3all:\37 slower, and PRESERVED_PLACEHOLDER_3all:\3all:\38–PRESERVED_PLACEHOLDER_3all:\3all:\3 more energy, while floating-point multipliers are smaller in area, slightly faster, but still PRESERVED_PLACEHOLDER_3all:\3 OR title:\3query3–PRESERVED_PLACEHOLDER_3all:\3 OR title:\3all:\3^ more energy (&&&3all:\3&&&).

Application-level studies refine that picture. For 8-bit 3 OR title:\3-D K-Means, the floating-point format PRESERVED_PLACEHOLDER_3all:\3 OR title:\3 OR title:\3^ is approximately PRESERVED_PLACEHOLDER_3all:\3 OR title:\33^ bigger and its distance kernel consumes approximately PRESERVED_PLACEHOLDER_3all:\3 OR title:\34 more energy than PRESERVED_PLACEHOLDER_3all:\3 OR title:\35, yet it converges in half as many iterations and yields PRESERVED_PLACEHOLDER_3all:\3 OR title:\36 lower CMSE and roughly half the misclassification rate. For 3all:\36-bit K-Means, the converse holds: fixed-point PRESERVED_PLACEHOLDER_3all:\3 OR title:\37 is both smaller, uses PRESERVED_PLACEHOLDER_3all:\3 OR title:\38 less energy overall, and reaches equal or better CMSE and ER than floatPRESERVED_PLACEHOLDER_3all:\3 OR title:\39. For FFT-3all:\36, fixed-point consistently beats floating-point for any equal total width, both in energy per operation and in MSE, because the dynamic range is modest and a 7-bit exponent in floating-point wastes bits that could have boosted the mantissa (&&&3all:\3&&&).

These results support an application-contingent view of arithmetic width. Fixed-point is almost always smaller, faster, and lower-energy per operation than floating-point when uniform quantization and explicit scaling are acceptable. Floating-point becomes advantageous when automatic scaling over a wide dynamic range materially improves convergence or prevents saturation and overflow (&&&3all:\3&&&).

5. Customized precision and non-native-width arithmetic

When the desired arithmetic width is not supported natively, recent work treats width itself as a software- or compiler-controlled design variable. The SAMD method, “Scalar Arithmetic Multiple Data,” embeds a vector architecture with custom bit-width lanes inside fixed-width scalar arithmetic. For native register width PRESERVED_PLACEHOLDER_3all:\33query3, virtual lane width PRESERVED_PLACEHOLDER_3all:\33all:\3, and PRESERVED_PLACEHOLDER_3all:\33 OR title:\3^ spacer bits per lane, one uses

PRESERVED_PLACEHOLDER_3all:\333^

lanes. Spacer bits prevent carries or overflows from bleeding between lanes, and ordinary scalar arithmetic plus bitwise operations implements lane-wise add, sub, mul, and even 3all:\3-D convolution. The paper reports a maximum speedup of approximately PRESERVED_PLACEHOLDER_3all:\334 on an Intel platform and approximately PRESERVED_PLACEHOLDER_3all:\335 on an ARM platform versus quantization to native 8-bit integers; on Cortex-A57, the temporary-spacer configuration reaches approximately PRESERVED_PLACEHOLDER_3all:\336 at 6-bit, PRESERVED_PLACEHOLDER_3all:\337 at 4-bit, PRESERVED_PLACEHOLDER_3all:\338 at 3-bit, and PRESERVED_PLACEHOLDER_3all:\339 at 3 OR title:\3-bit, while permanent spacer reaches up to PRESERVED_PLACEHOLDER_3all:\3max_results3query3^ at 3 OR title:\3-bit (&&&3 OR title:\3all:\3&&&).

The essential point is that width below 8 bits can be exploited without specialized hardware. One scalar multiply can simultaneously compute all PRESERVED_PLACEHOLDER_3all:\3max_results3all:\3^ lanes of partial products, so for small convolution kernels the overall multiply-plus-add count can drop substantially. The method therefore targets a regime in which custom widths are not merely storage formats but active algorithmic accelerants (&&&3 OR title:\3all:\3&&&).

A different non-native-width strategy appears in multi-word modular arithmetic for cryptographic kernels. In MoMA, if PRESERVED_PLACEHOLDER_3all:\3max_results3 OR title:\3^ is the machine-word size and PRESERVED_PLACEHOLDER_3all:\343 the required integer bit-width, each PRESERVED_PLACEHOLDER_3all:\344-bit integer is represented as

PRESERVED_PLACEHOLDER_3all:\345

A rewrite system recursively splits every PRESERVED_PLACEHOLDER_3all:\346-bit type into two PRESERVED_PLACEHOLDER_3all:\347-bit halves, replaces high-level add, sub, mul, and mod by multi-word sequences, prunes limbs known to be zero for non-power-of-two widths, and continues until all operations fit in the native word size. The implementation integrates this pass into SPIRAL and combines it with Barrett reduction, fused add-with-carry when PRESERVED_PLACEHOLDER_3all:\348, and rule selection between schoolbook and Karatsuba multiplication (&&&3 OR title:\33&&&).

The reported GPU results make arithmetic width an explicit performance knob. For BLAS Level-3all:\3^ operations on 3all:\3 OR title:\38–3all:\3query3 OR title:\34-bit widths, MoMA outperforms GRNS by PRESERVED_PLACEHOLDER_3all:\349–PRESERVED_PLACEHOLDER_3all:\3sort_by3query3^ and GMP by PRESERVED_PLACEHOLDER_3all:\3sort_by3all:\3–PRESERVED_PLACEHOLDER_3all:\3sort_by3 OR title:\3^ for add/sub, and by PRESERVED_PLACEHOLDER_3all:\353–PRESERVED_PLACEHOLDER_3all:\3 for mul/axpy. For NTT, 3 OR title:\356-bit performance on RTX 43query3max_results3query3^ is PRESERVED_PLACEHOLDER_3all:\355 faster than ICICLE on H3all:\3query3query3^ and within PRESERVED_PLACEHOLDER_3all:\356 of an ASIC; for 3all:\3 OR title:\38-bit NTT, MoMA on RTX 43query3max_results3query3^ outperforms an FHE ASIC by PRESERVED_PLACEHOLDER_3all:\357, and on H3all:\3query3query3^ by PRESERVED_PLACEHOLDER_3all:\358. The paper also notes that non-power-of-two widths such as 383all:\3^ and 753 bits avoid zero-padding and gain PRESERVED_PLACEHOLDER_3all:\359–PRESERVED_PLACEHOLDER_3all:\3relevance3query3^ speedups over naively padded code, while dynamically varying precision is not supported because kernels must fix PRESERVED_PLACEHOLDER_3all:\3relevance3all:\3^ at compile time (&&&3 OR title:\33&&&).

Taken together, SAMD and MoMA show that arithmetic width can be virtualized. In one case, width is packed into scalar registers; in the other, it is decomposed into limbs and lowered by compiler rewrites. Both approaches treat width as a first-class compilation parameter rather than a fixed ISA feature (&&&3 OR title:\3all:\3&&&, &&&3 OR title:\33&&&).

6. Arithmetic width of convex bodies

In discrete geometry, arithmetic width has a different and more literal meaning. For a convex body PRESERVED_PLACEHOLDER_3all:\3relevance3 OR title:\3^ and a nonzero integer vector PRESERVED_PLACEHOLDER_3all:\363, let

PRESERVED_PLACEHOLDER_3all:\364

The arithmetic width of PRESERVED_PLACEHOLDER_3all:\365 in direction PRESERVED_PLACEHOLDER_3all:\366 is

PRESERVED_PLACEHOLDER_3all:\367

the number of distinct integer values attained by PRESERVED_PLACEHOLDER_3all:\368 on lattice points of PRESERVED_PLACEHOLDER_3all:\369. Equivalently, it is the number of integer-normal hyperplanes slicing PRESERVED_PLACEHOLDER_3all:\3query3query3^ that actually contain lattice points (Loera et al., 5 Sep 2025).

This definition refines lattice width. The classical directional lattice width is

PRESERVED_PLACEHOLDER_3all:\3query3all:\3^

and the global lattice width is PRESERVED_PLACEHOLDER_3all:\3query3 OR title:\3. Since PRESERVED_PLACEHOLDER_3all:\373 ranges over an integer interval, one has

PRESERVED_PLACEHOLDER_3all:\374

Equality holds exactly when there are no gaps: every integer PRESERVED_PLACEHOLDER_3all:\375 between PRESERVED_PLACEHOLDER_3all:\376 and PRESERVED_PLACEHOLDER_3all:\377 occurs as PRESERVED_PLACEHOLDER_3all:\378 for some lattice point PRESERVED_PLACEHOLDER_3all:\379. Arithmetic width is therefore sensitive to missing slices that the classical lattice width does not detect (Loera et al., 5 Sep 2025).

The main structural theorem states that for a convex body PRESERVED_PLACEHOLDER_3all:\3(Gurjar et al., 2016) arithmetic width ROABP (Sentieys et al., 2022) word-length (0907.3780) constant width arithmetic circuits (Loera et al., 5 Sep 2025) convex bodies3query3^ whose affine span contains a lattice point, and PRESERVED_PLACEHOLDER_3all:\3(Gurjar et al., 2016) arithmetic width ROABP (Sentieys et al., 2022) word-length (0907.3780) constant width arithmetic circuits (Loera et al., 5 Sep 2025) convex bodies3all:\3, there exist constants PRESERVED_PLACEHOLDER_3all:\3(Gurjar et al., 2016) arithmetic width ROABP (Sentieys et al., 2022) word-length (0907.3780) constant width arithmetic circuits (Loera et al., 5 Sep 2025) convex bodies3 OR title:\3^ depending only on PRESERVED_PLACEHOLDER_3all:\383 such that for all integers PRESERVED_PLACEHOLDER_3all:\384,

PRESERVED_PLACEHOLDER_3all:\385

is exactly an arithmetic progression of step PRESERVED_PLACEHOLDER_3all:\386, where

PRESERVED_PLACEHOLDER_3all:\387

All gaps occur in two bounded end intervals. For rational polytopes PRESERVED_PLACEHOLDER_3all:\388, the fixed-direction function PRESERVED_PLACEHOLDER_3all:\389 is eventually a quasipolynomial of degree PRESERVED_PLACEHOLDER_3all:\3max_results3query3^ and period dividing PRESERVED_PLACEHOLDER_3all:\3max_results3all:\3, and the minimized global arithmetic width PRESERVED_PLACEHOLDER_3all:\3max_results3 OR title:\3^ is again eventually quasipolynomial of degree PRESERVED_PLACEHOLDER_3all:\393, with optimizing directions recurring periodically (Loera et al., 5 Sep 2025).

The algorithmic theory is dimension-sensitive. For fixed dimension PRESERVED_PLACEHOLDER_3all:\394 and rational input size PRESERVED_PLACEHOLDER_3all:\395, PRESERVED_PLACEHOLDER_3all:\396 can be computed in polynomial time by first using Barvinok’s algorithm to compute a short rational generating function

PRESERVED_PLACEHOLDER_3all:\397

then applying the monomial substitution PRESERVED_PLACEHOLDER_3all:\398 to obtain a univariate rational function

PRESERVED_PLACEHOLDER_3all:\399

whose distinct exponents are exactly PRESERVED_PLACEHOLDER_3 OR title:\3query3query3. Computing the global minimum PRESERVED_PLACEHOLDER_3 OR title:\3query3all:\3^ is reduced to a finite, though at worst exponential, set of candidate directions orthogonal to differences of lattice points (Loera et al., 5 Sep 2025).

This notion should not be confused with mean width in convex geometry. Mean width of a convex body PRESERVED_PLACEHOLDER_3 OR title:\3query3 OR title:\3^ is defined by

PRESERVED_PLACEHOLDER_3 OR title:\3query33^

with PRESERVED_PLACEHOLDER_3 OR title:\3query34, a spherical average of directional widths. Arithmetic width of a convex body, by contrast, counts attained integer levels on lattice points. The two notions address different questions: one is an averaged metric of continuous support geometry, the other a discrete measure of lattice occupation (&&&33all:\3&&&, Loera et al., 5 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Arithmetic Width.