Taylor-Interpolated Cache in AATS
- Taylor-interpolated cache is a method that stores per-coordinate degree-K Taylor segments in fixed-size circular buffers for continuous history representation.
- It uses compile-time automatic differentiation to generate Taylor coefficients, ensuring C⁰ continuity while eliminating runtime heap allocations.
- The design supports exact, interpolation-free dense output and asynchronous, out-of-order evaluations with O(N) scaling in multi-rate delay differential equation solvers.
Searching arXiv for the cited paper and closely related work to ground the article. arXiv search: (Malik, 19 Jun 2026) The Taylor-interpolated cache is the history-storage and dense-output mechanism used inside the Asynchronous Adaptive Taylor Solver (AATS) for high-dimensional, multi-rate Delay Differential Equations (DDEs). In AATS, each coordinate is assigned an independent local clock, and its recent history is represented not by dynamically allocated samples but by fixed-size circular buffers of degree- Taylor segments. The resulting design combines statically allocated storage, compile-time Automatic Differentiation (AD), interpolation-free continuous dense-output evaluation, and asynchronous event scheduling. Within the formulation given for AATS, the cache provides exact continuous dense output, zero-allocation history storage, and time scaling for large, sparse delay systems (Malik, 19 Jun 2026).
1. Concept and functional role
Within AATS, the Taylor-interpolated cache is the per-coordinate repository of local Taylor expansions that represent the continuous trajectory required by delayed evaluations. For each coordinate , AATS allocates at startup a fixed-size circular buffer of capacity segments. No further heap allocations occur at runtime, and the buffer never grows beyond size . This design is introduced specifically to avoid the dynamic memory allocation otherwise required for continuous history tracking in DDE solvers (Malik, 19 Jun 2026).
Each stored segment represents the degree- Taylor expansion of about a local update time ,
The cache is therefore not a store of raw state snapshots alone; it is a store of local polynomial models that can be queried continuously in time. Because delayed-argument evaluations can read directly from earlier segments of a neighbor’s ring buffer, no global synchronization is needed during asynchronous multi-rate integration.
This organization implies that the cache is not an auxiliary optimization layered on top of a synchronous solver. It is part of the solver’s core representation of history and is coupled directly to local clocks, event scheduling, and adaptive stepping.
2. Storage layout and segment representation
The data structure is a statically allocated ring buffer indexed by coordinate and segment slot. The description given for AATS is:
4
For each coordinate 0, an integer 1 points to the most recently written slot. Once the buffer has filled, the oldest valid segment is 2. When a new segment is pushed,
3
then ringbuf[v] [i_new] = new_segment, and head[v] = i_new.
| Component | Stored data | Role |
|---|---|---|
ringbuf[v] [m] |
t0, c[0..K] |
One Taylor segment for coordinate 4 |
Segment.t0 |
Start time of the segment | Reference point for local evaluation |
Segment.c[0..K] |
5 Taylor coefficients | Polynomial representation of history |
head[v] |
Integer in 6 | Most recently written slot |
The coefficients are stored in scaled form. On a subinterval 7, AATS states the local Taylor expansion
8
and stores
9
so that evaluation becomes
0
The representation is significant because it converts history storage from a potentially open-ended sequence of time-stamped values into a bounded sequence of polynomial segments with fixed memory cost per coordinate.
3. Coefficient generation and continuity properties
AATS generates the 1 coefficients of each segment using a forward-mode AD graph that was built once at compile time for the right-hand side 2. Because 3 is fixed at compile time, the AD engine unrolls all loops and generates a sequence of fused multiply–add (FMA) instructions to compute the 4 derivatives in 5 time, entirely in registers (Malik, 19 Jun 2026).
The segment-generation procedure is specified as:
5
AATS further states that at each update the true coordinate value 6 is sampled exactly from its previous segment, or from initial data, and then the 7 coefficients are recomputed to match 8 and its derivatives at 9. As a consequence, the piecewise-polynomial trajectory is 0 continuous at each boundary.
That continuity result is central to the cache’s role in delayed systems. The cache is not merely a piecewise approximation assembled from independent local fits; each segment begins from an exactly sampled boundary value, so adjacent segments meet without a jump.
4. Querying history and asynchronous evaluation
The evaluation procedure, named EvaluateHistory, locates the segment whose time interval contains the query time and then evaluates the stored polynomial in Horner form:
6
This procedure is used both for present-time local updates and for delayed evaluations. In the asynchronous multi-rate design, each coordinate 1 maintains its own local clock, and events 2 are scheduled into a global RadixHeap in 3 amortized time. When an event 4 is popped, AATS performs the following steps: it calls EvaluateHistory(v,t) to obtain 5; it generates a new segment about 6 using GenerateSegment(v,t,x_v(t)); it computes the next step size 7 via the adaptive controller, using the highest-order AD coefficient as an error estimate; and it re-inserts 8 into the heap (Malik, 19 Jun 2026).
A notable property of the cache is that out-of-order evaluations are explicitly supported. When 9 lies in the past relative to ringbuf[j] head, EvaluateHistory searches the correct historical segment and evaluates that polynomial. This is essential in asynchronous DDE integration because delayed dependencies are not aligned to a single global clock.
5. Eviction policy, memory cost, and complexity
The eviction rule is implicit in the circular-buffer design: when head[v] wraps, the oldest segment is simply overwritten. AATS chooses 0 so that no needed segment, for delays up to 1, is ever overwritten prematurely. Since 2 and 3 are chosen once at startup based on 4 and minimal step size, no runtime heap allocations ever occur; the formulation describes this as zero allocation.
Memory per coordinate is
5
which yields total static memory
6
The time per event for coordinate 7 with local in-degree 8 is given as:
- 9 to pop/push RadixHeap.
- 0 to compute
ForwardAD(linear in 1 for linear terms, 2 for nonlinearity). - 3 to binomial-shift (time-alignment of neighbor polynomials, if this is done).
- 4 for each delay lookup.
The stated overall per-event cost is therefore
5
independent of 6. Over a time interval 7, the global work is
8
under sparse coupling, with 9 constant, and adaptive step control. The paper then concludes 0 scaling in 1, with no dynamic memory overhead. Benchmarks against state-of-the-art synchronous solvers are reported as empirically consistent with 2 execution scaling, including large-scale benchmarks up to 3 coordinates (Malik, 19 Jun 2026).
This suggests that the Taylor-interpolated cache is not only a storage optimization. Its bounded-memory design and locality properties are part of the mechanism by which AATS reduces redundant evaluations and preserves linear scaling under sparse coupling.
6. Distinction from other Taylor-based caches and extrapolators
The term “cache” appears in multiple computational settings, but the Taylor-interpolated cache in AATS is specific to continuous history representation in DDE solvers. A useful contrast appears in diffusion acceleration, where feature-caching methods reuse or extrapolate intermediate activations across timesteps rather than store delayed state histories. In that separate setting, polynomial-based extrapolators such as TaylorSeer are described as suffering from error accumulation and trajectory drift in low-step regimes, motivating the use of rational Padé approximants in TC-Padé (Cui et al., 3 Mar 2026).
That contrast helps delimit the concept. In AATS, the cache stores piecewise Taylor segments of coordinate trajectories, supports out-of-order historical queries, and is designed around zero-allocation history storage. In diffusion acceleration, cached objects are raw features or residuals, and the central issue is predictive extrapolation across denoising steps rather than exact continuous dense output. A plausible implication is that the shared language of “Taylor” and “cache” can obscure fundamentally different numerical roles: in AATS the cache is an integral component of asynchronous DDE integration, whereas in diffusion sampling it is part of an inference-acceleration strategy.
A related misconception is that a Taylor-based cache must entail interpolation overhead or unbounded memory growth. The AATS formulation explicitly rejects both assumptions: the design uses statically allocated circular buffers, and evaluation is interpolation-free because each history query is answered directly by evaluating the appropriate stored Taylor polynomial.