---
title: Taylor-Interpolated Cache in AATS
url: https://www.emergentmind.com/topics/taylor-interpolated-cache
type: topic
---

# Taylor-Interpolated Cache in AATS

Searching arXiv for the cited paper and closely related work to ground the article.
arXiv search: 2606.21044
The Taylor-interpolated cache is the history-storage and dense-output mechanism used inside the Asynchronous Adaptive Taylor Solver (AATS) for high-dimensional, multi-rate Delay Differential Equations (DDEs). In AATS, each coordinate is assigned an independent local clock, and its recent history is represented not by dynamically allocated samples but by fixed-size circular buffers of degree-\(K\) Taylor segments. The resulting design combines statically allocated storage, compile-time Automatic Differentiation (AD), interpolation-free continuous dense-output evaluation, and asynchronous event scheduling. Within the formulation given for AATS, the cache provides exact continuous dense output, zero-allocation history storage, and \(O(N)\) time scaling for large, sparse delay systems [2606.21044].

## 1. Concept and functional role

Within AATS, the Taylor-interpolated cache is the per-coordinate repository of local Taylor expansions that represent the continuous trajectory required by delayed evaluations. For each coordinate \(v=1\ldots N\), AATS allocates at startup a fixed-size circular buffer of capacity \(M\) segments. No further heap allocations occur at runtime, and the buffer never grows beyond size \(M\). This design is introduced specifically to avoid the dynamic memory allocation otherwise required for continuous history tracking in DDE solvers [2606.21044].

Each stored segment represents the degree-\(K\) Taylor expansion of \(x_v(t)\) about a local update time \(t_0\),
\[
x_v(t) \approx \sum_{k=0}^{K} c_k (t-t_0)^k.
\]
The cache is therefore not a store of raw state snapshots alone; it is a store of local polynomial models that can be queried continuously in time. Because delayed-argument evaluations \(x_j(t-\tau)\) can read directly from earlier segments of a neighbor’s ring buffer, no global synchronization is needed during asynchronous multi-rate integration.

This organization implies that the cache is not an auxiliary optimization layered on top of a synchronous solver. It is part of the solver’s core representation of history and is coupled directly to local clocks, event scheduling, and adaptive stepping.

## 2. Storage layout and segment representation

The data structure is a statically allocated ring buffer indexed by coordinate and segment slot. The description given for AATS is:

```text
struct Segment {
  double t0;                  // start time of the segment
  double c[0..K];             // K+1 Taylor coefficients c₀…c_K
};
Segment ringbuf[N][M];        // statically allocated at initialization
```

For each coordinate \(v\), an integer \(\text{head}[v] \in [0,M-1]\) points to the most recently written slot. Once the buffer has filled, the oldest valid segment is \((\text{head}[v]+1) \bmod M\). When a new segment is pushed,
\[
i_{\text{new}} = (\text{head}[v] + 1) \bmod M,
\]
then `ringbuf[v][i_new] = new_segment`, and `head[v] = i_new`.

| Component | Stored data | Role |
|---|---|---|
| `ringbuf[v][m]` | `t0`, `c[0..K]` | One Taylor segment for coordinate \(v\) |
| `Segment.t0` | Start time of the segment | Reference point for local evaluation |
| `Segment.c[0..K]` | \(K+1\) Taylor coefficients | Polynomial representation of history |
| `head[v]` | Integer in \([0,M-1]\) | Most recently written slot |

The coefficients are stored in scaled form. On a subinterval \([t_n,t_{n+1}]\), AATS states the local Taylor expansion
\[
f_v(t) \approx \sum_{k=0}^{K} \frac{f_v^{(k)}(t_n)}{k!}(t-t_n)^k,
\]
and stores
\[
c_k = \frac{f_v^{(k)}(t_n)}{k!},
\]
so that evaluation becomes
\[
x_v(t) \approx \sum_{k=0}^{K} c_k \,\Delta t^k,\qquad \Delta t=t-t_n.
\]

The representation is significant because it converts history storage from a potentially open-ended sequence of time-stamped values into a bounded sequence of polynomial segments with fixed memory cost per coordinate.

## 3. Coefficient generation and continuity properties

AATS generates the \(K+1\) coefficients of each segment using a forward-mode AD graph that was built once at compile time for the right-hand side \(f\). Because \(K\) is fixed at compile time, the AD engine unrolls all loops and generates a sequence of fused multiply–add (FMA) instructions to compute the \(K+1\) derivatives in \(O(K)\) time, entirely in registers [2606.21044].

The segment-generation procedure is specified as:

```text
procedure GenerateSegment(v, tₙ, x_v(tₙ)):
  // 1. Produce Taylor coefficients by compile-time AD
  c[0..K] ← ForwardAD_graph_v(x_v(tₙ))

  // 2. Package new segment
  seg.t0 = tₙ;
  seg.c[0..K] = c[0..K];

  // 3. Push into ring buffer
  i_new = (head[v] + 1) mod M
  ringbuf[v][i_new] = seg
  head[v] = i_new
  return
```

AATS further states that at each update the true coordinate value \(x_v(t_n)\) is sampled exactly from its previous segment, or from initial data, and then the \(K+1\) coefficients are recomputed to match \(f_v\) and its derivatives at \(t_n\). As a consequence, the piecewise-polynomial trajectory is \(C^0\) continuous at each boundary.

That continuity result is central to the cache’s role in delayed systems. The cache is not merely a piecewise approximation assembled from independent local fits; each segment begins from an exactly sampled boundary value, so adjacent segments meet without a jump.

## 4. Querying history and asynchronous evaluation

The evaluation procedure, named `EvaluateHistory`, locates the segment whose time interval contains the query time and then evaluates the stored polynomial in Horner form:

```text
function EvaluateHistory(v, t_query):
  // 1. Locate the segment s in ringbuf[v] so that
  //    ringbuf[v][s].t0 ≤ t_query < ringbuf[v][next].t0
  //    (since M is small and segments are in time order modulo M, one can do a binary or linear scan of M entries in O(log M) or even O(M)).
  find index s such that ringbuf[v][s].t0 ≤ t_query < ringbuf[v][(s+1) mod M].t0
  Δt = t_query − ringbuf[v][s].t0

  // 2. Horner-style evaluation in O(K)
  result = 0
  for k = K down to 0:
    result = result*Δt + ringbuf[v][s].c[k]
  return result
```

This procedure is used both for present-time local updates and for delayed evaluations. In the asynchronous multi-rate design, each coordinate \(v\) maintains its own local clock, and events \((t_n,v)\) are scheduled into a global RadixHeap in \(O(1)\) amortized time. When an event \((t,v)\) is popped, AATS performs the following steps: it calls `EvaluateHistory(v,t)` to obtain \(x_v(t)\); it generates a new segment about \(t\) using `GenerateSegment(v,t,x_v(t))`; it computes the next step size \(\Delta t\) via the adaptive controller, using the highest-order AD coefficient as an error estimate; and it re-inserts \((t+\Delta t,v)\) into the heap [2606.21044].

A notable property of the cache is that out-of-order evaluations are explicitly supported. When \(t_{\text{query}}\) lies in the past relative to `ringbuf[j]` head, `EvaluateHistory` searches the correct historical segment and evaluates that polynomial. This is essential in asynchronous DDE integration because delayed dependencies are not aligned to a single global clock.

## 5. Eviction policy, memory cost, and complexity

The eviction rule is implicit in the circular-buffer design: when `head[v]` wraps, the oldest segment is simply overwritten. AATS chooses \(M\) so that no needed segment, for delays up to \(\tau_{\max}\), is ever overwritten prematurely. Since \(M\) and \(K\) are chosen once at startup based on \(\tau_{\max}\) and minimal step size, no runtime heap allocations ever occur; the formulation describes this as zero allocation.

Memory per coordinate is
\[
M \text{ segments} \times \bigl[(K+1)\text{ doubles for } c\bigr] + 1 \text{ double for } t_0 + 1 \text{ int head}[v],
\]
which yields total static memory
\[
O(N\cdot M\cdot K).
\]

The time per event for coordinate \(v\) with local in-degree \(E_v\) is given as:

- \(O(1)\) to pop/push RadixHeap.
- \(O(E_v\cdot K + K^2)\) to compute `ForwardAD` (linear in \(E_v\) for linear terms, \(K^2\) for nonlinearity).
- \(O(E_v\cdot K^2)\) to binomial-shift (time-alignment of neighbor polynomials, if this is done).
- \(O(\log M + K)\) for each delay lookup.

The stated overall per-event cost is therefore
\[
O(E_v\cdot K^2 + K^2),
\]
independent of \(N\). Over a time interval \([0,T]\), the global work is
\[
\sum_{v=1}^{N} N_v \cdot O(E_v\cdot K^2 + K^2) = O\!\left(N\cdot \epsilon^{-1/(K+1)}\right)
\]
under sparse coupling, with \(\max E_v\) constant, and adaptive step control. The paper then concludes \(O(N)\) scaling in \(N\), with no dynamic memory overhead. Benchmarks against state-of-the-art synchronous solvers are reported as empirically consistent with \(O(N)\) execution scaling, including large-scale benchmarks up to \(N=10000\) coordinates [2606.21044].

This suggests that the Taylor-interpolated cache is not only a storage optimization. Its bounded-memory design and locality properties are part of the mechanism by which AATS reduces redundant evaluations and preserves linear scaling under sparse coupling.

## 6. Distinction from other Taylor-based caches and extrapolators

The term “cache” appears in multiple computational settings, but the Taylor-interpolated cache in AATS is specific to continuous history representation in DDE solvers. A useful contrast appears in diffusion acceleration, where feature-caching methods reuse or extrapolate intermediate activations across timesteps rather than store delayed state histories. In that separate setting, polynomial-based extrapolators such as TaylorSeer are described as suffering from error accumulation and trajectory drift in low-step regimes, motivating the use of rational Padé approximants in TC-Padé [2603.02943].

That contrast helps delimit the concept. In AATS, the cache stores piecewise Taylor segments of coordinate trajectories, supports out-of-order historical queries, and is designed around zero-allocation history storage. In diffusion acceleration, cached objects are raw features or residuals, and the central issue is predictive extrapolation across denoising steps rather than exact continuous dense output. A plausible implication is that the shared language of “Taylor” and “cache” can obscure fundamentally different numerical roles: in AATS the cache is an integral component of asynchronous DDE integration, whereas in diffusion sampling it is part of an inference-acceleration strategy.

A related misconception is that a Taylor-based cache must entail interpolation overhead or unbounded memory growth. The AATS formulation explicitly rejects both assumptions: the design uses statically allocated circular buffers, and evaluation is interpolation-free because each history query is answered directly by evaluating the appropriate stored Taylor polynomial.

Source: https://www.emergentmind.com/topics/taylor-interpolated-cache