---
title: Time-Since-Last-Reward (TSLR) Dynamics
url: https://www.emergentmind.com/topics/time-since-last-reward-tslr
type: topic
---

# Time-Since-Last-Reward (TSLR) Dynamics

Searching arXiv for the cited TSLR-related papers to ground the article in published work.
{"queries":[{"query":"Time-Since-Last-Reward combinatorial multi-armed bandit 2509.12457"},{"query":"GenARM Autoregressive Reward Model 2410.08193"},{"query":"Demystifying the Recency Heuristic in Temporal-Difference Learning 2406.12284"}]}
Time-Since-Last-Reward (TSLR) denotes a history-dependent temporal variable that measures how long it has been since a reward event last occurred. In its explicit combinatorial multi-armed bandit formulation, TSLR is a per-arm discrete counter with reset-on-success dynamics and an interpretation closely aligned with age-of-information [2509.12457]. Closely related temporal constructions also appear in autoregressive test-time alignment, sequential recommendation, temporal-difference credit assignment, quantitative reward monitoring for non-Markovian reinforcement learning, non-stationary bandits with last-switch dependence, and stochastic-process record statistics, although several of those works use analogous quantities rather than the exact term “TSLR” [2410.08193], [2208.04760], [2406.12284], [2511.12808], [2110.11819], [1811.00827].

## 1. Formal definition and state dynamics

In the explicit formulation of TSLR, the quantity is defined per arm. For each arm \(n \in \{1,\dots,N\}\) and round \(t\), the TSLR is denoted by \(Z_n(t)\), where \(S_n(t)\in\{0,1\}\) indicates whether the arm is pulled and \(X_n(t)\in\{0,1\}\) is the realized Bernoulli reward with unknown mean \(\mu_n\). Its dynamics are
\[
Z_n(t+1)=
\begin{cases}
Z_n(t)+1, & \text{if } S_n(t)X_n(t)=0,\\
1, & \text{if } S_n(t)X_n(t)=1.
\end{cases}
\]
Thus, if an arm is not rewarded in round \(t\)—either because it was not pulled or because it was pulled and failed—its age increases by one; if it receives a successful reward, the counter resets to \(1\) rather than \(0\) [2509.12457].

This reset convention is a recurrent source of confusion. The paper notes that one could also use \(0\), but the analysis is written with \(Z_n(0)=0\) and reset-on-success to \(1\). TSLR is also strictly per arm: arms may be activated in combinatorial subsets, but no separate TSLR is defined for a subset or joint action. Regularity is assessed at the arm level and then aggregated across arms [2509.12457].

The same paper makes clear that TSLR is not merely a bookkeeping device. It is a state variable whose dynamics are intentionally parallel to age processes in scheduling and queuing. That interpretation matters because it determines both the objective being optimized and the analytical tools used to study the algorithmic consequences of reward recency [2509.12457].

## 2. Reward regularity, age-of-information, and what TSLR measures

TSLR is explicitly connected to age-of-information (AoI): it is “essentially the same as the time since the last service” and “age of information,” but with the notion of service replaced by receipt of reward. The principal regularity metric is the running average of total expected TSLR,
\[
\frac{1}{T}\sum_{t=0}^{T-1}\sum_{n=1}^{N} E[Z_n(t)],
\]
and the paper states that smaller TSLR implies more regular reward arrivals [2509.12457].

This interpretation is sharpened by a renewal-theoretic identity: if \(\tau_n\) denotes the inter-reward time of arm \(n\), then
\[
E[Z_n] \propto \frac{E[\tau_n^2]}{2E[\tau_n]}.
\]
Accordingly, controlling mean TSLR controls not only average inter-reward time but also its variability. TSLR therefore measures short-term temporal regularity rather than cumulative reward alone [2509.12457].

That distinction separates TSLR objectives from ordinary regret minimization. A standard CMAB algorithm seeks to maximize
\[
\sum_{t=0}^{T-1}E[R(t)] = \sum_{t=0}^{T-1}\sum_{n=1}^{N}E[X_n(t)S_n(t)],
\]
which can produce starvation of low-mean arms and bursty reward sequences even when fairness constraints require minimum average service. TSLR penalizes such irregularity because long gaps without reward cause \(Z_n(t)\) to grow, which in turn increases the priority of the neglected arm. A common misconception is therefore to treat TSLR as a proxy for reward magnitude; in the formulation above, it is a proxy for reward spacing and regularity [2509.12457].

## 3. TSLR in regular-and-fair combinatorial bandits

The most direct algorithmic use of TSLR appears in the Regular and Fair Learning (RFL) algorithm. RFL combines three signals: virtual queues \(Q_n(t)\) for fairness, TSLR \(Z_n(t)\) for regularity, and UCB weights \(w_n(t)\) for exploration and exploitation. The per-arm UCB term is defined from the pull count
\[
H_n(t)\triangleq \sum_{\tau=0}^{t-1} S_n(\tau),\qquad H_n(0)=0,
\]
the sample mean
\[
\mu_n(t)\triangleq
\begin{cases}
\frac{\sum_{\tau=0}^{t-1} X_n(\tau)S_n(\tau)}{H_n(t)}, & H_n(t)>0,\\
1, & H_n(t)=0,
\end{cases}
\]
and the truncated confidence bound
\[
w_n(t)\triangleq
\begin{cases}
\min\left\{\mu_n(t)+\sqrt{\frac{3\log t}{2H_n(t)}},\,1\right\}, & H_n(t)>0,\\
1, & H_n(t)=0.
\end{cases}
\]
Fairness is encoded by virtual queues
\[
Q_n(t+1)=\big(Q_n(t)+\lambda_n-S_n(t)X_n(t)+\epsilon\big)^+,\qquad Q_n(0)=0,
\]
with \(\epsilon\in(0,1)\) chosen so that \(\lambda_n+\epsilon<\mu_n\). RFL then selects the feasible activation vector solving
\[
S(t)\in \arg\max_{S\in\mathcal S}\sum_{n=1}^{N}\big(Q_n(t)+\alpha Z_n(t)+\beta w_n(t)\big)S_n,
\]
where \(\alpha\ge 0\) and \(\beta\ge 0\) tune the regularity–regret tradeoff [2509.12457].

The key structural lemma is
\[
1+Q_n(t)\ge \lambda_n Z_n(t),\qquad \forall t\ge 0,
\]
assuming \(Q_n(0)=Z_n(0)=0\). This sample-path inequality links fairness debt to reward age: large TSLR necessarily implies a large queue up to scaling. The paper uses this connection inside Lyapunov drift arguments with
\[
V_1(t)=\sum_{n=1}^{N}\frac{Q_n^2(t)}{\mu_n}+4\alpha\sum_{n=1}^{N}\frac{Z_n(t)}{\mu_n}
\]
to obtain finite-time fairness and regularity guarantees [2509.12457].

The resulting performance statements make the tradeoff explicit. Proposition 1 states that, under \(\epsilon\le \delta/2\), there exists
\[
t_0=\frac{g_0(\alpha,\beta)}{\epsilon}=O\!\left(\frac{\alpha^2\log\alpha+\beta}{\epsilon}\right)
\]
after which zero cumulative fairness violation holds. Proposition 2 bounds the running average of total expected TSLR by the minimum of a queue-based term and an intrinsic regularity term; the paper summarizes the second as
\[
\text{Regularity} \approx O\!\left(\frac{N^2}{\alpha}\right)\quad\text{when } \alpha\gg \beta,
\]
so increasing \(\alpha\) improves regularity and increasing \(\beta\) worsens it. Proposition 3 gives a cumulative regret bound
\[
\text{Reg}(T)\le \min\left\{S_{\max}\mu_{\max}T,\; \frac{NT}{\mu_{\min}\left(\frac{\alpha+1}{\beta}\right)} + 2\sqrt{6NS_{\max}T\log T} + N\left(1+\frac{5\pi^2}{12}\right)\right\},
\]
which makes the same tradeoff visible on the regret side: larger \(\alpha\) improves reward regularity but worsens regret, while larger \(\beta\) improves regret but worsens regularity [2509.12457].

## 4. Token-level and representation-level analogues of TSLR

Several recent sequence-modeling papers introduce temporally structured quantities that are not named TSLR but are directly analogous to it.

In GenARM, the original paper does not define TSLR explicitly, but the autoregressive reward model (ARM) yields a natural token-level reward process because
\[
r(x,y)=\log \pi_r(y\mid x)=\sum_t \log \pi_r(y_t\mid x,y_{<t}).
\]
The supplied formulation therefore defines a thresholded reward event
\[
\mathbb I_t=\mathbf 1\{R_t\ge \tau\},\qquad R_t=\log \pi_r(y_t\mid x,y_{<t}),
\]
and a token-level TSLR
\[
d_t=
\begin{cases}
0, & \text{if } \mathbb I_t=1,\\
d_{t-1}+1, & \text{if } \mathbb I_t=0,
\end{cases}
\qquad d_0=0.
\]
This quantity can be used to modulate reward guidance through a dynamic weight \(\lambda_t=f(d_t)\), yielding
\[
\pi^{\text{GenARM-TSLR}}(y_t\mid x,y_{<t}) \propto \pi_{\text{base}}(y_t\mid x,y_{<t})\cdot \big(\pi_r(y_t\mid x,y_{<t})\big)^{\lambda_t/\beta}.
\]
The proposal is explicitly marked as a natural extension rather than a definition from the original GenARM paper, but it shows how dense token-level reward signals make “time since last meaningful reward” operational at decoding time [2410.08193].

TLSRec provides a different but structurally similar mechanism. Its central temporal quantity is the lag
\[
\Delta t=t_{\text{now}}-t_{\text{last}},
\]
discretized as
\[
\delta=\min\left(\left\lceil\frac{\Delta t}{\Delta_{\min}}\right\rceil,\, C\right),
\]
embedded by
\[
\mathbf y=\mathbf Y\mathbf \delta,
\]
and fed to a neural time gate
\[
\mathbf g=\text{sigmoid}\big(\mathbf W_l \mathbf z^u_{\text{long}}+\mathbf W_s \mathbf z^u_{\text{short}}+\mathbf W_\delta \mathbf y+\mathbf b_g\big).
\]
The final fusion is
\[
\mathbf z_u=\mathbf g\otimes \mathbf z^u_{\text{short}}+(\mathbf 1-\mathbf g)\otimes \mathbf z^u_{\text{long}}.
\]
The paper is about time since last interaction rather than reward, but it explicitly states that in a TSLR setting the same machinery can be applied by replacing the last interaction with the last reward event. The case study further reports that the average gate value decreases with lag, so larger lag shifts weight from short-term to long-term information [2208.04760].

| Setting | Temporal quantity | Functional role |
|---|---|---|
| GenARM | \(d_t\) | Adaptive reward temperature in decoding |
| TLSRec | \(\Delta t,\delta\) | Time-lag-sensitive fusion of long/short-term preferences |

A plausible implication is that TSLR can serve either as a hard counter with reset-and-increment dynamics, as in bandits, or as a learned latent control signal, as in neural sequence models. The distinction is methodological rather than conceptual: in both cases, elapsed time since a reward-relevant event modulates present decisions.

## 5. Credit assignment and non-Markovian reward specification

TSLR also has a direct interpretation as a temporal credit-assignment kernel. In the generalized TD return
\[
\hat{G}_t = V_t + \sum_{i=0}^{\infty} h_i \gamma^i \delta_{t+i},
\]
the weight sequence \(h_i\) measures how much a future TD error \(\delta_{t+i}\) credits an earlier stimulus at time \(t\). The weak recency heuristic requires
\[
h_{i-1}\ge h_i\ge 0 \quad \text{for all } i\ge 1,
\]
and the strong recency heuristic requires
\[
h_{i-1}>h_i>0 \quad \text{for all } i\ge 1.
\]
TD\((\lambda)\) is the canonical strong-recency case because
\[
G_t^\lambda = V_t + \sum_{i=0}^{\infty} (\gamma\lambda)^i \delta_{t+i},
\]
so the influence of a TD error decays exponentially with the time lag between the state and the error. The paper proves that any return estimator satisfying the weak recency heuristic is exactly a convex combination of \(n\)-step returns, hence yields a contraction mapping with the correct fixed point in the on-policy tabular setting; it also provides a counterexample where a delayed, non-monotone weighting diverges [2406.12284].

This places TSLR-like constructions inside a broader theory of admissible temporal weighting. Monotone nonnegative lag-weighting is theoretically safe; non-monotone weighting may be intuitive in delayed-reward settings, but it can destroy the contraction property. The same paper shows that long-tailed convex mixtures can retain a long window of effective credit assignment while keeping worst-case variance bounded through the contraction modulus \(\beta\) [2406.12284].

A complementary line of work approaches the issue from temporal logic rather than TD weighting. Quantitative reward monitors synthesized from \(\text{LTL}_f[\mathcal F]\) maintain real-valued registers over finite traces and emit dense stepwise rewards for non-Markovian specifications. The paper does not define TSLR explicitly, but it states that TSLR is a particular kind of history-dependent reward mechanism: a signal that depends on how long ago the agent last received a reward. Its monitor framework naturally supports such constructions because rewards are functions of monitor state and registers, and the product MDP \(\mathcal M\otimes \mathcal A\) converts non-Markovian reward objectives into Markovian control over an extended state space. The paper therefore provides a principled route to implementing TSLR-like reward shaping as a quantitative monitor with a reset-on-event register and a reward function decreasing in elapsed time since that event [2511.12808].

## 6. Generalizations, analogues, and open tradeoffs

Beyond the explicit CMAB formulation, TSLR-like quantities appear in several adjacent settings. In Last Switch Dependent (LSD) bandits, each arm has a signed state \(\tau_a(t)\in\mathbb Z\): positive values encode time since the arm was last played, while negative values encode the length of the current play streak. The expected reward when playing arm \(a\) at time \(t\) is
\[
r_t=\mu_{a_t}(\tau_{a_t}(t)).
\]
The positive branch behaves like a time-since-last-play variable, while the negative branch captures satiation under repeated use. The paper proves that computing the optimal policy is NP-hard, proposes the block-based ISI-CombUCB1 algorithm, and obtains regret bounds of order \(\tilde O(KT^{3/4})\) by balancing approximation and estimation. This suggests that TSLR-like dependence can be generalized from reward regularity to novelty, recovery, satiation, and seasonality, but at a substantial algorithmic cost [2110.11819].

An even broader analogue appears in stochastic-process record statistics. For a process \((X_t)_{t\ge 0}\) on \([0,T]\), the drawdown time
\[
\tau=\sup_{t\le T}\{T-t:\, M_T=M_t\},
\qquad M_t=\max_{0\le t'\le t} X_{t'}
\]
is the time since the process last achieved its running maximum. The paper explicitly identifies this as a TSLR-type quantity. For Brownian motion with drift and for completely asymmetric Lévy processes, it derives exact density formulas and factorization results of the form
\[
f_\tau(t)=\text{(function of }t)\times \text{(function of }T-t).
\]
This shows that “time since last event” variables can be studied independently of control or learning, as intrinsic functionals of path-dependent stochastic dynamics [1811.00827].

Across these literatures, the main open issues are tradeoff and structure rather than definition. In the CMAB setting, the exact optimality of the regret–regularity tradeoff remains open, and low-complexity variants of RFL are identified as future work. In LSD bandits, the optimal cyclic-policy problem is itself NP-hard, and the structure of optimal non-stationary policies is unresolved. In TSLR-aware neural and monitor-based systems, the supplied formulations suggest several natural constructions, but in GenARM and quantitative reward monitoring they remain extensions of the underlying framework rather than original named definitions [2509.12457], [2110.11819], [2410.08193], [2511.12808].

Taken together, these results support a unified view. TSLR is a temporally local summary of reward history whose formal role depends on the problem class: an age process for regularity in bandits, a lag-weighting kernel in TD learning, a latent gating signal in sequence models, a monitor register in non-Markovian RL, a signed recovery-or-satiation state in non-stationary bandits, or a last-record functional in stochastic processes. The common invariant is the same: elapsed time since the most recent reward-relevant event is treated as a state variable with algorithmic, statistical, or probabilistic consequences.

Source: https://www.emergentmind.com/topics/time-since-last-reward-tslr