---
title: 'SATER: Self-Aware Token-Efficient Routing'
url: https://www.emergentmind.com/topics/sater
type: topic
---

# SATER: Self-Aware Token-Efficient Routing

SATER, short for **“A Self-Aware and Token-Efficient Approach to Routing and Cascading,”** is a framework for large language model deployment that addresses the trade-off between performance and cost by training a small language model (SLM) to both compress its outputs and explicitly refuse queries it is unlikely to solve. It is designed to be compatible with the two dominant routing paradigms in the literature—**pre-generation routing** and **cascade routing**—and is presented as a dual-mode approach in which the SLM acts not only as a generator but also as a capability-aware routing component [2510.05164].

## 1. Routing problem and formal objective

SATER is motivated by a central systems problem in LLM inference: whether a query should be answered by a budget-friendly SLM or escalated to a stronger large language model (LLM). The paper frames this as a per-query routing decision governed by an SLM-predicted confidence score \(s_i\) and a routing threshold \(\tau \in [0,1]\):
\[
r(i) =
\begin{cases}
1, & \text{if } s_i < \tau \text{ (routed to } M_l\text{)} \\
0, & \text{if } s_i \geq \tau \text{ (routed to } M_s\text{)}
\end{cases}
\]
A larger \(\tau\) sends more queries to the LLM, improving quality but increasing cost [2510.05164].

The paper distinguishes two routing paradigms. In **pre-generation routing**, the system decides before generation whether the SLM or LLM should answer. In **cascade routing**, the SLM generates first, and the system falls back to the LLM if the SLM output is low-confidence or explicitly rejected. SATER is intended to improve both settings by reducing redundant outputs and avoiding wasted generation on queries that will ultimately be escalated.

The cost model is token-level. For question \(i\), with input length \(t_i^{\text{in}}\), SLM output length \(t_i^s\), and LLM output length \(t_i^l\), the paper defines
\[
C_i^s = c_s^{\text{in}} \times t_i^{\text{in}} + c_s^{\text{out}} \times t_i^s
\]
and
\[
C_i^l = c_l^{\text{in}} \times t_i^{\text{in}} + c_l^{\text{out}} \times t_i^l.
\]
The normalized total cost for pre-generation routing is
\[
\tilde{C}_{pre} =
\frac{\sum_{i=1}^N \left[(1-r(i))C_i^s + r(i)C_i^l\right]}
{\sum_{i=1}^N C_i^l},
\]
whereas for cascade routing it is
\[
\tilde{C}_{cascade} =
\frac{\sum_{i=1}^N \left[K \cdot C_i^s + r(i)\cdot C_i^l\right]}
{\sum_{i=1}^N C_i^l},
\]
with \(K\) denoting the number of samples. Average quality is defined as
\[
\tilde{P} = \frac{1}{N}\sum_{i=1}^N \left[(1-r(i))p_i^s + r(i)p_i^l\right].
\]
These definitions make explicit that routing quality depends not only on accuracy but also on output length, sampling multiplicity, and fallback behavior.

## 2. Two-stage training recipe

The core of SATER is a two-stage training recipe: **shortest-response preference optimization** followed by a **confidence-aware rejection mechanism**. The first stage is called **Long to Short Training**, and it combines **Direct Preference Optimization (DPO)** with supervised fine-tuning (SFT). The total loss is
\[
\mathcal{L}_{\text{Total}} = \mathcal{L}_{\text{DPO}} + \lambda \mathcal{L}_{\text{SFT}}.
\]
The DPO term is
\[
-\mathbb{E}_{(x,y_w,y_l)\sim\mathcal{D}}
\left[
\log \sigma \left(
\beta \log \frac{\pi_\theta(y_w\mid x)}{\pi_{\text{ref}}(y_w\mid x)}
-
\beta \log \frac{\pi_\theta(y_l\mid x)}{\pi_{\text{ref}}(y_l\mid x)}
\right)
\right].
\]
Here \(x\) is the input, \(y_w\) the preferred response, \(y_l\) the dispreferred response, \(\pi_\theta\) the policy model, \(\pi_{\text{ref}}\) the reference model, \(\beta\) the temperature coefficient, and \(\sigma\) the sigmoid [2510.05164].

Training pairs are built by sampling each question **10 times**, then selecting the **shortest correct response** as the positive example and the **longest incorrect response** as the negative example, with the additional constraint that the negative response must be at least **1.5×** the length of the positive sample. The stated aim is to teach the model to prefer concise correct answers and reject verbose incorrect ones. The paper reports that DPO is sensitive to length, and sets
\[
\beta = 1,\quad \lambda = 0.2
\]
to keep training stable. It also reports that including the longest correct response as a negative example reduces average accuracy by over **2%**.

The second stage is **Refusal Training**. After Stage I, the model again resamples each question **10 times** and computes its accuracy on a **0 to 1.0** scale. Confidence thresholds
\[
0.1, 0.2, \dots, 1.0
\]
are then used to construct prompt-conditioned targets. For each question and threshold, the model is given the instruction: **“Please respond with a confidence level of [threshold]:”** If the question’s accuracy exceeds the threshold, a correct answer is sampled; otherwise the target becomes the rejection template: **“Sorry, I can't answer that.”** Standard SFT is then used to train the model to condition its behavior on the requested confidence level.

The paper attributes two distinct effects to this pipeline. Stage I reduces token count by **over 40%** in the appendix tables and is summarized in the main narrative as reducing redundant tokens by **over 50% with minimal performance degradation**. Stage II makes the model “self-aware” in the routing sense by enabling explicit refusal on uncertain or difficult queries.

## 3. Pre-generation routing as self-aware model selection

In the pre-generation setting, SATER uses the SLM itself as the router. After refusal training, the model can emit an explicit reject response on difficult queries, and those rejected queries are routed to the LLM. The paper emphasizes that this removes the need for an external classifier-based router, because the model becomes both the generator and the routing mechanism [2510.05164].

This framing is important to the paper’s comparison with earlier routing methods such as **HybridLLM**, **KNN routing**, and **BERT-based routing**. The authors argue that traditional routers may overestimate performance on some benchmarks because of **“performance gap bias.”** To address this, they introduce more robust metrics: **ToA**, **ToGR**, **ToA-100**, and **ToGA-100**. The purpose of these metrics is to evaluate the cost-performance trade-off without conflating routing quality with benchmark-specific separation between easy and hard questions.

The reported result is that SATER improves **ToA-100** and **ToGR** over baselines across **all three SLMs and six datasets**, with especially strong gains on harder tasks and out-of-domain benchmarks. In representative comparisons, the authors report that SATER’s **ToGR** is around **0.2 higher than BERT** on both **ReClor** and **ARC-Easy**. This suggests that the internalized routing signal learned by refusal training captures task difficulty at a finer granularity than an external router in the examined setting.

## 4. Cascade routing, voting, and latency control

In cascade routing, the SLM generates first, and the system escalates to the LLM if the answer is low-confidence or refused. SATER is designed to reduce the inefficiencies of this mode by making the SLM either answer quickly or refuse early. The paper characterizes the baseline problem as one of latency polarization: easy questions are completed quickly, whereas hard questions may incur long generations, repeated sampling, and then fallback to the LLM [2510.05164].

To evaluate this setting, the paper introduces two latency metrics. **AGL (Average Generation Latency)** measures latency when the response is ultimately completed by \(M_s\). **AROL (Average Routing Overhead Latency)** measures the extra latency incurred when \(M_s\) fails and the system must fall back to \(M_l\). These metrics are intended to capture the concrete effect of output length and fallback behavior.

For cascade routing, SATER uses a confidence-aware dynamic weighted voting scheme. For question \(i\), with \(K\) samples and candidate answers \(A=\{A_1,\dots,A_M\}\), each sample \(k\) has a discretized confidence \(p_k \in \{0.1,0.2,\dots,1.0\}\), and the weight is
\[
w_k = 0.55 + \alpha(p_k - 0.55),
\qquad \alpha = 0.5.
\]
The score for answer \(A_m\) is
\[
\delta(A_m) =
\frac{\sum_{k=1}^K w_k \cdot \mathbb{I}(a_k = A_m)}
{\sum_{k=1}^K w_k},
\]
and the answer with the highest score is selected. The paper defines two voting variants. **RCV (Ranged Confidence Voting)** samples confidences uniformly from **0.1 to 1.0** for 10 times, whereas **FCV (Fixed Confidence Voting)** samples only at confidence **1.0** for 10 times. RCV is described as often better for stronger SLMs because it preserves voting diversity, while FCV is described as tending to work better for weaker SLMs because it quickly rejects uncertain queries and avoids wasted sampling.

At \(\tau = 0.6\), the paper reports the following average latency figures for the vanilla cascade baseline (**SC**) and an FCV-based cascade variant (**SC/FCV**):

| Model | SC | SC/FCV |
|---|---|---|
| Llama-3.1-8B-Instruct | AGL 199, AROL 304 | AGL 73, AROL 6 |
| Qwen2.5-7B-Instruct | AGL 260, AROL 422 | AGL 150, AROL 13 |
| Qwen2.5-3B-Instruct | AGL 330, AROL 466 | AGL 150, AROL 5 |

The same pattern is reported to persist at \(\tau = 1.0\), with AROL dropping to single digits or near single digits in many cases. These figures are the empirical basis for the paper’s claim that SATER reduces cascade latency by **over 80%**.

## 5. Experimental setting and quantitative profile

The paper evaluates SATER using **three SLMs**—**Llama-3.1-8B-Instruct**, **Qwen2.5-7B-Instruct**, and **Qwen2.5-3B-Instruct**—and one LLM, **DeepSeek-V3-0324** [2510.05164]. Evaluation is conducted on **six benchmarks**: **MMLU**, **MATH-500**, **GSM8K**, **ARC-Challenge**, **ReClor**, and **ARC-Easy**. Training uses the training sets of **MMLU**, **ARC-Challenge**, **GSM8K**, and **MATH-500**, while **ARC-Easy** and **ReClor** are treated as out-of-domain.

For pre-generation routing, the main metrics are **ToA-100** and **ToGR**. For cascade routing, the main metrics are **AGL** and **AROL**. Cost-accuracy curves are compared under several token price ratios: **1:13.75** as the default, plus **1:25**, **1:50**, and **1:100**. The default ratio is derived from pricing assumptions in which SLM output tokens cost **\$0.08 per million** and DeepSeek output tokens cost **\$1.10**, with input prices set at one-quarter of output prices.

The headline result is that SATER achieves **comparable performance** while consistently reducing **computational cost by over 50%** and **cascade latency by over 80%**. The appendix further reports substantial reductions in average chain-of-thought token count: for **Llama-3.1-8B-Instruct**, average tokens drop from **202.5** to **93.3**; for **Qwen2.5-7B-Instruct**, from **254.5** to **170.5**; and for **Qwen2.5-3B-Instruct**, from **319.8** to **195.0**. The accompanying accuracy changes are described as modest, usually around **1–3 points**.

The paper also compares SATER against **FrugalGPT**, **Automix**, **Margin Sampling**, the vanilla cascade baseline (**SC**), and an intermediate version trained only with Stage I (**SC/TE**). The comparison is meant to isolate the contributions of token compression, refusal behavior, and confidence-aware voting.

## 6. Ablations, threshold effects, and limitations

Several ablations are central to the paper’s interpretation of why SATER works. The **SC/TE** baseline, which isolates Stage I, shows that shortest-response preference optimization alone already provides major latency savings, whereas full SATER improves both routing quality and cascade efficiency through the addition of refusal training and voting [2510.05164].

Threshold sensitivity is task dependent. The appendix notes that for mathematical reasoning tasks, \(\tau = 0.6\) usually works well, whereas other tasks may benefit from thresholds in the range **0.6 to 1.0**. Very low thresholds are harder to use effectively, and the paper states that refusal training struggles below **0.1**. Cost-ratio analysis further differentiates routing regimes: at low cost ratios, pre-generation routing can be more cost-effective because cascade sampling is expensive; as the LLM becomes relatively more expensive, cascade routing becomes more attractive. The paper states that at **1:25**, **RCV** and **FCV** outperform pre-generation routing; at **1:50**, **SC/TE** matches **BERT**; and at **1:100**, it surpasses it.

The paper also notes a practical limitation of refusal training. It may not be ideal when a response is mandatory but routing is impossible; in such cases, one might need to keep an untrained fallback copy. This limitation follows directly from SATER’s design choice: over-rejection is treated differently in routing systems than in standalone LLM use, because in the routing setting every query is still expected to receive an answer from some model in the system.

Taken together, these analyses position SATER as a routing and cascading framework in which the SLM is trained to identify its own competence boundary, generate more compact outputs, and stop early when escalation is preferable. Within the experimental scope reported in the paper, the combination of shortest-response preference optimization and confidence-aware refusal is the mechanism by which SATER improves both pre-generation routing and cascade routing under explicit cost and latency objectives.

Source: https://www.emergentmind.com/topics/sater