Papers
Topics
Authors
Recent
Search
2000 character limit reached

SATER: Self-Aware Token-Efficient Routing

Updated 14 July 2026
  • SATER is a self-aware, token-efficient framework that employs a small language model to compress outputs and dynamically route queries based on confidence levels.
  • It features a two-stage training process combining shortest-response preference optimization with confidence-aware refusal, significantly reducing token count and cascade latency.
  • The framework demonstrates robust performance across multiple benchmarks while achieving over 50% cost savings and over 80% reduction in latency compared to conventional methods.

SATER, short for “A Self-Aware and Token-Efficient Approach to Routing and Cascading,” is a framework for LLM deployment that addresses the trade-off between performance and cost by training a small LLM (SLM) to both compress its outputs and explicitly refuse queries it is unlikely to solve. It is designed to be compatible with the two dominant routing paradigms in the literature—pre-generation routing and cascade routing—and is presented as a dual-mode approach in which the SLM acts not only as a generator but also as a capability-aware routing component (Shen et al., 4 Oct 2025).

1. Routing problem and formal objective

SATER is motivated by a central systems problem in LLM inference: whether a query should be answered by a budget-friendly SLM or escalated to a stronger LLM. The paper frames this as a per-query routing decision governed by an SLM-predicted confidence score sis_i and a routing threshold τ∈[0,1]\tau \in [0,1]: r(i)={1,if si<τ (routed to Ml) 0,if si≥τ (routed to Ms)r(i) = \begin{cases} 1, & \text{if } s_i < \tau \text{ (routed to } M_l\text{)} \ 0, & \text{if } s_i \geq \tau \text{ (routed to } M_s\text{)} \end{cases} A larger τ\tau sends more queries to the LLM, improving quality but increasing cost (Shen et al., 4 Oct 2025).

The paper distinguishes two routing paradigms. In pre-generation routing, the system decides before generation whether the SLM or LLM should answer. In cascade routing, the SLM generates first, and the system falls back to the LLM if the SLM output is low-confidence or explicitly rejected. SATER is intended to improve both settings by reducing redundant outputs and avoiding wasted generation on queries that will ultimately be escalated.

The cost model is token-level. For question ii, with input length tiint_i^{\text{in}}, SLM output length tist_i^s, and LLM output length tilt_i^l, the paper defines

Cis=csin×tiin+csout×tisC_i^s = c_s^{\text{in}} \times t_i^{\text{in}} + c_s^{\text{out}} \times t_i^s

and

Cil=clin×tiin+clout×til.C_i^l = c_l^{\text{in}} \times t_i^{\text{in}} + c_l^{\text{out}} \times t_i^l.

The normalized total cost for pre-generation routing is

τ∈[0,1]\tau \in [0,1]0

whereas for cascade routing it is

τ∈[0,1]\tau \in [0,1]1

with τ∈[0,1]\tau \in [0,1]2 denoting the number of samples. Average quality is defined as

τ∈[0,1]\tau \in [0,1]3

These definitions make explicit that routing quality depends not only on accuracy but also on output length, sampling multiplicity, and fallback behavior.

2. Two-stage training recipe

The core of SATER is a two-stage training recipe: shortest-response preference optimization followed by a confidence-aware rejection mechanism. The first stage is called Long to Short Training, and it combines Direct Preference Optimization (DPO) with supervised fine-tuning (SFT). The total loss is

τ∈[0,1]\tau \in [0,1]4

The DPO term is

τ∈[0,1]\tau \in [0,1]5

Here τ∈[0,1]\tau \in [0,1]6 is the input, τ∈[0,1]\tau \in [0,1]7 the preferred response, τ∈[0,1]\tau \in [0,1]8 the dispreferred response, τ∈[0,1]\tau \in [0,1]9 the policy model, r(i)={1,if si<τ (routed to Ml) 0,if si≥τ (routed to Ms)r(i) = \begin{cases} 1, & \text{if } s_i < \tau \text{ (routed to } M_l\text{)} \ 0, & \text{if } s_i \geq \tau \text{ (routed to } M_s\text{)} \end{cases}0 the reference model, r(i)={1,if si<τ (routed to Ml) 0,if si≥τ (routed to Ms)r(i) = \begin{cases} 1, & \text{if } s_i < \tau \text{ (routed to } M_l\text{)} \ 0, & \text{if } s_i \geq \tau \text{ (routed to } M_s\text{)} \end{cases}1 the temperature coefficient, and r(i)={1,if si<τ (routed to Ml) 0,if si≥τ (routed to Ms)r(i) = \begin{cases} 1, & \text{if } s_i < \tau \text{ (routed to } M_l\text{)} \ 0, & \text{if } s_i \geq \tau \text{ (routed to } M_s\text{)} \end{cases}2 the sigmoid (Shen et al., 4 Oct 2025).

Training pairs are built by sampling each question 10 times, then selecting the shortest correct response as the positive example and the longest incorrect response as the negative example, with the additional constraint that the negative response must be at least 1.5× the length of the positive sample. The stated aim is to teach the model to prefer concise correct answers and reject verbose incorrect ones. The paper reports that DPO is sensitive to length, and sets

r(i)={1,if si<τ (routed to Ml) 0,if si≥τ (routed to Ms)r(i) = \begin{cases} 1, & \text{if } s_i < \tau \text{ (routed to } M_l\text{)} \ 0, & \text{if } s_i \geq \tau \text{ (routed to } M_s\text{)} \end{cases}3

to keep training stable. It also reports that including the longest correct response as a negative example reduces average accuracy by over 2%.

The second stage is Refusal Training. After Stage I, the model again resamples each question 10 times and computes its accuracy on a 0 to 1.0 scale. Confidence thresholds

r(i)={1,if si<τ (routed to Ml) 0,if si≥τ (routed to Ms)r(i) = \begin{cases} 1, & \text{if } s_i < \tau \text{ (routed to } M_l\text{)} \ 0, & \text{if } s_i \geq \tau \text{ (routed to } M_s\text{)} \end{cases}4

are then used to construct prompt-conditioned targets. For each question and threshold, the model is given the instruction: “Please respond with a confidence level of [threshold]:” If the question’s accuracy exceeds the threshold, a correct answer is sampled; otherwise the target becomes the rejection template: “Sorry, I can't answer that.” Standard SFT is then used to train the model to condition its behavior on the requested confidence level.

The paper attributes two distinct effects to this pipeline. Stage I reduces token count by over 40% in the appendix tables and is summarized in the main narrative as reducing redundant tokens by over 50% with minimal performance degradation. Stage II makes the model “self-aware” in the routing sense by enabling explicit refusal on uncertain or difficult queries.

3. Pre-generation routing as self-aware model selection

In the pre-generation setting, SATER uses the SLM itself as the router. After refusal training, the model can emit an explicit reject response on difficult queries, and those rejected queries are routed to the LLM. The paper emphasizes that this removes the need for an external classifier-based router, because the model becomes both the generator and the routing mechanism (Shen et al., 4 Oct 2025).

This framing is important to the paper’s comparison with earlier routing methods such as HybridLLM, KNN routing, and BERT-based routing. The authors argue that traditional routers may overestimate performance on some benchmarks because of “performance gap bias.” To address this, they introduce more robust metrics: ToA, ToGR, ToA-100, and ToGA-100. The purpose of these metrics is to evaluate the cost-performance trade-off without conflating routing quality with benchmark-specific separation between easy and hard questions.

The reported result is that SATER improves ToA-100 and ToGR over baselines across all three SLMs and six datasets, with especially strong gains on harder tasks and out-of-domain benchmarks. In representative comparisons, the authors report that SATER’s ToGR is around 0.2 higher than BERT on both ReClor and ARC-Easy. This suggests that the internalized routing signal learned by refusal training captures task difficulty at a finer granularity than an external router in the examined setting.

4. Cascade routing, voting, and latency control

In cascade routing, the SLM generates first, and the system escalates to the LLM if the answer is low-confidence or refused. SATER is designed to reduce the inefficiencies of this mode by making the SLM either answer quickly or refuse early. The paper characterizes the baseline problem as one of latency polarization: easy questions are completed quickly, whereas hard questions may incur long generations, repeated sampling, and then fallback to the LLM (Shen et al., 4 Oct 2025).

To evaluate this setting, the paper introduces two latency metrics. AGL (Average Generation Latency) measures latency when the response is ultimately completed by r(i)={1,if si<τ (routed to Ml) 0,if si≥τ (routed to Ms)r(i) = \begin{cases} 1, & \text{if } s_i < \tau \text{ (routed to } M_l\text{)} \ 0, & \text{if } s_i \geq \tau \text{ (routed to } M_s\text{)} \end{cases}5. AROL (Average Routing Overhead Latency) measures the extra latency incurred when r(i)={1,if si<τ (routed to Ml) 0,if si≥τ (routed to Ms)r(i) = \begin{cases} 1, & \text{if } s_i < \tau \text{ (routed to } M_l\text{)} \ 0, & \text{if } s_i \geq \tau \text{ (routed to } M_s\text{)} \end{cases}6 fails and the system must fall back to r(i)={1,if si<τ (routed to Ml) 0,if si≥τ (routed to Ms)r(i) = \begin{cases} 1, & \text{if } s_i < \tau \text{ (routed to } M_l\text{)} \ 0, & \text{if } s_i \geq \tau \text{ (routed to } M_s\text{)} \end{cases}7. These metrics are intended to capture the concrete effect of output length and fallback behavior.

For cascade routing, SATER uses a confidence-aware dynamic weighted voting scheme. For question r(i)={1,if si<τ (routed to Ml) 0,if si≥τ (routed to Ms)r(i) = \begin{cases} 1, & \text{if } s_i < \tau \text{ (routed to } M_l\text{)} \ 0, & \text{if } s_i \geq \tau \text{ (routed to } M_s\text{)} \end{cases}8, with r(i)={1,if si<τ (routed to Ml) 0,if si≥τ (routed to Ms)r(i) = \begin{cases} 1, & \text{if } s_i < \tau \text{ (routed to } M_l\text{)} \ 0, & \text{if } s_i \geq \tau \text{ (routed to } M_s\text{)} \end{cases}9 samples and candidate answers τ\tau0, each sample τ\tau1 has a discretized confidence τ\tau2, and the weight is

τ\tau3

The score for answer τ\tau4 is

τ\tau5

and the answer with the highest score is selected. The paper defines two voting variants. RCV (Ranged Confidence Voting) samples confidences uniformly from 0.1 to 1.0 for 10 times, whereas FCV (Fixed Confidence Voting) samples only at confidence 1.0 for 10 times. RCV is described as often better for stronger SLMs because it preserves voting diversity, while FCV is described as tending to work better for weaker SLMs because it quickly rejects uncertain queries and avoids wasted sampling.

At τ\tau6, the paper reports the following average latency figures for the vanilla cascade baseline (SC) and an FCV-based cascade variant (SC/FCV):

Model SC SC/FCV
Llama-3.1-8B-Instruct AGL 199, AROL 304 AGL 73, AROL 6
Qwen2.5-7B-Instruct AGL 260, AROL 422 AGL 150, AROL 13
Qwen2.5-3B-Instruct AGL 330, AROL 466 AGL 150, AROL 5

The same pattern is reported to persist at τ\tau7, with AROL dropping to single digits or near single digits in many cases. These figures are the empirical basis for the paper’s claim that SATER reduces cascade latency by over 80%.

5. Experimental setting and quantitative profile

The paper evaluates SATER using three SLMs—Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, and Qwen2.5-3B-Instruct—and one LLM, DeepSeek-V3-0324 (Shen et al., 4 Oct 2025). Evaluation is conducted on six benchmarks: MMLU, MATH-500, GSM8K, ARC-Challenge, ReClor, and ARC-Easy. Training uses the training sets of MMLU, ARC-Challenge, GSM8K, and MATH-500, while ARC-Easy and ReClor are treated as out-of-domain.

For pre-generation routing, the main metrics are ToA-100 and ToGR. For cascade routing, the main metrics are AGL and AROL. Cost-accuracy curves are compared under several token price ratios: 1:13.75 as the default, plus 1:25, 1:50, and 1:100. The default ratio is derived from pricing assumptions in which SLM output tokens cost $\tau$81.10, with input prices set at one-quarter of output prices.

The headline result is that SATER achieves comparable performance while consistently reducing computational cost by over 50% and cascade latency by over 80%. The appendix further reports substantial reductions in average chain-of-thought token count: for Llama-3.1-8B-Instruct, average tokens drop from 202.5 to 93.3; for Qwen2.5-7B-Instruct, from 254.5 to 170.5; and for Qwen2.5-3B-Instruct, from 319.8 to 195.0. The accompanying accuracy changes are described as modest, usually around 1–3 points.

The paper also compares SATER against FrugalGPT, Automix, Margin Sampling, the vanilla cascade baseline (SC), and an intermediate version trained only with Stage I (SC/TE). The comparison is meant to isolate the contributions of token compression, refusal behavior, and confidence-aware voting.

6. Ablations, threshold effects, and limitations

Several ablations are central to the paper’s interpretation of why SATER works. The SC/TE baseline, which isolates Stage I, shows that shortest-response preference optimization alone already provides major latency savings, whereas full SATER improves both routing quality and cascade efficiency through the addition of refusal training and voting (Shen et al., 4 Oct 2025).

Threshold sensitivity is task dependent. The appendix notes that for mathematical reasoning tasks, $\tau$9 usually works well, whereas other tasks may benefit from thresholds in the range 0.6 to 1.0. Very low thresholds are harder to use effectively, and the paper states that refusal training struggles below 0.1. Cost-ratio analysis further differentiates routing regimes: at low cost ratios, pre-generation routing can be more cost-effective because cascade sampling is expensive; as the LLM becomes relatively more expensive, cascade routing becomes more attractive. The paper states that at 1:25, RCV and FCV outperform pre-generation routing; at 1:50, SC/TE matches BERT; and at 1:100, it surpasses it.

The paper also notes a practical limitation of refusal training. It may not be ideal when a response is mandatory but routing is impossible; in such cases, one might need to keep an untrained fallback copy. This limitation follows directly from SATER’s design choice: over-rejection is treated differently in routing systems than in standalone LLM use, because in the routing setting every query is still expected to receive an answer from some model in the system.

Taken together, these analyses position SATER as a routing and cascading framework in which the SLM is trained to identify its own competence boundary, generate more compact outputs, and stop early when escalation is preferable. Within the experimental scope reported in the paper, the combination of shortest-response preference optimization and confidence-aware refusal is the mechanism by which SATER improves both pre-generation routing and cascade routing under explicit cost and latency objectives.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to SATER.