- The paper introduces a four-stage Explore, Compare, Distill, and Reinject framework that learns from execution quality rather than completion alone while keeping the base model frozen.
- TimeClaw achieves the best result on 27 of 30 metrics across finance and weather tasks, including improvements in weather forecasting, trend prediction, and MACD analysis.
- The paper shows that task-aware tool dropout and hierarchical, conflict-aware memory reduce tool-prior collapse and improve experience reuse, although long-horizon finance reasoning remains a weakness.
TimeClaw is a framework for time-series LLM agents that shifts the learning objective from solving individual task instances to accumulating reusable experience from exploratory execution. The authors identify two failure modes of the prevailing execution-centric paradigm: completion-only supervision, which treats all task-valid executions as equivalent despite large differences in quantitative quality, and tool-prior collapse, in which an agent prematurely concentrates tool usage on a narrow subset during exploration, degrading the informativeness of subsequent comparisons. TimeClaw addresses both through a four-stage loop—Explore, Compare, Distill, Reinject—that keeps the base model frozen and requires no online test-time adaptation (2605.10038).
Each task instance is x=(z,c,τ,s), where z is the time-series input, c optional multimodal context, τ the task type, and s the scope under which experience is stored. An agent produces candidate executions Π(x)={π1,…,πK}; each is scored by q(πk;x)=−Lτ(y^k,y⋆) under task-dependent losses such as MAE or RMSE, and only task-valid candidates compete for selection as π⋆(x). Because agents cannot backpropagate through discrete tool calls, learning proceeds by explicit comparison among candidates—what the authors call execution-quality supervision. Episodes with zero or one valid candidate are not discarded: they contribute single-execution or failure evidence with weaker signal.
Method
Exploration-time learning. A constrained executor generates at least two distinct, task-valid candidate executions per instance using exploration-only orchestration tools (spawn_subagent, evaluate_* against ground truth), then compares them before terminating with a learning_summary. The summary records the selected strategy, decisive evidence, and failure signals; it is cleaned to remove answer leakage and non-transferable details before updating the task-local experience state. Intermediate tool outputs are tracked as typed artifacts with coordinate provenance so that downstream operations condition on transformed series correctly.
Task-aware tool dropout. To counteract tool-prior collapse, TimeClaw assigns each non-protected tool a keep probability anchored at the least-explored tool:
pkeep(i∣s)=(1+ni(s)1+nmin(s))α,
with protected (task-critical or hinted) tools exempted. A monotonicity proposition guarantees that historically dominant tools are suppressed more strongly than less-used ones. The authors distinguish this exploration-phase pathology from legitimate inference-time tool preference learned from evidence—a distinction prior time-series agent work has not operationalized.
Hierarchical distilled experience. Learning is externalized into five layers: Soul (fixed behavioral anchor), Notes (append-only sample-level distillation source), Memory (structured 9-tuple rules with kind, applicability conditions, preferred/avoided tool sets, rationale, evidence, confidence, and injection status), Tools (per-tool usage advice), and Skills (task-local SOPs). Updates are conflict-aware rather than append-only: rules may be merged, strengthened, or registered as conflicting, and unstable rules can be stored but marked non-injectable. At inference, retrieval filters memory by task, applicability, and injection status; exploration-only tools are removed, and the frozen agent solves tasks under prompts enriched with distilled experience.
Experimental results
Evaluation uses a fixed 17-task MTBench suite spanning finance and weather, aggregated into 30 metrics, with baselines including GPT-4o, Gemini, Claude, DeepSeek, GPT-5.4, Qwen-3, and Llama-3 variants. Exploration runs on a disjoint, schema-aligned corpus with source-level disjointness checks (non-overlapping finance URLs and weather station IDs), so gains reflect transfer rather than memorization.
The headline claim is strong: TimeClaw achieves the best result on 27 of 30 metrics, including 9/10 core numeric prediction metrics, 6/6 indicator metrics, and 12/14 reasoning metrics. Representative numbers:
| Task |
Metric |
Best baseline |
TimeClaw |
| Weather forecast 7D |
MSE ↓ |
14.197 (GPT-5.4) |
10.719 |
| Weather forecast 14D |
MSE ↓ |
20.984 (GPT-5.4) |
13.525 |
| Finance trend 7D |
Acc.-5 ↑ |
0.569 (GPT-5.4) |
0.696 |
| Weather past trend |
Acc. ↑ |
0.980 (GPT-5.4) |
0.994 |
| Finance MACD 7D |
MSE ↓ |
0.352 (DeepSeek) |
0.201 |
The main weakness is long-horizon finance correlation, where Gemini outperforms TimeClaw on both label granularities; 30-day MACD MSE also favors Llama-3. These exceptions indicate that distilled experience does not uniformly transfer to long-horizon relational reasoning.
Ablations separate confounded effects. Replacing GPT-5.4 with the same backbone without experience injection ("w/o Exp.") improves weather short-horizon MSE from 14.197 to 13.058, isolating the value of agent runtime and tool access; adding full TimeClaw further reduces it to 10.719, isolating the contribution of learned experience. Removing higher-level distilled layers while retaining tool routines ("w/o Mem.") consistently falls between the baseline and full system, supporting the claim that hierarchical organization—not merely retention—drives retrieval quality. Tool-dropout analysis shows reduced Top-5 tool concentration and increased early coverage across four tasks, consistent with mitigation of premature convergence. A qualitative case study on post-shock MACD forecasting shows a 5.5× error reduction after exploration, driven by a learned rule to treat news as directional bias only and let post-shock price structure determine the trajectory.
Limitations and open questions
The paper concedes several constraints plainly. Empirical claims are limited to benchmarked, verifiable finance and weather tasks with the given tool interfaces; results depend on a frozen base LLM, an externally served API stack, and a predefined tool library. The protocol is offline rather than continual, with a fixed budget of 300 exploration samples per task, and sensitivity to exploration budget, domain shift beyond MTBench-style settings, and deployment under changing distributions remain unexamined. Exploration is expensive: a full run takes roughly one day on two RTX 3090 GPUs plus substantial token expenditure, scaling linearly with tool-library size. No confidence intervals or significance tests are reported, so close comparisons should be read cautiously. Open questions include whether metric-supervised comparison transfers to tasks lacking verifiable ground truth, and how conflict-aware memory behaves over much longer horizons than 30 rules per task.
Conclusion
TimeClaw reframes time-series agent learning around exploratory execution: comparing task-valid candidates under verifiable metrics, distilling evidence into hierarchical external experience, and reinjecting it at inference without parameter updates. Its consistent advantage over strong closed-model baselines—best on 27/30 metrics—supports the thesis that for scientific agents the bottleneck lies in how exploratory experience is compared, organized, and reused, not solely in execution-time capability.