Papers
Topics
Authors
Recent
Search
2000 character limit reached

TimeClaw: A Time-Series AI Agent with Exploratory Execution Learning

Published 11 May 2026 in cs.AI | (2605.10038v1)

Abstract: Time series analysis underpins forecasting, monitoring, and decision making in domains such as finance and weather, where solving a task often requires both numerical accuracy and contextual reasoning. Recent progress has moved from specialized neural predictors to approaches built on LLMs and foundation models that can reason over time series inputs and use external tools. However, most such systems remain execution-centric: they focus on solving the current instance but learn little from exploratory execution. This is especially limiting in verifiable numeric settings, where multiple candidate executions and tool-use procedures may all be task-valid yet differ sharply in quantitative quality, and where early success can trigger tool-prior collapse that suppresses further exploration. To address this limitation, we present TimeClaw, an exploratory execution learning framework that turns exploratory execution into reusable hierarchical distilled experience through a four-stage loop: Explore, Compare, Distill, and Reinject. TimeClaw combines metric-supervised exploratory execution learning, task-aware tool dropout, and hierarchical distilled experience for inference-time reinjection, while keeping the base model frozen and avoiding online test-time adaptation. In an MTBench-aligned evaluation with 17 tasks that span finance and weather prediction and reasoning tasks, TimeClaw delivers consistent gains over the baselines. These results suggest that, for scientific systems, the bottleneck is not only execution-time capability, but how exploratory experience is compared, distilled, and reused.

Summary

  • The paper introduces a four-stage Explore, Compare, Distill, and Reinject framework that learns from execution quality rather than completion alone while keeping the base model frozen.
  • TimeClaw achieves the best result on 27 of 30 metrics across finance and weather tasks, including improvements in weather forecasting, trend prediction, and MACD analysis.
  • The paper shows that task-aware tool dropout and hierarchical, conflict-aware memory reduce tool-prior collapse and improve experience reuse, although long-horizon finance reasoning remains a weakness.

TimeClaw is a framework for time-series LLM agents that shifts the learning objective from solving individual task instances to accumulating reusable experience from exploratory execution. The authors identify two failure modes of the prevailing execution-centric paradigm: completion-only supervision, which treats all task-valid executions as equivalent despite large differences in quantitative quality, and tool-prior collapse, in which an agent prematurely concentrates tool usage on a narrow subset during exploration, degrading the informativeness of subsequent comparisons. TimeClaw addresses both through a four-stage loop—Explore, Compare, Distill, Reinject—that keeps the base model frozen and requires no online test-time adaptation (2605.10038).

Problem formulation

Each task instance is x=(z,c,τ,s)x=(z,c,\tau,s), where zz is the time-series input, cc optional multimodal context, τ\tau the task type, and ss the scope under which experience is stored. An agent produces candidate executions Π(x)={π1,…,πK}\Pi(x)=\{\pi_1,\dots,\pi_K\}; each is scored by q(πk;x)=−Lτ(y^k,y⋆)q(\pi_k;x)=-\mathcal{L}_\tau(\hat{y}_k,y^\star) under task-dependent losses such as MAE or RMSE, and only task-valid candidates compete for selection as π⋆(x)\pi^\star(x). Because agents cannot backpropagate through discrete tool calls, learning proceeds by explicit comparison among candidates—what the authors call execution-quality supervision. Episodes with zero or one valid candidate are not discarded: they contribute single-execution or failure evidence with weaker signal.

Method

Exploration-time learning. A constrained executor generates at least two distinct, task-valid candidate executions per instance using exploration-only orchestration tools (spawn_subagent, evaluate_* against ground truth), then compares them before terminating with a learning_summary. The summary records the selected strategy, decisive evidence, and failure signals; it is cleaned to remove answer leakage and non-transferable details before updating the task-local experience state. Intermediate tool outputs are tracked as typed artifacts with coordinate provenance so that downstream operations condition on transformed series correctly.

Task-aware tool dropout. To counteract tool-prior collapse, TimeClaw assigns each non-protected tool a keep probability anchored at the least-explored tool:

pkeep(i∣s)=(1+nmin⁡(s)1+ni(s))α,p_{\mathrm{keep}}(i\mid s)=\left(\frac{1+n_{\min}(s)}{1+n_i(s)}\right)^{\alpha},

with protected (task-critical or hinted) tools exempted. A monotonicity proposition guarantees that historically dominant tools are suppressed more strongly than less-used ones. The authors distinguish this exploration-phase pathology from legitimate inference-time tool preference learned from evidence—a distinction prior time-series agent work has not operationalized.

Hierarchical distilled experience. Learning is externalized into five layers: Soul (fixed behavioral anchor), Notes (append-only sample-level distillation source), Memory (structured 9-tuple rules with kind, applicability conditions, preferred/avoided tool sets, rationale, evidence, confidence, and injection status), Tools (per-tool usage advice), and Skills (task-local SOPs). Updates are conflict-aware rather than append-only: rules may be merged, strengthened, or registered as conflicting, and unstable rules can be stored but marked non-injectable. At inference, retrieval filters memory by task, applicability, and injection status; exploration-only tools are removed, and the frozen agent solves tasks under prompts enriched with distilled experience.

Experimental results

Evaluation uses a fixed 17-task MTBench suite spanning finance and weather, aggregated into 30 metrics, with baselines including GPT-4o, Gemini, Claude, DeepSeek, GPT-5.4, Qwen-3, and Llama-3 variants. Exploration runs on a disjoint, schema-aligned corpus with source-level disjointness checks (non-overlapping finance URLs and weather station IDs), so gains reflect transfer rather than memorization.

The headline claim is strong: TimeClaw achieves the best result on 27 of 30 metrics, including 9/10 core numeric prediction metrics, 6/6 indicator metrics, and 12/14 reasoning metrics. Representative numbers:

Task Metric Best baseline TimeClaw
Weather forecast 7D MSE ↓ 14.197 (GPT-5.4) 10.719
Weather forecast 14D MSE ↓ 20.984 (GPT-5.4) 13.525
Finance trend 7D Acc.-5 ↑ 0.569 (GPT-5.4) 0.696
Weather past trend Acc. ↑ 0.980 (GPT-5.4) 0.994
Finance MACD 7D MSE ↓ 0.352 (DeepSeek) 0.201

The main weakness is long-horizon finance correlation, where Gemini outperforms TimeClaw on both label granularities; 30-day MACD MSE also favors Llama-3. These exceptions indicate that distilled experience does not uniformly transfer to long-horizon relational reasoning.

Ablations separate confounded effects. Replacing GPT-5.4 with the same backbone without experience injection ("w/o Exp.") improves weather short-horizon MSE from 14.197 to 13.058, isolating the value of agent runtime and tool access; adding full TimeClaw further reduces it to 10.719, isolating the contribution of learned experience. Removing higher-level distilled layers while retaining tool routines ("w/o Mem.") consistently falls between the baseline and full system, supporting the claim that hierarchical organization—not merely retention—drives retrieval quality. Tool-dropout analysis shows reduced Top-5 tool concentration and increased early coverage across four tasks, consistent with mitigation of premature convergence. A qualitative case study on post-shock MACD forecasting shows a 5.5× error reduction after exploration, driven by a learned rule to treat news as directional bias only and let post-shock price structure determine the trajectory.

Limitations and open questions

The paper concedes several constraints plainly. Empirical claims are limited to benchmarked, verifiable finance and weather tasks with the given tool interfaces; results depend on a frozen base LLM, an externally served API stack, and a predefined tool library. The protocol is offline rather than continual, with a fixed budget of 300 exploration samples per task, and sensitivity to exploration budget, domain shift beyond MTBench-style settings, and deployment under changing distributions remain unexamined. Exploration is expensive: a full run takes roughly one day on two RTX 3090 GPUs plus substantial token expenditure, scaling linearly with tool-library size. No confidence intervals or significance tests are reported, so close comparisons should be read cautiously. Open questions include whether metric-supervised comparison transfers to tasks lacking verifiable ground truth, and how conflict-aware memory behaves over much longer horizons than 30 rules per task.

Conclusion

TimeClaw reframes time-series agent learning around exploratory execution: comparing task-valid candidates under verifiable metrics, distilling evidence into hierarchical external experience, and reinjecting it at inference without parameter updates. Its consistent advantage over strong closed-model baselines—best on 27/30 metrics—supports the thesis that for scientific agents the bottleneck lies in how exploratory experience is compared, organized, and reused, not solely in execution-time capability.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.