- The paper introduces Milkyway, a frozen-backbone agent that uses repeated predictions and temporal internal feedback to evolve reusable procedures for factor tracking, evidence gathering, and uncertainty handling.
- Milkyway raises scores from 44.07 to 60.90 on FutureX and from 62.22 to 77.96 on FutureWorld, outperforming the strongest baselines by 3.40 and 5.99 points, respectively.
- Ablation results show that harness evolution adds 5.00–11.43 points beyond the agent scaffold, while limitations include sparse pre-resolution evidence, possible harness drift, limited evaluation windows, and an unspecified cohort size.
Milkyway is a self-evolving agent system for future prediction that improves without modifying model parameters. Its central observation is that when an unresolved question is revisited over time, the contrast between earlier and later prediction attempts exposes process-level failures — missed factors, unexamined sources, premature uncertainty collapse — before the outcome is known. The authors call this signal internal feedback, and they use it to update a persistent, editable "harness" governing factor tracking, evidence gathering, and uncertainty handling. On FutureX and FutureWorld, the system raises the overall score from 44.07 to 60.90 and from 62.22 to 77.96 respectively, relative to a direct GPT-5.4 reference (2604.15719).
Motivation: supervision arrives too late
Future prediction differs from standard agent benchmarks in that the target answer does not exist at inference time. The evidence base is partial and evolving, while the realized outcome y becomes available only at resolution, and even then it is a coarse signal: it says whether the prediction was right but not what the agent should have tracked, verified, or kept uncertain. Existing improvement methods — Outcome-based Reinforcement Learning and Future-as-Label — learn only from resolved outcomes (2604.15719). The paper's argument is that the pre-resolution trajectory itself contains richer diagnostic signal, and that this signal is currently unused.
The formalization makes the temporal structure explicit: a question q issued at time i is revisited at checkpoints τ1​<⋯<τT​<r, where the agent produces predictions zt​=π(q,τt​,St​) from the evidence state St​ it actually surfaces. Because the evidence surfaced at each checkpoint depends on the agent's current procedure, later runs can reveal what earlier runs failed to retrieve or interpret — the basis of internal feedback.
System design
Milkyway keeps the backbone fixed and adapts a harness Ht​=(Ft​,Et​,Ut​) with three components: factor tracking, evidence gathering and interpretation, and uncertainty handling. Two modules interact through it. A ReAct-style BaseAgent performs each checkpoint prediction; a Harness Editor never answers the question but revises the harness across checkpoints. Each run produces a compact checkpoint note nt​ covering question state, key evidence, judgment trajectory, and unresolved risks; raw traces are retained only as bounded fallback.
For t≥2, internal feedback It​=InternalFeedback(Ht​,N≤t​) is synthesized from the accumulated note history, and the reusable part is written back: q0. Crucially, the harness stores only procedural, reusable guidance — question-specific facts stay in the notes. Evolution operates at two timescales: within a question, updates are provisional and immediately available to later checkpoints; across questions, carryover occurs only after resolution, when the realized outcome serves as a retrospective check that validates, refines, or corrects provisional revisions. The implementation uses a staged-then-promote writeback policy with structural validation but no human review gate.
Benchmark results
All methods use GPT-5.4 as backbone under a forward-only protocol, with shared ceilings of 256k context tokens and 100 tool calls per question. The authors concede the search stack is not fully unified, so they frame comparisons as controlled evaluations under shared protocol and bounded compute rather than identical realized compute.
| Method |
FutureX (Ovr.) |
FutureWorld (mean±std) |
| GPT-5.4 (web search) |
44.07 |
62.22 ± 2.79 |
| MiroFlow |
57.50 |
71.97 ± 3.45 |
| Flash-Searcher |
45.34 |
63.52 ± 2.81 |
| MemEvolve+Flash-Searcher |
48.06 |
66.12 ± 2.39 |
| AgentKB+smolagents |
46.65 |
64.88 ± 2.04 |
| Milkyway |
60.90 |
77.96 ± 3.73 |
Milkyway leads the strongest baseline by 3.40 points on FutureX and 5.99 on FutureWorld. It does not dominate every subset — MiroFlow is slightly higher on FutureX Level 4 — so the authors emphasize aggregates. Against external experience-reuse baselines, which are the closest structural alternatives, the result supports the claim that a structured harness updated from temporal internal feedback is more useful than generic memory accumulation.
A rolling daily evaluation across five FutureWorld windows shows Milkyway first on all five days, with margins of 3.52–8.89 points, and leading on Day 1 before any within-window retrospective update is available. The authors appropriately caution that five windows cannot establish long-horizon temporal stability.
Mechanism analysis
A fixed-cohort ablation re-runs predictions daily from five days before resolution (T−5d) to one day before (T−1d), comparing three settings:
| Setting |
T−5d |
T−3d |
T−1d |
| GPT-5.4 (web search) |
57.48 |
63.21 |
65.54 |
| Milkyway (No Harness) |
61.78 |
68.75 |
71.60 |
| Milkyway (Full Harness) |
69.03 |
73.75 |
76.98 |
Two findings stand out. First, the agent scaffold alone (No Harness) contributes 3.31–6.06 points over the base model, but enabling harness evolution adds a further 5.00–11.43 points, isolating the harness as the dominant source of gain. Second, the Full Harness advantage is visible from T−5d and grows toward resolution, particularly at T−2d, which is consistent with the proposed mechanism: as resolution nears, later evidence makes earlier omissions more diagnosable. The trajectory is not perfectly monotonic, which the authors attribute to live-evaluation variance. Notably, the paper's cohort size for this ablation is left as an unfilled placeholder (q1), so the ablation's statistical basis is not fully specified in the provided text.
On general-purpose benchmarks (GAIA, HLE) with harness evolution disabled, Milkyway's scaffold remains competitive with MiroFlow and Flash-Searcher (82.4 vs. 84.2 GAIA overall), indicating the architecture does not materially degrade broader agent capability, though the authors present this only as supporting evidence.
Limitations and open questions
The paper is explicit about its boundaries. The method requires questions that can be revisited before resolution and public evidence that evolves enough for temporal contrasts to be informative; it is unclear how much benefit accrues for questions with sparse pre-resolution signal. The textual harness may accumulate redundancy or drift over long horizons, and no mechanism is proposed for pruning or consolidation. Evaluation covers two benchmarks and limited online windows. The ablation's fixed cohort size is not reported. Open questions include whether harness updates transfer across question families, how the retrospective check scales as the number of accumulated notes grows, and whether the observed advantage persists over substantially longer evaluation horizons.
Conclusion
The paper's contribution is a reframing of supervision in future prediction: temporal contrasts across repeated predictions on an unresolved question constitute a usable feedback signal available before the outcome is known. Milkyway operationalizes this by evolving a persistent prediction harness with a frozen backbone, using realized outcomes only as retrospective checks. The benchmark gains, daily stability, and ablation evidence jointly indicate that process-level internal feedback is a practical complement to outcome-based learning, within the stated limits of the evaluated settings.