---
title: Direct Future Prediction (DFP)
url: https://www.emergentmind.com/topics/direct-future-prediction-dfp
type: topic
---

# Direct Future Prediction (DFP)

Direct Future Prediction (DFP) is a supervised learning paradigm in sequential decision-making, forecasting, and control that replaces value or policy optimization with direct multivariate prediction of future measurements conditioned on current observations, goals, and (optionally) actions. DFP enables agents and models to flexibly pursue multiple objectives, incorporate dynamic preferences, and decouple policy learning from sparse scalar rewards. DFP has been foundational in visual sensorimotor control [1611.01779], multi-objective resource allocation [2403.16298], and more recently, outcome-driven fine-tuning of large language models for forecasting [2502.05253].

## 1. Formal Definition and Principles

In the classical DFP framework, the agent at each discrete time step $t$ is presented with:
- A high-dimensional sensory observation, $x_t$ (e.g., a raw image, structured state vector)
- A low-dimensional measurement vector, $m_t$ (e.g., health, queue utilization)
- A goal or preference vector, $g$ or $g_t$, encoding the importance of each measurement dimension and/or prediction horizon

For each candidate action $a \in \mathcal{A}$, the agent predicts a vector of future measurement differences for designated offsets $\tau_1, ..., \tau_n$:
$$
f_t = \left[ m_{t+\tau_1} - m_t,\; ...,\; m_{t+\tau_n} - m_t \right] \in \mathbb{R}^d
$$
where $d = \dim(m)\times n$. The utility of a hypothetical future $f$ is defined as $u(f;g) = g^\top f$. The network $f_\theta$ predicts $y_t^{(a)} = f_\theta(x_t, m_t, g, a)$. Control is effected by selecting
$$
a^*_t = \arg\max_{a} g^\top y_t^{(a)}
$$
DFP thus converts value estimation into direct regression over multi-timescale, multi-dimensional future outcomes.

## 2. Network Architecture and Training Procedures

DFP implementations share several architectural elements:
- **Perception module:** Processes $x_t$ via a convolutional or MLP backbone; e.g., [1611.01779] uses an 84×84 image through a 3-layer CNN, flattening to a 512-d vector.
- **Measurement and Goal Embedding:** $m_t$ and $g$ pass through parallel MLPs (e.g., three-layer, 128 units [1611.01779], or 128 for MRSch [2403.16298]).
- **Joint Representation:** The embeddings are concatenated into a joint latent vector.
- **Dueling Streams:** The joint vector is processed by an *expectation* stream $E(\cdot)$ and an *action/advantage* stream $A(\cdot)$ as in the dueling-DQN architecture, with normalization across actions to enforce zero mean per measurement dimension.
- **Prediction Heads:** Produce, for each $a\in\mathcal{A}$, the predicted $d$-dimensional future measurement vector.

Supervised learning minimizes the mean squared error between predicted and observed future measurement differences:
$$
L(\theta) = \sum_{i=1}^N \| f_i - f_\theta(x_i, m_i, g_i, a_i) \|^2
$$
Data are generated by an $\epsilon$-greedy or random-exploration policy, storing $(x_t, m_t, g_t, a_t, f_t)$ in an experience buffer. No external reward model or human demonstration is required.

## 3. Applications and Domain Adaptations

### 3.1 Sensorimotor Control

The original DFP formulation was applied to first-person Doom-based environments [1611.01779]. Measurements included health, ammo, and frag count; goals were arbitrary linear combinations across time scales (e.g., maximizing frags over multiple future steps). DFP agents outperformed DQN, A3C, and Deep Successor Representation in challenging 3D navigation and combat, demonstrating strong transfer to unseen goals and environments.

### 3.2 Multi-Resource Scheduling

"MRSch" extends DFP to high-performance computing (HPC) cluster scheduling with multiple resources (CPU, burst buffer, power) [2403.16298]. State encoding aggregates queued job descriptors and per-resource-unit status. The goal vector $g_t$ is recomputed at each scheduling event to dynamically prioritize resources under contention:
$$
r_j = \frac{\sum_{i=1}^N P_{ij} t_i}{\sum_{\ell=1}^R \sum_{i=1}^N P_{i\ell} t_i}
$$
where $P_{ij}$ is the fraction of resource $j$ requested by job $i$, and $t_i$ its estimated runtime. MRSch achieved up to 48% higher node utilization and reduced average job wait and slowdown compared to fixed-weight RL and heuristics, highlighting the utility of multi-objective, goal-adaptive forecasting.

### 3.3 LLM Forecasting via Direct Future Prediction

Outcome-driven fine-tuning (ODFT) adapts DFP to large language models for probabilistic forecasting [2502.05253]. Rather than acting in an environment, the model generates multiple reasoning/forecast trajectories for each real-world question via self-play, ranks them by proximity to eventual ground-truth outcome, and applies Direct Preference Optimization (DPO) to preference pairs. For binary-resolution questions, the error is $r(p, o) = |p - o|$ for the predicted probability $p$ and outcome $o$.

The DPO objective is:
$$
\mathcal{L}_{\mathrm{DPO}} = -\mathbb{E}_{(x,y^+,y^-)} \left[ \log \sigma(\beta[\log\pi_\theta(y^+|x) - \log\pi_\theta(y^-|x)]) \right]
$$
where $y^+$ and $y^-$ are more and less accurate predictions, and $\beta$ is a temperature parameter. Fine-tuning small models (Phi-4 14B, DeepSeek-R1 14B) with this self-generated supervision improved forecast accuracy (Brier score reduction of 6.6–9.5%) to rival that of much larger frontier models.

## 4. Inference Procedures and Goal Adaptivity

DFP models, both in control and scheduling, compute per-action forecasts for all candidate actions using a single forward pass. The action with maximal predicted utility under the current $g$ is selected:
$$
a^*_t = \arg\max_a g^\top f_\theta(x_t, m_t, g, a)
$$
In multi-resource or dynamic-goal settings, $g_t$ is recomputed to reflect momentary preferences or congestion. Such adaptivity is crucial for robust performance under changing objectives or workload characteristics, as fixed-goal RL and scalar-reward optimization fail to adapt to resource imbalances [2403.16298].

## 5. Empirical Findings and Ablation Studies

Key empirical results include:
- In Doom-based control scenarios, DFP matched or outperformed RL baselines: 84% health in navigation (vs 59% A3C, 25% DQN); 33 frags in D3 (vs 5.6 A3C, 1.2 DQN) [1611.01779].
- In HPC resource scheduling, MRSch delivered up to 48% higher node utilization, 30% higher burst-buffer utilization, 48% lower average job wait, and 41% lower slowdown versus heuristic and RL baselines [2403.16298].
- In LLM forecasting, outcome-driven DFP fine-tuning closed the performance gap between 14B parameter models and much larger models such as GPT-4o, with fine-tuned models achieving Brier scores (0.200, 0.197) statistically indistinguishable from GPT-4o (0.196) [2502.05253].

Ablation studies revealed:
- Predicting multiple measurements across multiple time scales substantially boosts performance over scalar, single-offset prediction [1611.01779].
- Skip identical-forecast questions in DPO-based LLM fine-tuning to ensure only divergent reasoning informs the updates [2502.05253].
- Dynamic goal weighting outperforms fixed-scalar RL in resource scheduling, especially under non-uniform workload patterns [2403.16298].
- LoRA adaptation rank tuning showed 16 as optimal (8 underfit, 32 no further gain) for LLMs [2502.05253].

## 6. Limitations, Extensions, and Current Research Trajectories

DFP, while effective for dense measurement streams and explicit goal representation, faces notable challenges:
- Extension to longer horizons, multi-way outcome spaces, and continuous-event forecasting requires appropriate distance metrics and scaling of prediction heads [2502.05253].
- Interpretability remains limited; deep DFP architectures for infrastructure scheduling make black-box decisions, impeding production verification [2403.16298].
- Starvation mitigation in scheduling necessitates mechanisms (windowed reservation, backfilling) not inherent to base DFP [2403.16298].
- Calibration-aware variants (e.g., adding scoring-rule minimization or post-hoc re-ranking) represent open directions [2502.05253].

Ongoing work explores chaining DFP-style self-play and DPO for sequential or multi-step forecasting, richer measurement sets, and generalization to tasks where direct reward supervision is poorly specified or unreliable. DFP’s central paradigm—multivariate, goal-conditional supervised prediction—continues to underpin advances in both online decision systems and offline outcome modeling.

Source: https://www.emergentmind.com/topics/direct-future-prediction-dfp