DreamPRM-1.5: Fine-Grained Reweighted PRMs
- The paper presents an instance‐reweighted bi-level optimization framework that assigns adaptive weights to individual training samples.
- It introduces two parameterizations – Instance Table for maximum expressiveness on small/medium datasets and Instance Net for scalability on large ones.
- DreamPRM-1.5 enhances multimodal PRM performance, achieving 84.6% accuracy on the MMMU benchmark and improving test-time inference.
DreamPRM-1.5 is an instance-reweighted framework for training multimodal process reward models (PRMs) under distribution shifts and noisy data. Its central idea is to adaptively adjust the importance of each training example through bi-level optimization, replacing uniform treatment of samples with meta-learned instance weights aligned to validation objectives that resemble test-time inference. The framework introduces two complementary parameterizations of instance weighting—Instance Table and Instance Net—and is designed to plug directly into test-time scaling pipelines such as reranking and self-consistency. On the MMMU benchmark, the reported best configuration reaches 84.6 accuracy, surpassing GPT-5 according to the paper’s comparison table (Cao et al., 5 Sep 2025).
1. Position within multimodal PRM research
Process Reward Models provide fine-grained evaluation of intermediate reasoning steps and guide the reasoning process. In multimodal settings, their training is complicated by a broader task distribution, more severe train–test distribution shift, and substantial quality imbalance in available reasoning data. The earlier DreamPRM framework addressed this by introducing domain-reweighted training: instead of weighting every sample equally, it learned domain weights over multiple datasets through bi-level optimization, using a separate meta-learning dataset to improve generalization (Cao et al., 26 May 2025).
DreamPRM-1.5 moves this reweighting granularity from the domain level to the instance level. Rather than assigning one weight per dataset or domain, it assigns and dynamically learns weights for each individual training sample. The intended effect is to amplify informative or clean examples while down-weighting noisy or trivial samples. In the terminology of the paper, this is a fine-grained instance-reweighted training paradigm for multimodal PRMs, proposed as a response to the limitations of domain-level reweighting when data quality varies substantially within a domain rather than only across domains (Cao et al., 5 Sep 2025).
This shift in granularity is consequential because multimodal reasoning corpora often mix difficult expert-level problems with easier or noisier instances. A plausible implication is that domain-level averaging can obscure useful within-domain distinctions, whereas per-instance weighting allows the training procedure to respond to heterogeneity at the scale where it actually occurs.
2. Bi-level optimization and the training objective
The core mechanism of DreamPRM-1.5 is a bi-level optimization structure in which the PRM parameters and the instance weights are optimized at different levels. Let denote a training sample, its learnable instance weight, the PRM with parameters , and the step-wise supervision for a solution prefix . The instance-reweighted loss for a sample is given as
and the lower-level optimization updates the PRM parameters for fixed weights:
The upper-level optimization meta-learns on a held-out meta dataset . For a sample 0, let 1 be the generated solution and 2 the ground truth outcome indicator. Step-wise PRM scores are aggregated by a function 3, with the summary noting that mean aggregation is one example. The meta-objective is
4
where 5 is a sigmoid and 6 is the mean squared error (Cao et al., 5 Sep 2025).
The lower level therefore trains the PRM under the current weighting scheme, while the upper level adjusts the weights so that the resulting PRM performs better on a meta-validation objective that simulates inference and candidate selection. The paper characterizes this as meta-learning of instance weights directly guided by validation meta-objectives closely aligned with test-time inference. A common simplification is to treat DreamPRM-1.5 as merely a reranking method; the formulation shows that its principal novelty lies in training-time weight adaptation, with reranking appearing later as an application of the trained PRM.
3. Instance Table and Instance Net
DreamPRM-1.5 proposes two complementary parameterizations for instance weighting. The first, Instance Table, assigns each training sample 7 a distinct learnable weight 8. It functions as an explicit lookup table, is described as best for small or medium datasets, and is said to offer maximum expressiveness. The summary additionally states that the stored values are clipped to a range to avoid extremes (Cao et al., 5 Sep 2025).
The second, Instance Net, replaces per-example storage with a compact shared MLP 9 that predicts a weight from the sample representation:
0
Here, 1 is the instance representation from the PRM, and 2 is sigmoid. Instance Net has a fixed number of parameters, is scalable to large datasets, and is reported to generalize weighting to unseen or new instances (Cao et al., 5 Sep 2025).
The two designs make different trade-offs:
| Strategy | Weight parameterization | Stated use case |
|---|---|---|
| Instance Table | Explicit per-example weight 3 | Small/medium datasets; maximum expressiveness |
| Instance Net | Shared MLP predicts 4 from 5 | Large datasets; generalization to unseen data |
Both strategies plug directly into the same bi-level optimization framework. The reported MMMU results indicate that scalability does not imply superior peak accuracy in every setting: Instance Table reaches 84.6, while Instance Net reaches 83.6. This suggests that explicit per-instance storage may remain advantageous when dataset size permits it, whereas Instance Net is motivated primarily by scalability and transfer to unseen instances rather than by universal dominance in benchmark accuracy.
4. Training pipeline and integration into test-time scaling
The training pipeline is described as pretraining plus supervised fine-tuning as a cold start, followed by bi-level optimization with either Instance Table or Instance Net. Instance weights are learned via a meta-objective on a carefully curated meta set designed to cover real-world, validation-like distributions. This choice is integral to the framework: the upper-level objective is not an abstract regularizer but a direct proxy for downstream reward-model utility under inference conditions (Cao et al., 5 Sep 2025).
At inference time, the trained instance-reweighted PRM is used as a judge or reward model. It can select, rerank, or aggregate candidate solutions generated by base multimodal LLMs. The summary explicitly places reranking and self-consistency among the compatible test-time scaling or selection pipelines. In this usage, the PRM provides improved reward-modeling of intermediate steps, which then translates into better final answer selection (Cao et al., 5 Sep 2025).
This architecture preserves a conceptual distinction between generation and evaluation. The base MLLM proposes candidate solutions, while DreamPRM-1.5 evaluates their step-wise reasoning quality and supports final selection. The paper states that this outperforms both naive selection and prior reasoning or reranking methods. A plausible implication is that the framework can be used as a modular component in systems where the generator itself is not retrained, provided that candidate reasoning traces are available.
5. Distribution shift, noise, and the move beyond domain-level reweighting
The stated motivation for DreamPRM-1.5 is that multimodal PRM training is challenged by distribution shifts and noisy data. Instance-level weighting is presented as the mechanism by which the model adapts to sample quality and informativeness, rather than treating each domain—or each sample within a domain—equally. The paper further states that the meta-optimization loop allows the model to “learn to ignore” outlier, noisy, or non-representative data, directly confronting mismatch between synthetic or noisy training data and real meta or test data (Cao et al., 5 Sep 2025).
This design should be understood in relation to the earlier DreamPRM framework. DreamPRM used domain weights over multiple datasets in a bi-level setup, with the PRM evaluated on a separate meta-learning dataset and the feedback used to update domain weights through an aggregation loss function. That earlier system was introduced to alleviate dataset quality imbalance and improve generalization capability in multimodal PRMs (Cao et al., 26 May 2025). DreamPRM-1.5 preserves the same high-level meta-learning logic while changing the unit of adaptation from domain to instance.
One recurring misconception is that instance reweighting simply replicates data selection heuristics at a finer scale. The framework is more specific than that. It does not merely discard or retain examples according to a precomputed rule; rather, it continuously meta-learns weights according to validation feedback tied to downstream inference. This distinguishes it from static filtering and from uniform sampling.
6. Benchmarks, reported performance, and interpretation
The headline experimental result is on MMMU, described as a benchmark for multimodal, college-level expert reasoning spanning six core disciplines, 30 subjects, and 30 image types. The reported accuracy figures are as follows (Cao et al., 5 Sep 2025):
| Model | MMMU Accuracy | Absolute Gain |
|---|---|---|
| GPT-5 with thinking | 84.2 | --- |
| Gemini 2.5 Pro Deep-Think | 84.0 | --- |
| o3 | 82.9 | --- |
| GPT-5-mini w/ thinking (BASELINE) | 80.0 | --- |
| Vanilla PRM (No Selection) | 79.1 | -0.9 |
| Self-consistency | 81.4 | +1.4 |
| VisualPRM | 80.5 | +0.5 |
| DreamPRM-1.5 – Instance Table | 84.6 | +4.6 |
| DreamPRM-1.5 – Instance Net | 83.6 | +3.6 |
The paper’s stated takeaways are that DreamPRM-1.5 with Instance Table achieves 84.6%, surpassing GPT-5 and setting a new state-of-the-art for open (non-proprietary) PRMs; that even with far fewer training examples, DreamPRM-1.5 outperforms training on much larger datasets with uniform weights; and that both Instance Table and Instance Net significantly improve over all baselines, including self-consistency and prior PRMs (Cao et al., 5 Sep 2025).
In interpretive terms, these results are presented as evidence for the superiority of fine-grained instance reweighting over dataset-level or domain-level reweighting, and more sharply over uniform sampling. The paper also frames the framework as efficient and generalizable, emphasizing that Instance Net provides a scalable route for large datasets and for unseen instances. Code and resources are listed at https://github.com/coder-qicao/DreamPRM-1.5 (Cao et al., 5 Sep 2025).
The naming history is mildly nontrivial. In the earlier DreamPRM paper, “DreamPRM-1.5” is described as the current, improved iteration of DreamPRM rather than a fundamentally different variant (Cao et al., 26 May 2025). In the later paper titled “DreamPRM-1.5,” however, the term is attached to a distinct instance-reweighted framework. The consistent through-line across both uses is bi-level optimization for improving multimodal PRM generalization; the principal architectural change is the transition from domain-reweighted to instance-reweighted training.