---
title: 'RAD-DPO: Dual Domain Methods'
url: https://www.emergentmind.com/topics/rad-dpo
type: topic
---

# RAD-DPO: Dual Domain Methods

Searching arXiv for both uses of “RAD-DPO” to ground the article in the cited papers.
RAD-DPO is an acronym with two unrelated technical uses in arXiv literature. In planetary radiation science, it denotes the **RAD Dose Prediction Operator**, an empirically calibrated operator derived from Mars Science Laboratory Radiation Assessment Detector measurements to predict galactic-cosmic-ray-induced surface dose rate and dose equivalent on Mars as functions of atmospheric pressure and heliospheric modulation [1507.03473]. In recommender and retrieval modeling, it denotes **Robust Adaptive Denoising Direct Preference Optimization**, a preference-optimization objective for generative retrieval in e-commerce that modifies DPO to accommodate hierarchical Semantic IDs, noisy implicit feedback, and multi-label relevance [2602.23964]. The shared acronym masks a complete separation of domain, objective, mathematical structure, and operational interpretation.

## 1. Terminological scope and disambiguation

The ambiguity of RAD-DPO is substantive rather than notational. In the Martian-radiation setting, “RAD” refers to the **Radiation Assessment Detector** on board Curiosity, and “DPO” refers to a **Dose Prediction Operator** fitted from one-sol-averaged surface data. In the e-commerce retrieval setting, “RAD-DPO” expands to **Robust Adaptive Denoising Direct Preference Optimization**, where “RAD” characterizes a set of denoising and robustness mechanisms layered onto DPO for alignment of generative retrieval models [1507.03473] [2602.23964].

A common misconception is that RAD-DPO names a single optimization framework with cross-domain reuse. The literature in fact uses the acronym for two independent constructs. One is a radiation-environment estimator operating on physical variables such as pressure $P$ and modulation potential $\Phi$; the other is a training objective operating on queries $x$, positives $\mathcal{Y}_{\rm pos}$, negatives $\mathcal{Y}_{\rm neg}$, and sequence probabilities under $\pi_\theta$. The only commonality is acronymic coincidence.

## 2. The RAD Dose Prediction Operator for the Martian surface

In Guo et al., the RAD Dose Prediction Operator is introduced as an empirical representation of how Martian-surface radiation dose varies with atmospheric and heliospheric conditions during the first MSL Martian year. The underlying physical definition of the surface dose rate is written as
$$
D(\Phi,P) =
\sum_{j} \iint_{E,\epsilon} \lambda_{j}(E,\epsilon)\,F_{j}(\Phi,P,E)\,dE\,d\epsilon / m,
$$
where $j$ labels particle species, $E$ is particle kinetic energy, $\epsilon$ the energy deposited in the detector, $F_j(\Phi,P,E)$ the surface spectral fluence at atmospheric pressure $P$ and heliospheric modulation $\Phi$, $\lambda_j(E,\epsilon)$ the detector yield, and $m$ the detector mass [1507.03473].

Within the range of the first MSL Martian year, the two dominant, independent drivers of $D$ were found to be the surface pressure $P$ and the modulation potential $\Phi$. Two empirical forms were fitted from one-sol-averaged RAD data:
$$
D(\Phi,P)=D_{01}+\kappa\cdot P+\beta\cdot\Phi
$$
for **Model I**, and
$$
D(\Phi,P)=D_{02}+\kappa\cdot P+\frac{\alpha_1}{\Phi+\alpha_2}
$$
for **Model II**. Here $P$ is in Pa, $\Phi$ in MV, and $D$ in $\mu{\rm Gy/day}$; the fitted parameters are detector-specific. The paper attributes short-term diurnal variations to daily thermal tides, long-term seasonal changes to Martian atmospheric pressure changes, and long-term solar-cycle and solar-rotation effects to heliospheric modulation of the primary GCR flux [1507.03473].

This operator is explicitly empirical. It does not replace transport modeling; rather, it quantitatively demonstrates how long-term influences of pressure and solar modulation are related to measured dose rates, enabling predictions of Martian surface radiation environment under different solar modulations and atmospheric conditions.

## 3. Calibration, fitted coefficients, corrections, and validity limits

All fits for the Martian RAD Dose Prediction Operator were carried out for two RAD dose channels—silicon detector B and plastic, tissue-equivalent detector E—after removing SEP excursions and diurnal-pressure oscillations. A bootstrap Monte Carlo was used to propagate measurement and binning uncertainties [1507.03473].

For the diurnal pressure coefficient $\kappa_d$, obtained by fitting one-Mars-sol binned perturbations $\delta D_h$ versus $\delta P_h$ through
$$
\delta D_h=\kappa_d\cdot\delta P_h,
$$
the reported values are as follows:

| Parameter | Detector E | Detector B |
|---|---:|---:|
| $\kappa_d$ | $(-0.1306 \pm 0.0176)\ \mu{\rm Gy/day/Pa}$ | $(-0.100 \pm 0.060)\ \mu{\rm Gy/day/Pa}$ |

In the long-term DPO formulae, $\kappa \equiv \kappa_d$ was used for both Models I and II. After normalizing to the reference pressure $P_0=840\ {\rm Pa}$, the fitted Model I coefficients are $D'_{01}=(267.17\pm16.29)\ \mu{\rm Gy/day}$ and $\beta=(-0.11\pm0.03)\ \mu{\rm Gy/day/MV}$ for detector E, and $D'_{01}=(224.10\pm16.12)\ \mu{\rm Gy/day}$ and $\beta=(-0.12\pm0.03)\ \mu{\rm Gy/day/MV}$ for detector B. For Model II, the corresponding normalized coefficients are $D'_{02}=(149.6\pm15.3)\ \mu{\rm Gy/day}$, $\alpha_1=(3.3\pm0.8)\times10^4\ ({\rm MV}\cdot\mu{\rm Gy/day})$, and $\alpha_2=(1.3\pm0.1)\times10^{-3}\ {\rm MV}$ for detector E, and $D'_{02}=(89.4\pm14.0)\ \mu{\rm Gy/day}$, $\alpha_1=(3.7\pm0.8)\times10^4\ ({\rm MV}\cdot\mu{\rm Gy/day})$, and $\alpha_2=(9.9\pm0.1)\times10^{-4}\ {\rm MV}$ for detector B [1507.03473].

The correction conventions are equally explicit. Detector E is inherently tissue-equivalent and includes both charged and neutral contributions, whereas detector B sees mostly charged. To convert silicon-detector dose $D_B$ into water-equivalent dose, the prescribed factor is $1.38$. Dose equivalent is then obtained from
$$
H=Q\cdot D,
$$
with mean quality factor $Q=3.05\pm0.26$, adopted from Hassler et al. (2014). Operational use proceeds by selecting $P$ and $\Phi$, choosing Model I or Model II, selecting detector channel, computing
$$
D_{\rm raw}=D_0+\kappa\cdot P+\{\beta\cdot\Phi\ {\rm\ or\ }\alpha_1/(\Phi+\alpha_2)\},
$$
and optionally converting to dose equivalent with the stated $Q$ [1507.03473].

The stated validity range is limited to fits based on data with $P\in[\sim700,1000]\ {\rm Pa}$ and $\Phi\in[\sim550,800]\ {\rm MV}$. Extrapolation well outside those windows carries large model uncertainty, and divergence of Model I versus II at $\Phi\approx250\ {\rm MV}$ is specifically noted. Both models also assume linear independence of $P$ and $\Phi$ effects; simulations are said to hint that at very low $P$ or very high $\Phi$, second-order coupling may appear. SEP events are excluded from the operator and must be added separately, while dust opacity and atmospheric composition changes are described as negligible at RAD’s sensitivity [1507.03473].

## 4. Robust Adaptive Denoising Direct Preference Optimization in generative retrieval

In the 2026 e-commerce literature, RAD-DPO addresses alignment of **Generative Retrieval (GR)** models that retrieve items by autoregressive decoding of structured **Semantic IDs (SIDs)**. Given a query or context $x$, the model $\pi_\theta$ generates an SID
$$
y=(y^1,y^2,\dots,y^L),
$$
where each token corresponds to one level of a hierarchical taxonomy, such as category, subcategory, and item code. The generated SID is then mapped back to one or more SKUs for ranking or display [2602.23964].

The work situates itself as a modification of Direct Preference Optimization. The classic pairwise DPO loss, ignoring SFT, is given by
$$
\mathcal{L}_{\rm DPO}(\theta)=-\log\sigma\!\bigl(\beta\cdot(\hat r_\theta(x,y_w)-\hat r_\theta(x,y_\ell))\bigr),\qquad
\hat r_\theta(x,y)=\frac{\log\pi_\theta(y\mid x)}{|y|}.
$$
Three failure modes are identified when this is applied directly to structured SIDs. First, standard DPO induces **gradient conflicts on shared prefixes** because positives and negatives often share the first $k$ tokens. Second, it is **vulnerable to pseudo-negatives** arising from noisy implicit feedback. Third, in queries with multiple relevant items, it creates **probability squeezing** among valid candidates, degrading recall [2602.23964].

The proposed RAD-DPO therefore combines three mechanisms: token-level gradient detachment to protect shared hierarchical prefixes, similarity-based dynamic reward weighting to mitigate label noise, and a multi-label global contrastive objective integrated with global SFT loss to expand positive coverage. This suggests that the method is less a minor loss reweighting than a composite redesign of the preference-learning pipeline around the structural peculiarities of SIDs.

## 5. Objective construction and optimization mechanics

The token-level detachment mechanism begins from the longest common prefix between a positive-negative pair $(y_w,y_\ell)$:
$$
k=\max\{t:y_w^{1:t}=y_\ell^{1:t}\}.
$$
For the negative sequence, standard log-likelihood is
$$
\log\pi_\theta(y_\ell\mid x)=\sum_{t=1}^{L}\log\pi_\theta(y_\ell^t\mid y_\ell^{<t},x).
$$
RAD-DPO defines
$$
1_{\rm diff}(t)=
\begin{cases}
0,& t\le k,\\
1,& t>k,
\end{cases}
$$
and applies stop-gradient to shared-prefix terms:
$$
\log\hat\pi_\theta(y_\ell\mid x)=\sum_{t=1}^{L}\Bigl[\mathrm{SG}\bigl((1-1_{\rm diff}(t))\log\pi_\theta(y_\ell^t\mid y_\ell^{<t},x)\bigr)+1_{\rm diff}(t)\log\pi_\theta(y_\ell^t\mid y_\ell^{<t},x)\Bigr].
$$
The resulting gradient is
$$
\nabla_\theta\log\hat\pi_\theta(y_\ell\mid x)=\sum_{t=1}^{L}1_{\rm diff}(t)\,\nabla_\theta\log\pi_\theta(y_\ell^t\mid y_\ell^{<t},x),
$$
so the first $k$ tokens incur zero gradient from the negative sample, while the positive sequence remains fully differentiable [2602.23964].

For noise mitigation, the method uses cosine similarity between final hidden-state vectors at EOS:
$$
\mathrm{sim}(y_w,y_\ell)=
\frac{h(y_w,x)\cdot h(y_\ell,x)}
{\|h(y_w,x)\|\;\|h(y_\ell,x)\|}.
$$
The first $N=4096$ similarity scores are collected into $\mathcal{S}$, quartiles $Q_{25}$, $Q_{50}$, and $Q_{75}$ are computed, and anchors are updated periodically. The negative-sample weight is then defined piecewise as
$$
w(\mathrm{sim})=
\begin{cases}
1.0,& \mathrm{sim}<Q_{25},\\
0.5+\dfrac{0.5}{1+\exp\bigl(\lambda(\mathrm{sim}-Q_{50})\bigr)},& Q_{25}\le \mathrm{sim}\le Q_{75},\\
0.5,& \mathrm{sim}>Q_{75},
\end{cases}
\qquad \lambda=12.
$$
This gives softer penalties to highly ambiguous negatives [2602.23964].

For multi-label settings, the positive set is
$$
\mathcal{Y}_{\rm pos}=\{y_{w1},y_{w2},\dots,y_{wM}\}.
$$
The global SFT term is
$$
\mathcal{L}_{\rm SFT}=-\frac{1}{|\mathcal{Y}_{\rm pos}|}\sum_{y\in\mathcal{Y}_{\rm pos}}\log\pi_\theta(y\mid x),
$$
while the global contrastive preference term selects the highest-score positive $y_w^*$ under the current model and contrasts it against the negative pool. The unweighted form is
$$
\mathcal{L}_{\rm PL}=-\log\sigma\!\Bigl(\beta\cdot\log\sum_{y_j\in\mathcal{Y}_{\rm neg}}
\exp\bigl(\hat r_\theta(x,y_w^*)-\hat r_\theta(x,y_j)\bigr)\Bigr),
$$
and the weighted form incorporates both detachment and $w(\mathrm{sim}(y_w^*,y_j))$. The total RAD-DPO objective combines $\mathcal{L}_{\rm SFT}$ and $\mathcal{L}_{\rm PL}$; the training algorithm updates $\theta$ using the gradient of $\mathcal{L}_{\rm SFT}+\mathcal{L}_{\rm PL}$ [2602.23964].

## 6. Training procedure, inference profile, and reported results

The reported training algorithm takes a pretrained SFT model $\pi_\theta^{(0)}$ and a preference dataset $\{(x_i,\mathcal{Y}_{\rm pos}^{(i)},\mathcal{Y}_{\rm neg}^{(i)})\}_{i=1}^N$. The listed hyperparameters are learning rate $10^{-6}$, batch size $64$, temperature $\beta=0.1$ (tuned), smoothness $\lambda=12$, and warm-up on the first $N_w=4096$ pairs. Each minibatch requires forward passes for all $y$, extraction of hidden states $h(y,x)$, quartile updates beyond warm-up, computation of $\mathcal{L}_{\rm SFT}$ and the detached, weighted $\mathcal{L}_{\rm PL}$, and gradient descent on their sum [2602.23964].

At inference, the method uses beam-search decoding with beam size $128$, generates top-$K$ SIDs, maps each to candidate SKUs, and merges them with other retrieval branches. The stated serving profile is latency below $150\,{\rm ms}$ at $15\,{\rm QPS}$ on a single RTX 4090 [2602.23964].

The reported offline evaluation is on JD.com’s 700 M-log dataset, test day $T+1$, for a 1.7 B model. The SFT baseline reports Halluc. $0.0575$, $R@8$ $0.3755$, $R@64$ $0.6139$, $R@128$ $0.6775$, and MRR $0.2859$. Standard DPO reports Halluc. $0.0838$, $R@8$ $0.3648$, $R@64$ $0.6001$, $R@128$ $0.6642$, and MRR $0.2747$. RAD-DPO reports Halluc. $0.0652$, $R@8$ $0.3888$, $R@64$ $0.6247$, $R@128$ $0.6864$, and MRR $0.2947$. The summary statement is that RAD-DPO reduces hallucination by approximately $22\%$ relative to DPO and boosts $R@128$ by $+2.2\%$ absolute; it is also reported to scale across model sizes from $0.6$ B to $8$ B and to remain robust when preference data is reduced from $10$ M to $50$ M [2602.23964].

Online A/B testing places the 1.7 B RAD-DPO model as a parallel generative branch in JD.com’s live search engine. The paper reports a $+0.34\%$ absolute lift in User Conversion Rate over one week of testing on hundreds of millions of users, with the same sub-$150\,{\rm ms}$ latency at $15\,{\rm QPS}$ [2602.23964].

## 7. Comparative interpretation and domain-specific significance

The two RAD-DPO formulations exemplify how identical acronyms can encode entirely different epistemic roles. The Martian RAD Dose Prediction Operator is a **measurement-driven empirical predictor**: its parameters are fitted coefficients, its inputs are environmental variables, and its output is a dose or dose-equivalent rate under quiet-Sun conditions. Robust Adaptive Denoising Direct Preference Optimization is a **training-time alignment objective**: its parameters are neural-network weights, its inputs are preference-annotated retrieval instances, and its output is an updated generative retrieval model [1507.03473] [2602.23964].

Their methodological contrast is similarly sharp. The Martian operator assumes linear independence of pressure and solar-modulation effects over a bounded validity window and explicitly excludes SEP events. The retrieval objective, by contrast, is built around sequence-level and token-level differentiable optimization, representation-space similarity weighting, and multi-label contrastive learning. A plausible implication is that the acronym overlap is most usefully handled as a disambiguation problem in bibliographic and indexing systems, because the surrounding terminology alone—Mars surface dose versus Semantic IDs and SKUs—determines the intended referent.

In both cases, however, RAD-DPO denotes an attempt to make a complex observational or behavioral system operational. In planetary science, that operationalization supports empirical prediction of the Martian surface radiation environment under varying pressure and heliospheric conditions. In e-commerce retrieval, it supports alignment of generative retrieval models with structured preference data while reducing gradient conflicts, pseudo-negative sensitivity, and multi-label probability squeezing [1507.03473] [2602.23964].

Source: https://www.emergentmind.com/topics/rad-dpo