Papers
Topics
Authors
Recent
Search
2000 character limit reached

LLM Output Detectability and Task Performance Can be Jointly Optimized

Published 2 May 2026 in cs.CL | (2605.01350v1)

Abstract: Detecting machine-generated text is essential for transparency and accountability when deploying LLMs. Among detection approaches, watermarking is a statistically reliable method by design -- it embeds detectable signals into LLM outputs by biasing their token distributions. However, it has been reported that watermarked LLMs often perform worse on downstream tasks. We propose PUPPET, a framework that fine-tunes an LLM via reinforcement learning to generate text that is both more detectable and better performing on downstream tasks. We use two reward functions: a detector that outputs a machine-class likelihood and an evaluator that measures a task-specific metric. Experiments on long-form QA, summarization, and essay writing show that LLMs trained with PUPPET achieve high detectability competitive with watermarking methods while outperforming them on downstream tasks. The analysis shows that this optimization can be performed efficiently with only a few thousand samples in 1--2 GPU hours. Moreover, these gains are consistent across out-of-domain tasks, different LLM families, and model sizes, and are even robust to paraphrasing attacks.

Summary

  • The paper introduces PUPPET, a DPO framework that combines detector and task-quality rewards to train LLMs whose outputs are both more detectable and more effective.
  • PUPPET improves AUROC by up to 12.3 points, ROUGE-L by 2.9 points, and IELTS judge scores by 0.42 points while matching or exceeding watermarking baselines.
  • The method generalizes across tasks, model families, detectors, and paraphrasing attacks, converging with roughly 1,000–2,000 samples and 1–2 hours of training on one A6000 GPU.

Overview

The paper introduces PUPPET, a framework that fine-tunes a LLM via Direct Preference Optimization (DPO) so that its outputs are simultaneously more detectable as machine-generated text and stronger on downstream tasks. The work is motivated by a known limitation of watermarking: because watermarking methods such as KGW, SynthID, Unigram, EXPGumbel, and K-SemStamp bias token selection at decoding time to embed detectable signals, they optimize detectability alone and can degrade task performance (2605.01350). PUPPET instead internalizes detection-relevant features into the model's parameters through preference learning, using two reward signals: a detector score SDet=p(Machine∣t)S^{\text{Det}} = p(\text{Machine} \mid t) from an external classifier, and an evaluator score SEvalS^{\text{Eval}} computed by a task-specific metric such as ROUGE-L or LLM-as-a-Judge.

Method

For each query qq in a task-specific dataset QQ, PUPPET samples k=5k=5 candidate responses from the base model π\pi, scores each candidate with both rewards, combines them via z-score normalization within the sample set,

Sj=α⋅z(SjDet)+(1−α)⋅z(SjEval),S_j = \alpha \cdot z(S^{\text{Det}}_j) + (1-\alpha)\cdot z(S^{\text{Eval}}_j),

and selects the highest-scoring response as "chosen" and the lowest-scoring as "rejected" for DPO training with LoRA. The hyperparameter α\alpha controls the trade-off; all main experiments use α=0.5\alpha = 0.5. The authors adopt DPO over policy-gradient alternatives such as PPO for training stability.

Experiments use Llama-3-8B-Instruct as the primary base model, with Qwen3-4B/8B/14B for robustness checks. Training data comprise 5,000 instances each of ELI5 (long-form QA), Multi-News (summarization), and IELTS essay writing, with 200 evaluation samples per task. The OpenAI Detector serves as the primary detector reward, with FakeSpotAI and MAGE evaluated as alternatives.

Main results

Across the three benchmarks, PUPPET-trained models match or exceed watermarking baselines on detection while outperforming them on task quality. Relative to the Vanilla baseline, PUPPET improves AUROC by up to +12.3 points (IELTS: 87.0 → 99.3), ROUGE-L by up to +2.9 points (ELI5: 22.9 → 25.8), and the IELTS judge score by up to +0.42 points (6.77 → 7.19). Against watermarking baselines, PUPPET achieves a task-performance gap of up to 6% on IELTS (SynthID: 6.76 vs. PUPPET: 7.19) while surpassing KGW—the strongest detector baseline—on 2 of 3 benchmarks. On ELI5, PUPPET reaches AUROC 100.0, exceeding KGW's 99.6.

A SHAP-based span analysis supports the mechanism: after PUPPET training, the coverage of spans contributing to detector salience increases by 1.8% and coverage of reference-overlapping (ROUGE-L-relevant) spans increases by 13%, indicating that the model acquires features serving both objectives rather than trading one off against the other.

Notably, the paper concedes that the task-performance degradation of watermarked models reported in prior literature was not consistently reproduced in its own settings, conjecturing—without dedicated analysis—that recent LLMs remain fluent under distributional constraints. The paper's contribution therefore does not depend on reproducing that degradation: PUPPET actively raises task performance above what watermarking achieves.

Generalization and robustness

Out-of-domain transfer: models trained on a single task improve detection on unseen tasks without degrading their performance. A model trained only on ELI5 lifts Multi-News AUROC from 92.8 to 97.3 and IELTS AUROC from 87.0 to 94.9, suggesting the learned detection features generalize across tasks rather than overfitting to the training distribution.

Base model family and size: on Qwen3-8B, PUPPET improves average detection from 81.2 to 94.1 and task performance from 44.4 to 46.0. Across sizes, vanilla detectability decreases monotonically with parameter count (88.7 → 81.2 → 76.0 for 4B → 8B → 14B), yet PUPPET reverses this trend, achieving 97.1, 94.1, and 97.4 respectively—with the largest absolute gain (+21.4 points) on the largest, hardest-to-detect model. The single exception to simultaneous improvement across all experiments is a small, statistically insignificant IELTS judge-score drop for Qwen3-14B (7.86 → 7.81).

Detector choice: gains hold across OpenAI Detector, FakeSpotAI, and MAGE when each detector serves as both reward and evaluation signal. One minor exception is MAGE on Multi-News with Llama-3, where detection drops slightly (97.2 → 95.4), hinting at mild tension between text-quality improvement and MAGE-detectable features.

Paraphrasing attacks: under Dipper paraphrasing, all five watermarking baselines suffer severe AUROC degradation—up to 49.2 points—whereas PUPPET loses only 1.2 points on average. The authors attribute this robustness to learned features (sentence structure, stylistic patterns) that survive surface-level rewriting, though this explanation remains a hypothesis rather than a demonstrated mechanism. Results vary across detectors under attack: FakeSpotAI saturates near 100 throughout, while MAGE shows notable post-attack drops (e.g., 95.4 → 84.9 on Multi-News), indicating that paraphrase robustness partly depends on how sensitive each detector's features are to surface perturbation.

Efficiency and ablation

PUPPET is markedly lightweight. Both objectives converge within roughly 1k–2k samples on ELI5 (AUROC reaches 100.0 at 1k samples), and each run completes in 1–2 hours on a single NVIDIA A6000 (48GB)—a small budget compared to large-scale DPO pipelines using 150K–200K pairs. On harder settings (Qwen3 on IELTS, where vanilla AUROC is only 66.6), convergence requires more samples and the task-performance curve is unstable, which the authors link to the weak initial detection signal.

The reward-composition ablation confirms that both components are necessary. Optimizing detection alone (α=1\alpha=1) yields large AUROC gains (up to +12.9) but degrades task performance (up to −1.5 ROUGE-L); optimizing task performance alone (SEvalS^{\text{Eval}}0) improves tasks but provides no reliable detection gain. Full PUPPET recovers 95.3–100% of the detection-only gain while also improving task performance over Vanilla. For Qwen3 on IELTS, recovery falls to 59.4%, exposing a harder trade-off when the base model's detectability is low and SEvalS^{\text{Eval}}1 is fixed at 0.5—systematic tuning of SEvalS^{\text{Eval}}2 is left as future work.

Model attribution

As an exploratory analysis, the authors show that PUPPET's controlled upward shift in machine-class likelihood enables model attribution—distinguishing outputs of a PUPPET-trained model from other models via a thresholded detector likelihood. Attribution AUROC rises from near-chance or moderate levels at baseline (58.99–77.71 between vanilla Llama-3 and vanilla Qwen3) to 94.47–99.14 for PUPPET Llama-3 versus vanilla Qwen3, with Cohen's SEvalS^{\text{Eval}}3 reaching 4.62 on IELTS. When both models are PUPPET-trained, attribution collapses toward chance where their likelihood distributions converge (SEvalS^{\text{Eval}}4 on ELI5). The authors explicitly frame these findings as preliminary evidence, and the asymmetry between perspectives (Qwen3-based attribution is weaker because vanilla Llama-3 already occupies a high machine-likelihood region) underscores that attribution depends on baseline distributional separation rather than a property intrinsic to PUPPET.

Limitations and open questions

Several limitations are stated directly in the paper. First, the failure to reproduce prior reports of watermark-induced task degradation is unexplained; no dedicated analysis substantiates the authors' fluency conjecture. Second, hyperparameters—including SEvalS^{\text{Eval}}5 and SEvalS^{\text{Eval}}6—are held fixed without systematic search, and the IELTS results for Qwen3 show that a fixed SEvalS^{\text{Eval}}7 may underperform when the detection signal is weak. Third, the mechanistic nature of the learned detection-relevant features is not established; the paraphrase-robustness explanation rests on the assumption that such features are structural rather than surface-level. Fourth, the model-attribution result is exploratory and contingent on baseline likelihood separation between models. Finally, the ethics statement cautions that all detection approaches, including PUPPET, remain research-stage techniques whose false positives can cause harm in high-stakes deployments.

Conclusion

PUPPET demonstrates that detectability and downstream task performance need not be traded off: a DPO-based fine-tuning procedure with composite detector-plus-evaluator rewards matches watermarking-level detection while exceeding watermarking methods on task quality, generalizes across tasks, model families, sizes, and detectors, and withstands paraphrasing attacks that cripple watermarking—all within a few thousand samples and 1–2 GPU hours. The open questions the paper leaves are specific: a mechanistic account of the acquired detection features, principled tuning of the reward weighting, and validation of the attribution capability beyond the preliminary evidence presented here.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.