- The paper introduces Fusian, a two-stage method that collects stable LoRA checkpoints from the SFT trajectory and uses a Dirichlet policy with REINFORCE to fuse them for continuous MBTI trait control.
- Fusian achieves 6.79 overall MAE and 0.88 Pearson correlation on Qwen3-14B, outperforming prompting, standard LoRA checkpoint selection, PISF, and Personality Vector baselines.
- Ablations show that nonlinear dynamic fusion, stable-basis filtering, and aggressive reward shaping are important, while limitations include single-dimension control, costly offline RL, and evaluation on one model and dataset family.
Motivation and problem statement
Existing personality control in LLMs treats traits as discrete categories. Prompt-based approaches ("act 70% extraverted") are interpreted inconsistently, and standard SFT converges to a single static trait extreme, so neither supports precise control over trait intensity on a continuous spectrum. The paper identifies this granularity gap as a practical limitation for applications such as counseling agents that must shift gradually along a Thinking–Feeling axis or tutors that adjust extraversion to student engagement. Prior parameter-space approaches, notably the Personality Vector method that scales difference vectors with manually chosen coefficients [(Chen et al., 2024), 2503.xxxx-style work of Sun et al.], assume a linear relationship between parameter scaling and semantic output, which the authors argue does not hold in general.
The Fusian framework
Fusian operates in two stages on top of Qwen3-14B, using LoRA (rank 8, α=32, applied to q_proj and v_proj) as the adaptation mechanism.
Stage 1: Trajectory collection. The central hypothesis is that the SFT trajectory itself traverses a continuous "personality manifold": intermediate checkpoints encode intermediate trait intensities. The authors save a LoRA adapter after every gradient update and evaluate each on a held-out set of standardized MBTI questionnaire items (50 per dimension), using a system prompt that casts the model as a human test subject. Likert responses (1–5) are mapped to cumulative points for each pole of a dimension pair, yielding a scalar trait percentage Pt per checkpoint. The raw trajectory is then filtered: a sliding window (size W=3) variance test removes unstable checkpoints, and uniform value sampling selects N=10 adapters evenly spaced across the observed intensity range, producing a basis set that covers the manifold without bias toward regions of slow convergence.
Stage 2: RL-based dynamic fusion. A small MLP policy network (two hidden layers of size 128, LayerNorm, Tanh) maps a normalized target intensity p^∈[−1,1] to Dirichlet concentration parameters via Softplus. Allowing αi<1 is a deliberate design choice: it pushes probability mass toward simplex corners so the policy can collapse onto a single adapter when that fits best. The sampled weight vector linearly combines the frozen basis adapters into a fused adapter. The reward is deliberately aggressive, combining an exponential decay term exp(−λ∣diff∣) with a bonus B granted only when the error falls below a tight threshold δ; optimization uses REINFORCE with an entropy bonus that decays over training. This shaping is motivated by the observation that a linear reward provides constant gradients near the target and fails to incentivize fine adjustment.
Main results
Evaluation uses Mean Absolute Error (MAE) between specified and measured trait intensity, and Pearson correlation between uniformly spaced targets and measured outputs, on the validation questionnaire from the Personality Control Datasets (Chen et al., 2024).
| Method |
Overall MAE ↓ |
Overall Pearson r ↑ |
| Prompt (gpt-5-mini) |
14.97 |
0.35 |
| Prompt (Qwen3-14B) |
24.46 |
0.17 |
| Standard LoRA (closest checkpoint) |
10.18 |
0.69 |
| PISF |
19.08 |
−0.02 |
| Personality Vector |
11.36 |
0.62 |
| Fusian |
6.79 |
0.88 |
Several findings stand out. First, prompt engineering fails even at frontier scale: gpt-5-mini achieves a correlation of only 0.35, indicating it grasps trait direction but cannot discriminate fine-grained intensities such as 60% versus 80%. Second, the Standard LoRA baseline (MAE 10.18, Pt0 = 0.69) substantiates the core hypothesis that the SFT trajectory encodes the personality spectrum — this is a notable result in itself, since it means intermediate checkpoints are semantically meaningful rather than merely undertrained states. Third, Fusian's advantage is largest on the Thinking dimension, where it more than halves the error relative to checkpoint selection (MAE 5.44 vs. 12.80). The Personality Vector method performs well on T, F, P, and J (correlations 0.75–0.84), confirming that linear parameter operations suffice for some traits, but it is unstable elsewhere.
A particularly instructive case is the Intuition (N) dimension, where prompting, PISF, and Personality Vector all yield negative correlations (−0.38, −0.17, −0.15), meaning no monotonic control relationship at all. The authors attribute this to N being linguistically subtle and less linearly separable in parameter space. Fusian retains a correlation of 0.79 on N, which the authors present as evidence that the RL-driven fusion can navigate a non-linear parameter-to-semantics mapping that static scaling cannot.
Ablations and analysis
Ablations on the Feeling dimension (MAE 4.95, Pt1 = 0.97 for the full system) isolate each component's contribution:
| Variant |
MAE ↓ |
Pt2 ↑ |
| Fusian (full) |
4.95 |
0.97 |
| w/o Dynamic Fusion (linear interpolation) |
10.02 |
0.53 |
| w/o Stable Basis |
9.21 |
0.77 |
| w/o Aggressive Reward |
7.08 |
0.86 |
Replacing the learned policy with deterministic linear interpolation between the two nearest basis adapters causes the largest degradation, directly supporting the claim that the LoRA-parameter-space-to-intensity mapping is non-linear. Removing the stability filter nearly doubles the MAE, indicating that noisy trajectory regions genuinely corrupt the interpolation space. The aggressive reward contributes a smaller but consistent gain.
The manifold analysis reports two properties that justify design decisions. Trajectory intensity is non-monotonic early in training — in the Feeling trajectory, step 3 reaches 61.22% while step 6 falls to 51.06% — which motivates variance-based rather than step-based checkpoint selection. It also exhibits diminishing returns near the extremes: the model reaches roughly 96% intensity by step 38 but needs until step 119 for the final increment, implying the long tail of SFT is disproportionately important for full-intensity control. A qualitative case study on the Feeling dimension shows the expected semantic shift: at 20% the model gives analytical problem-solving advice, at 50% a balanced acknowledgment-plus-advice response, and at 80% overtly empathetic emotional support.
Limitations and open questions
The paper is explicit that the framework controls a single MBTI dimension at a time. Composing a full personality profile requires fusing adapters across interacting dimensions (e.g., Extraversion and Thinking), and whether the Dirichlet fusion policy extends to that multi-dimensional setting is left untested. The RL phase is also computationally expensive, since each reward evaluation requires generating and scoring a batch of diagnostic responses, although training is offline. Two further assumptions deserve scrutiny: the trait-percentage metric depends entirely on the model's self-reported Likert answers to the questionnaire under a fixed system prompt, so "intensity alignment" is measured against the model's own test-taking behavior rather than independent behavioral annotation; and the evaluation is confined to one backbone (Qwen3-14B) and one dataset family, leaving the generality of the manifold hypothesis across base models and trait taxonomies (e.g., Big Five) open.
Conclusion
Fusian reframes personality control as a problem of navigating the SFT trajectory rather than converging within it. By treating per-step LoRA checkpoints as a basis spanning a continuous trait manifold and learning a Dirichlet-parameterized fusion policy via REINFORCE, the method achieves an overall MAE of 6.79 and correlation of 0.88 on Qwen3-14B, substantially outperforming prompting, checkpoint selection, and linear vector-scaling baselines, including on the Intuition dimension where competing approaches fail to establish any monotonic control. The ablations confirm that the learned non-linear fusion, stability filtering, and aggressive reward shaping are each necessary. The principal open question is whether the single-dimension fusion policy scales to composing multiple interacting traits into coherent, fully specified personalities.