Papers
Topics
Authors
Recent
Search
2000 character limit reached

Fusian: Multi-LoRA Fusion for Fine-Grained Continuous MBTI Personality Control in Large Language Models

Published 16 Mar 2026 in cs.CL | (2603.15405v1)

Abstract: LLMs have demonstrated impressive capabilities in simulating diverse human behaviors and personalities. However, existing methods for personality control, which include prompt engineering and standard Supervised Fine-Tuning (SFT), typically treat personality traits as discrete categories (e.g., "Extroverted" vs. "Introverted"), lacking the ability to precisely control the intensity of a trait on a continuous spectrum. In this paper, we introduce Fusian, a novel framework for fine-grained, continuous personality control in LLMs. Fusian operates in two stages: (1) Trajectory Collection, where we capture the dynamic evolution of personality adoption during SFT by saving a sequence of LoRA adapters, effectively mapping the continuous manifold of a trait; and (2) RL-based Dynamic Fusion, where we train a policy network using Reinforcement Learning to dynamically compute mixing weights for these frozen adapters. By sampling from a Dirichlet distribution parameterized by the policy network, Fusian fuses multiple adapters to align the model's output with a specific numerical target intensity. Experiments on the Qwen3-14B model demonstrate that Fusian achieves high precision in personality control, significantly outperforming baseline methods in aligning with user-specified trait intensities.

Authors (2)

Summary

  • The paper introduces Fusian, a two-stage method that collects stable LoRA checkpoints from the SFT trajectory and uses a Dirichlet policy with REINFORCE to fuse them for continuous MBTI trait control.
  • Fusian achieves 6.79 overall MAE and 0.88 Pearson correlation on Qwen3-14B, outperforming prompting, standard LoRA checkpoint selection, PISF, and Personality Vector baselines.
  • Ablations show that nonlinear dynamic fusion, stable-basis filtering, and aggressive reward shaping are important, while limitations include single-dimension control, costly offline RL, and evaluation on one model and dataset family.

Motivation and problem statement

Existing personality control in LLMs treats traits as discrete categories. Prompt-based approaches ("act 70% extraverted") are interpreted inconsistently, and standard SFT converges to a single static trait extreme, so neither supports precise control over trait intensity on a continuous spectrum. The paper identifies this granularity gap as a practical limitation for applications such as counseling agents that must shift gradually along a Thinking–Feeling axis or tutors that adjust extraversion to student engagement. Prior parameter-space approaches, notably the Personality Vector method that scales difference vectors with manually chosen coefficients [(Chen et al., 2024), 2503.xxxx-style work of Sun et al.], assume a linear relationship between parameter scaling and semantic output, which the authors argue does not hold in general.

The Fusian framework

Fusian operates in two stages on top of Qwen3-14B, using LoRA (rank 8, α=32\alpha=32, applied to q_proj and v_proj) as the adaptation mechanism.

Stage 1: Trajectory collection. The central hypothesis is that the SFT trajectory itself traverses a continuous "personality manifold": intermediate checkpoints encode intermediate trait intensities. The authors save a LoRA adapter after every gradient update and evaluate each on a held-out set of standardized MBTI questionnaire items (50 per dimension), using a system prompt that casts the model as a human test subject. Likert responses (1–5) are mapped to cumulative points for each pole of a dimension pair, yielding a scalar trait percentage PtP_t per checkpoint. The raw trajectory is then filtered: a sliding window (size W=3W=3) variance test removes unstable checkpoints, and uniform value sampling selects N=10N=10 adapters evenly spaced across the observed intensity range, producing a basis set that covers the manifold without bias toward regions of slow convergence.

Stage 2: RL-based dynamic fusion. A small MLP policy network (two hidden layers of size 128, LayerNorm, Tanh) maps a normalized target intensity p^∈[−1,1]\hat{p} \in [-1,1] to Dirichlet concentration parameters via Softplus. Allowing αi<1\alpha_i < 1 is a deliberate design choice: it pushes probability mass toward simplex corners so the policy can collapse onto a single adapter when that fits best. The sampled weight vector linearly combines the frozen basis adapters into a fused adapter. The reward is deliberately aggressive, combining an exponential decay term exp⁡(−λ∣diff∣)\exp(-\lambda |diff|) with a bonus BB granted only when the error falls below a tight threshold δ\delta; optimization uses REINFORCE with an entropy bonus that decays over training. This shaping is motivated by the observation that a linear reward provides constant gradients near the target and fails to incentivize fine adjustment.

Main results

Evaluation uses Mean Absolute Error (MAE) between specified and measured trait intensity, and Pearson correlation between uniformly spaced targets and measured outputs, on the validation questionnaire from the Personality Control Datasets (Chen et al., 2024).

Method Overall MAE ↓ Overall Pearson rr ↑
Prompt (gpt-5-mini) 14.97 0.35
Prompt (Qwen3-14B) 24.46 0.17
Standard LoRA (closest checkpoint) 10.18 0.69
PISF 19.08 −0.02
Personality Vector 11.36 0.62
Fusian 6.79 0.88

Several findings stand out. First, prompt engineering fails even at frontier scale: gpt-5-mini achieves a correlation of only 0.35, indicating it grasps trait direction but cannot discriminate fine-grained intensities such as 60% versus 80%. Second, the Standard LoRA baseline (MAE 10.18, PtP_t0 = 0.69) substantiates the core hypothesis that the SFT trajectory encodes the personality spectrum — this is a notable result in itself, since it means intermediate checkpoints are semantically meaningful rather than merely undertrained states. Third, Fusian's advantage is largest on the Thinking dimension, where it more than halves the error relative to checkpoint selection (MAE 5.44 vs. 12.80). The Personality Vector method performs well on T, F, P, and J (correlations 0.75–0.84), confirming that linear parameter operations suffice for some traits, but it is unstable elsewhere.

A particularly instructive case is the Intuition (N) dimension, where prompting, PISF, and Personality Vector all yield negative correlations (−0.38, −0.17, −0.15), meaning no monotonic control relationship at all. The authors attribute this to N being linguistically subtle and less linearly separable in parameter space. Fusian retains a correlation of 0.79 on N, which the authors present as evidence that the RL-driven fusion can navigate a non-linear parameter-to-semantics mapping that static scaling cannot.

Ablations and analysis

Ablations on the Feeling dimension (MAE 4.95, PtP_t1 = 0.97 for the full system) isolate each component's contribution:

Variant MAE ↓ PtP_t2 ↑
Fusian (full) 4.95 0.97
w/o Dynamic Fusion (linear interpolation) 10.02 0.53
w/o Stable Basis 9.21 0.77
w/o Aggressive Reward 7.08 0.86

Replacing the learned policy with deterministic linear interpolation between the two nearest basis adapters causes the largest degradation, directly supporting the claim that the LoRA-parameter-space-to-intensity mapping is non-linear. Removing the stability filter nearly doubles the MAE, indicating that noisy trajectory regions genuinely corrupt the interpolation space. The aggressive reward contributes a smaller but consistent gain.

The manifold analysis reports two properties that justify design decisions. Trajectory intensity is non-monotonic early in training — in the Feeling trajectory, step 3 reaches 61.22% while step 6 falls to 51.06% — which motivates variance-based rather than step-based checkpoint selection. It also exhibits diminishing returns near the extremes: the model reaches roughly 96% intensity by step 38 but needs until step 119 for the final increment, implying the long tail of SFT is disproportionately important for full-intensity control. A qualitative case study on the Feeling dimension shows the expected semantic shift: at 20% the model gives analytical problem-solving advice, at 50% a balanced acknowledgment-plus-advice response, and at 80% overtly empathetic emotional support.

Limitations and open questions

The paper is explicit that the framework controls a single MBTI dimension at a time. Composing a full personality profile requires fusing adapters across interacting dimensions (e.g., Extraversion and Thinking), and whether the Dirichlet fusion policy extends to that multi-dimensional setting is left untested. The RL phase is also computationally expensive, since each reward evaluation requires generating and scoring a batch of diagnostic responses, although training is offline. Two further assumptions deserve scrutiny: the trait-percentage metric depends entirely on the model's self-reported Likert answers to the questionnaire under a fixed system prompt, so "intensity alignment" is measured against the model's own test-taking behavior rather than independent behavioral annotation; and the evaluation is confined to one backbone (Qwen3-14B) and one dataset family, leaving the generality of the manifold hypothesis across base models and trait taxonomies (e.g., Big Five) open.

Conclusion

Fusian reframes personality control as a problem of navigating the SFT trajectory rather than converging within it. By treating per-step LoRA checkpoints as a basis spanning a continuous trait manifold and learning a Dirichlet-parameterized fusion policy via REINFORCE, the method achieves an overall MAE of 6.79 and correlation of 0.88 on Qwen3-14B, substantially outperforming prompting, checkpoint selection, and linear vector-scaling baselines, including on the Intuition dimension where competing approaches fail to establish any monotonic control. The ablations confirm that the learned non-linear fusion, stability filtering, and aggressive reward shaping are each necessary. The principal open question is whether the single-dimension fusion policy scales to composing multiple interacting traits into coherent, fully specified personalities.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.