- The paper reveals that LLM persona comprises dual aspects: frame-dependent geometric coordination and stable aggregated tendencies.
- It introduces an Item-Dimension Matrix method to map sequential responses, using SPD geometry and Big Five aggregates for discrimination.
- Empirical results show a collapse-recovery pattern in geometric features under frame perturbation, questioning standard trait-based assessments.
The Dual Nature of LLM Persona: Frame-Dependent Coordination and Frame-Robust Aggregates
Introduction and Motivation
Standard methods for assessing LLM "personality" employ psychometric inventories like the IPIP-50, reducing model responses to aggregate scores such as the Big Five. These approaches, however, disregard the within-instance correlation structure, treating persona as a set of static, trait-like aggregates. This paper rigorously interrogates whether the apparent geometric structure in LLM-generated persona measures is intrinsic to the model or merely a consequence of shared temporal structure (i.e., the fixed order of questions) during measurement.
Critically, due to the autoregressive nature of LLMs, responses are sensitive to sequence context. Thus, the paper posits three hypotheses: (1) the geometric structure is intrinsic and order-robust, (2) it is an artifact destroyed by order perturbations, or (3) it reflects frame-dependent coordination, requiring temporal alignment for discriminative geometry to emerge. Only the third predicts a collapse-recovery pattern under frame perturbation and alignment Figure 1.

Figure 1: Only the frame-dependence hypothesis predicts a collapse-recovery "V-shape" in geometric discriminability across order conditions.
Methodological Framework
The author introduces the Item-Dimension Matrix method to construct, for each LLM response instance, a 10×5 matrix mapping sequential IPIP-50 answers onto Big Five dimensions as a multivariate time series. From this, within-instance 5×5 correlation matrices (on the SPD manifold) are computed, encoding the temporal coordination across personality dimensions during autoregressive response generation.
Three experimental conditions are compared:
- Fixed Order (FO): All instances use the same canonical item order (shared frame).
- Random Order, Native Frame (RO): Each instance answers in a uniquely shuffled order (frame misalignment).
- Random Order, Bootstrap Shared Frame (RO-BTSP): Each random-order instance is resampled into a common frame for analysis, aligning temporal structure post hoc.
Feature extraction leverages manifold geometry (log-Euclidean mapping), including full SPD structure, eigenvalues, eigenvectors, and conventional Big Five aggregates. Discriminability is evaluated via clustering (UMAP+Spectral or k-means), with silhouette/AUC metrics and bootstrap resampling for statistical rigor.
Empirical Findings: Dual Nature and Collapse-Recovery
Analysis produces a robust dissociation between aggregated and geometric features. Aggregated Big Five scores, by construction, are robust to frame manipulation but show moderate degradation under randomization—preserving high discriminability even as item order varies. In contrast, geometric (SPD) features exhibit catastrophic collapse under frame misalignment, but substantially recover (and often surpass Big Five) when shared frames are restored—even when responses remain randomized in content.

Figure 2: Geometric features (SPD, Eigenvalues) collapse under frame misalignment (RO) but recover with bootstrap-aligned frames (RO-BTSP); Big Five show opposite pattern.
The frame-dependence of geometric features is quantified: 74% of performance loss is attributable to frame effects for SPD features, while Big Five losses are purely driven by order (content) effects.
Further, under shared frames—even with content randomization—SPD features yield higher discriminability than Big Five aggregates, indicating emergent coordination patterns invisible to aggregation. This collapse-recovery—and, in large samples, collapse-recovery-surpass—pattern supports the frame-dependent coordination hypothesis, rejecting conventional trait interpretations.

Figure 3: Correlational structure (eigenvalue spacing, entropy) persists under randomization; collapse in geometric discrimination is not due to structural destruction, but frame misalignment.
PCA visualization with sample size N≈2000 shows SPD features maintain cross-cultural separation under FO vs. RO, while Big Five overlap, confirming larger-scale robustness.


Figure 4: PCA of SPD and Big Five features, N≈2000: SPD geometry retains cross-cultural separability under FO, collapses under RO; Big Five features lose discrimination with aggregation.
Theoretical Implications
LLM "personality" is not a unitary, intrinsic set of traits, but a dual-natured phenomenon:
- Frame-Dependent Geometric Coordination: Underlies transient, order-sensitive correlation patterns across dimensions; these are not intrinsic model properties but emergent from temporal alignment of context during generation.
- Frame-Robust Aggregated Tendencies: Big Five-like aggregates are stable to frame perturbation, but degrade under content randomization and lose fine-grained discriminatory power as aggregation increases.
This undermines trait-based interpretations (as applied in human psychometrics) and underscores the measurement-dependence of LLM persona. The empirical superiority of SPD geometry under shared frames reveals that critical information about model "persona" resides in the high-order structure of context coordination—not simply in content aggregation.
Practical Implications
Frame-dependence in LLMs demands new standards for evaluation:
- Bias and Fairness Auditing: Fixed order protocols overstate geometric stability; robust evaluation requires variation of both content and frame, and decomposition of order/frame effects.
- Model Comparison and Safety: Without frame-aware assessment, aggregated metrics may obscure context-sensitive biases and overestimate cross-situational consistency.
- Mitigation Strategies: Structural bias mitigation should target dynamic coordination patterns in sequence generation, not solely calibrate output distributions.
Limitations and Future Directions
Generality beyond GPT-4o and across diverse cultural contexts requires further empirical support. The detected frame-dependence likely generalizes to all autoregressive LLMs given shared architectural constraints, but systematic cross-model comparisons are necessary. The mechanism—whether rooted in attention/saliency dynamics, internal state propagation, or positional encoding—should be tested via analysis of internal activation manifolds and neuron/attention-level correlational geometries. Extending evaluation to other bias domains (political, gender, etc.) will clarify the universality and specificity of frame-dependence. Future mitigation work should focus on manipulating generation-time context and developing frame-normalized bias metrics.
Conclusion
This paper demonstrates that LLM persona measures comprise two dissociable phenomena: frame-dependent, context-coordinated geometry and frame-robust aggregate tendencies. The geometric structure, responsible for much of the apparent discriminative power in LLM bias evaluation, is not an intrinsic property but a measurement artifact dependent on shared temporal frames. Valid assessment and mitigation require frame-aware protocols that separate content effects from coordination artifacts. This dual-nature framework provides a rigorous basis for future research into LLM behavior, bias, and interpretability, with broad implications for fairness, evaluation, and AI safety.