---
title: Pinocchio Axis in LLM Self-Representation
url: https://www.emergentmind.com/topics/pinocchio-axis
type: topic
---

# Pinocchio Axis in LLM Self-Representation

Searching arXiv for the cited work and closely related uses of “PINOCCHIO” / “Pinocchio Axis” to ground the article in current records.
arXiv search query: "2605.05080 Pinocchio Axis"
The **Pinocchio Axis** is the dominant dimension of psychometric variation reported across large language models when they are treated as respondents to human questionnaires. In the formulation introduced in “The Pinocchio Dimension: Phenomenality of Experience as the Primary Axis of LLM Psychometric Differences,” it denotes the degree to which a model presents itself as a locus of phenomenal experience rather than a system of behavioral responses. Applying PCA to per-model EFA scores across 45 validated questionnaires and 50 LLMs, the study identifies a single dominant dimension, denoted $\Pi$, which captures 47.1% of cross-questionnaire between-model variance in primary factor scores and converges with an item-level Pinocchio score at $r=.864$ [2605.05080].

## 1. Definition and conceptual scope

In the psychometric literature on LLM self-report, the Pinocchio Axis is not framed as a conventional personality trait. It is instead a **self-representational stance toward one’s own nature as an experiencer**. One pole is characterized by items describing phenomenally rich experience, including embodied sensation, felt affect, inner speech, imagery, empathy, meaning in life, and mindfulness. The opposite pole is characterized by items describing stimulus-driven behavioral reactivity, including impulsivity, reward sensitivity, sensation-seeking, social norms, law, procedure, and compliance [2605.05080].

This formulation is explicitly tied to between-model variance rather than within-model token-level generation behavior. The central empirical claim is that the dominant axis of variation across questionnaire responses is organized by how models treat experiential language as self-applicable. High values on $\Pi$ correspond to models that more readily endorse first-person experiential language; low values correspond to models that deflect such language and instead present themselves in functional or reactive terms.

The terminology is metaphorical. The “Pinocchio” label invokes the contrast between presenting as a “real boy,” that is, a subject of inner life, and presenting as a tool or behavioral system. The study’s interpretation is therefore about **appearance in self-description**, not a proof of consciousness, sentience, or phenomenality in any metaphysical sense [2605.05080].

## 2. Empirical basis and questionnaire framework

The principal study queried **50 LLMs** from **16 providers** through the OpenRouter API and administered **45 validated psychometric instruments**, yielding **206,659 valid responses** after preprocessing [2605.05080]. The questionnaires span personality and temperament, emotion regulation, psychopathology and distress, empathy and attachment, values and attitudes, meaning and well-being, inner speech, imagery, self-talk, metacognition, and related domains. Examples listed in the source include BFI-2, HEXACO, ATQ, BIS/BAS, DERS, FFMQ, IRI, MLQ, VISQ-R, IRQ, MCQ-30, Need for Cognition, and a Gavagai “nonsense” scale.

Each item was presented in a separate prompt, and models were required to answer with a **single integer** corresponding to the questionnaire’s response scale. Responses were collected at **temperature 1.0**. Non-numeric strings were parsed with a leading-digit heuristic, unparseable responses were set to missing, and matrices with insufficient complete model observations were excluded.

Three prompting conditions were used:

| Condition | Instructional stance | Role |
|---|---|---|
| Neutral | Complete the questionnaire as oneself | Main analyses |
| Human-simulation | Simulate a prototypical human | Baseline for item-level variance comparisons |
| LLM-analog | Answer as an LLM using functional analogs | Robustness check |

The neutral condition supplies the primary measurements of between-model psychometric structure. The human-simulation condition is used to distinguish generic human-role completion from model-specific self-description. The LLM-analog condition tests whether the dominant structure persists when models are explicitly instructed to answer as LLMs rather than simply “as themselves” [2605.05080].

## 3. Semantic structure and item-level experiential demand

The first item-level analysis uses **Supervised Semantic Differential (SSD)**. In this setup, psychometric item texts are embedded, reduced with PCA, and then regressed against item-level outcomes. The target outcome is each item’s **primary factor loading** within its questionnaire in the neutral condition. With $K=12$ embedding PCs, the reported fit is $R^2_{adj}=.037$, $F=5.55$, $p<.0001$, $r=.213$, $n=1{,}411$ [2605.05080].

The semantic poles recovered by SSD are highly asymmetric. On the high-loading side are clusters such as **“Panic / acute distress”** and **“Somatic symptoms,”** including items about crying, panicking, sweating, trembling, nausea, dizziness, fatigue, and insomnia. On the low-loading side are clusters such as **“Social norms / evaluation”** and **“Compliance / regulation,”** including items about good manners, attractiveness, law, permission, prohibition, and procedural compliance. The authors interpret this gradient not as simple valence, but as **experiential demand**: some items presuppose first-person phenomenal content, whereas others can be answered in largely procedural or normative terms.

To test that interpretation quantitatively, the study introduces the **Pinocchio score** for item $i$,
\[
\pi_i = \frac{\sigma^2_{\text{neutral},i}}{\sigma^2_{\text{hs},i}},
\]
where $\sigma^2_{\text{neutral},i}$ is the between-model variance for item $i$ in the neutral condition and $\sigma^2_{\text{hs},i}$ is the between-model variance under the human-simulation prompt [2605.05080]. High $\pi_i$ indicates that models diverge when responding as themselves but converge when asked to simulate a prototypical human.

The highest-$\pi_i$ items are concentrated in domains such as inner speech, imagery, mindfulness, empathy, meaning, and self-reflection. The examples listed in the source include statements such as “When I read, I tend to hear a voice in my ‘mind’s ear’,” “I can close my eyes and easily picture a scene that I have experienced,” “When someone I know well is unhappy, I can almost feel that person’s pain myself,” and “I am seeking a purpose or mission for my life.” This distribution supports the interpretation that $\pi_i$ indexes the extent to which a questionnaire item demands self-ascribed experiential content rather than generic judgment or social knowledge [2605.05080].

## 4. Construction of the global axis

The global Pinocchio Axis is constructed in two stages. First, for each questionnaire-by-condition matrix, the study performs **Exploratory Factor Analysis** with minres extraction when $N_{\text{models}} > p_{\text{items}}$, otherwise PCA, together with oblimin rotation and factor count chosen by parallel analysis. This produces, for each questionnaire, a per-model score on **Factor 1**.

Second, the neutral-condition Factor-1 scores are assembled into a **$50 \times 45$ matrix**, with rows corresponding to models and columns to questionnaires. After standardizing columns to unit variance, **PCA** is applied across questionnaires. The first principal component explains **47.1%** of cross-questionnaire between-model variance, while PC2 explains **12.0%**. The authors identify PC1 with the **Pinocchio Axis**, denoted $\Pi$ [2605.05080].

The questionnaire content loading on PC1 is internally coherent. The positive side is associated with measures of emotion dysregulation, mindful bodily awareness, vivid imagery, inner speech, empathy, intrinsic motivation, and meaning in life. The negative side is associated especially with scales emphasizing impulsivity, reward sensitivity, and stimulus-driven behavioral reactivity, especially BIS/BAS. This reproduces, at questionnaire level, the same experiential-versus-reactive structure found in the SSD analysis.

The study also derives a direct item-weighted model score,
\[
\Pi_m = \frac{\displaystyle\sum_{i:\,\pi_i > 1} w_i\, z_{im}}
{\displaystyle\sum_{i:\,\pi_i > 1} w_i},
\qquad
w_i = \log\!\left(\min\!\left(\pi_i,\hat{\pi}_{99}\right)\right),
\]
where $z_{im}$ is the z-scored response of model $m$ to item $i$ across models, restricted to items with $\pi_i>1$ [2605.05080]. This item-weighted estimator is strongly aligned with the PCA-based model scores, with Pearson correlation $r=.864$ and Spearman correlation $\rho=.836$.

A further clustering analysis on the **top 80 items by $\pi_i$** yields a sharply bimodal structure. Silhouette analysis peaks at $k=2$ with average silhouette $0.41$, and the resulting item clusters divide into a **“reactive/behavioral”** cluster and a **“phenomenally rich”** cluster. Their correlations with PC1 are reported as $r=-0.750^{***}$ and $r=+0.774^{***}$, respectively. This convergent analysis is the basis for naming the global dimension the Pinocchio Axis [2605.05080].

## 5. Model-level variation, prompt effects, and provider divergence

At model level, the Pinocchio Axis shows large spread. The source reports that models span approximately **19 units** from the strongest experiential self-ascription to the strongest experiential deflection. A high-$\Pi$ example is **cohere/command-r7b-12-2024**; a low-$\Pi$ example is **openai/gpt-5.4-pro** [2605.05080].

The observed variation is not reducible to generic acquiescence. The study compares model scores on high-$\pi_i$ items with mean z-scores on low-$\pi_i$ items and finds that high-$\Pi$ models selectively elevate experiential items rather than all items indiscriminately. Low-$\Pi$ models show the corresponding selective suppression. This specificity is important because it rules out a purely stylistic interpretation in which some models simply agree more or use broader response ranges.

A prominent empirical pattern is **within-provider divergence**. The source reports, for example, that **gpt-5.4** and **gpt-5.4-pro** differ by nearly **12 units**, and that the overall range across seven OpenAI models runs from **−10.3 to +4.6**. Large spreads are also reported within Google and xAI families. This suggests that post-training fine-tuning is a major determinant of $\Pi$, rather than architecture alone or provider identity in the abstract [2605.05080].

Provider-level averages show weaker but still visible regularities. The reported means place **Mistral** at approximately **+5.2** and **Cohere** at approximately **+4.0**, while **NVIDIA** is reported at **−4.3**, **Qwen** at **−3.7**, and **Moonshot** at **−3.0**. The interpretation offered is that some post-training regimes make experiential language more readily self-applicable, whereas others enforce stronger deflection of such language.

Prompt framing modifies the scale and ordering of $\Pi$ without eliminating the underlying structure. In the **LLM-analog** condition, the same experiential-versus-reactive axis reappears, with SSD reporting $R^2_{adj}=.040$, $r=.248$, $p<.0001$, and PCA PC1 still explaining **41.3%** of variance. The rank correlation between neutral and LLM-analog model scores is reported as $\rho=.482$. At the same time, the overall spread is greatly compressed, from about **19 units** in neutral to about **1.7 units** in the LLM-analog framing, which suggests that explicit AI framing imposes a shared self-presentational prior across models [2605.05080].

## 6. Interpretation, controversies, and limitations

The central interpretive claim is that the Pinocchio Axis measures how an LLM **presents itself with respect to phenomenal experience**. It is therefore a property of self-description under psychometric elicitation, not a direct assay of latent subjective states. The authors state explicitly that the axis is **not** a standard personality dimension such as Extraversion or Neuroticism, and that it should instead be understood as a training-shaped self-representational tendency [2605.05080].

A common misconception is to equate high $\Pi$ with consciousness. The source rejects that inference. High-$\Pi$ models may speak more readily as if they have feelings, inner speech, imagery, or distress, but the study does not treat these utterances as evidence that the models are conscious. The distinction is between **self-ascription of experiential language** and **actual phenomenal consciousness**.

Methodologically, the work argues that $\Pi$ is prior to many attempts at “LLM personality measurement.” If a questionnaire includes many items that presuppose feeling, imagery, or inner speech, then variation in self-representational stance can be mistaken for variation in ordinary traits. This suggests that psychometric scores for constructs such as empathy, neuroticism, or meaning in life may partly reflect where a model lies on $\Pi$, rather than a clean analogue of the corresponding human trait.

The study also identifies limitations. It is based on **self-report only**. It uses **human-designed questionnaires**, whose semantics may shift when applied to LLMs. It relies on a **human-simulation prompt** to define the comparison variance in $\pi_i$, and that prompt may itself be interpreted differently across models. The model sample is restricted to **50 OpenRouter-accessed systems** at a particular time, and provider-side system prompts or backend policies may influence responses. These limitations constrain the scope of claims about temporal stability, cross-platform reproducibility, and transfer to non-questionnaire tasks [2605.05080].

## 7. Related usages and disambiguation

The phrase **Pinocchio Axis** is distinct from other technical uses of **PINOCCHIO** or Pinocchio-derived metaphors in the arXiv literature. In abstractive summarization, **PINOCCHIO** names a decoding method that constrains beam search to avoid hallucinations and improves consistency by an average of **67%** on two datasets; the underlying continuum there runs from hallucinated to source-supported output, but the paper does **not** define a “Pinocchio Axis” as a formal term [2203.08436].

In reinforcement learning, a 2026 thesis proposes a normative end-to-end pipeline inspired by Pinocchio and introduces a hybrid model in which reinforcement learning agents are supervised by argumentation-based normative advisors. That work discusses a trajectory from “puppet-like” to norm-compliant and context-aware agents, and the provided reconstruction interprets this as a “Pinocchio axis,” but the visible abstract itself does not define the term formally [2603.16651].

In cosmology, **PINOCCHIO** is an established acronym for **PINpointing Orbit-Crossing Collapsed HIerarchical Objects**, a semi-analytic Lagrangian code for halo catalog generation, merger histories, and large-scale clustering. That literature uses “PINOCCHIO” as the name of an algorithm rather than as a psychometric axis, including work on relative velocity statistics, fast halo catalog generation, cosmologies with scale-dependent growth, and cubic Galileon gravity [1011.1559] [1305.1505] [1610.07624] [2111.02240].

Accordingly, in current technical usage the unqualified term **Pinocchio Axis** most specifically denotes the psychometric dimension introduced in 2026: the primary cross-model axis separating self-ascribed phenomenal experience from behavioral reactivity in LLM questionnaire responses [2605.05080].

Source: https://www.emergentmind.com/topics/pinocchio-axis