---
title: 'InMind: Cognitive & Neural Decoding Framework'
url: https://www.emergentmind.com/topics/inmind
type: topic
---

# InMind: Cognitive & Neural Decoding Framework

InMind encompasses both a cognitively anchored evaluation framework for Large Language Models (LLMs) in capturing and applying individualized human reasoning styles in social deduction games (SDGs), and the $i$MIND neural decoding architecture for subject-invariant decoding of visual signals from fMRI. Both approaches address distinct dimensions of subjectivity and individuation in cognition—one in behavioral strategy, the other in neural representation and decoding—exemplifying convergent themes in computational cognitive science and neural engineering [2508.16072][2509.17313].

## 1. Conceptual Foundation and Motivation

InMind (LLM): The InMind framework emerges from limitations in existing theory-of-mind (ToM) benchmarks that typically restrict evaluation to global plausibility of intent judgments or false-belief attribution. These settings fail to probe whether LLMs genuinely internalize the *style* and *trajectory* of a specific individual's reasoning as leveraged in real-time, sequential social contexts. SDGs like Avalon, with their transparency of utterances, sequential moves, and evolving private/public states, present a testbed for observing and evaluating actual individualized reasoning behaviors [2508.16072].

$i$MIND (Neural Decoding): In parallel, the $i$MIND (Insightful Multi-subject Invariant Neural Decoding) model is introduced to mitigate the challenge that cross-subject neural decoding from fMRI is dominated by subject-specific variability, which both limits generalization and occludes interpretation of brain-based visual processing. The $i$MIND architecture is designed to explicitly factorize, decode, and interpret both individual- and object-level components of neural representations, enabling scalable, interpretable, and generalizable neural decoding within and across subjects [2509.17313].

## 2. InMind LLM Framework: Methodology and Task Structure

The InMind framework is built around a formal structured game representation:
\[
\mathcal{G} = \langle mode,\,\mathcal{A},\,\{E_z\}_{z=1}^m,\,\mathcal{F}\rangle
\]
where $\mathcal{A}$ designates the mapping from players to roles, $E_z$ encodes round-level (utterances, game state, and strategy trace), and $\mathcal{F}$ is the global session reflection. Data is collected under two "modes": Observer (no direct participation) and Participant (active player).

Dual-layer cognitive annotations are attached: (1) round-level *strategy traces* $\{S_z\}$ capturing quasi-veridical records of evolving beliefs, intentions, and inference processes; and (2) a high-level reflective summary $\mathcal{F}$, interpreting key events and meta-strategic assessments.

Evaluation is operationalized as a two-stage pipeline:
- Capturing: Profile induction from observer-mode annotation using a “ProfilePrompt”—free-form profiles summarizing temporally diffuse reasoning styles.
- Applying: Given a subject profile $\mathcal S$ and a participant-mode session, the model is evaluated on four tasks:

| Task               | Description                                                    | Metric(s)                                   |
|--------------------|----------------------------------------------------------------|---------------------------------------------|
| Player Identification | Rank the true subject in anonymized session                  | Top-$k$ accuracy $\mathrm{Acc}_{ID^k}$      |
| Reflection Alignment | Fill in masked IDs in post-game reflection                   | Exact match $\mathrm{Acc}_{Ref}$            |
| Trace Attribution     | Map masked trace segments to correct IDs per round           | Match accuracy; adaptation $\Delta$         |
| Role Inference        | Infer player roles at each round, strict or grouped labeling | Strict/group accuracy                       |

General-purpose LLMs and reasoning-enhanced models (DeepSeek-R1, QwQ, O3-mini) are compared under zero-shot prompting regimens with enforced structured output, and all data is Mandarin voice-chat transcribed [2508.16072].

## 3. $i$MIND Neural Decoding Pipeline: Architecture and Objectives

The $i$MIND model is a three-stage end-to-end pipeline for multi-subject fMRI decoding:

1. **Self-supervised ViT-MAE Pretraining**  
   The input is a flattened, uniformly padded voxel vector $\mathbf V\in\mathbb R^L$, split into $N$ non-overlapping 64-voxel patches. The model employs a 12-layer ViT encoder and an 8-layer transformer decoder. Masked patches (75%) are reconstructed with a voxel-wise mean-squared error loss:
   \[
   \mathcal L_{\rm rec} = \frac1{|\Omega|}\sum_{i\in\Omega}\|\mathbf V_i-\widehat{\mathbf V}_i\|_2^2
   \]
   yielding shared neural features $\mathbf F\in\mathbb R^{N\times d}$, with $d=768$.

2. **Subject–Object Disentanglement**  
   Each patch embedding is factorized via a learned orthonormal basis $\mathbf B=[\mathbf B_{\rm subj}\;\mathbf B_{\rm obj}]\in\mathbb R^{d\times d}$, with explicit decomposition:
   \[
   \mathbf z = \mathbf B^\top\mathbf f = (\mathbf z_{\rm subj},\mathbf z_{\rm obj})
   \]
   where $\mathbf z_{\rm subj}\in\mathbb R^{d_{\rm subj}}$ and $\mathbf z_{\rm obj}\in\mathbb R^{d_{\rm obj}}$, $d_{\rm subj}+d_{\rm obj}=d$.

3. **Dual Decoding: Biometric and Semantic**  
   - Biometric (subject ID): Pool $\mathbf Z_{\rm subj}$, apply linear classifier to output logits for $S$ subjects.
   - Semantic (object classification): Freeze CLIP visual features; cross-attend CLIP tokens with $\mathbf Z_{\rm obj}$ via multi-head attention, and classify pooled outputs.

The objective function in the dual decoding phase combines:
\[
\mathcal L = \mathcal L_{\rm subj} + \mathcal L_{\rm obj} + \lambda\,\mathcal L_{\rm orth}
\]
with classification (cross-entropy or binary cross-entropy) for subject and object, and orthonormality regularization on $\mathbf B$ ($\lambda=0.1$) [2509.17313].

## 4. Empirical Evaluation and Comparative Results

**InMind (LLM):**  
The primary case study is on Avalon (6-player, Mandarin, transcribed voice chat). Dataset statistics: 30 sessions, 884 utterances, 160 traces, and 30 reflections. Quantitative results highlight several constraints:

- Player identification: General LLMs achieve top-1 accuracy near baseline (0.16 for GPT-4o), DeepSeek-R1 marginally higher (0.24). Top-3 accuracy remains modest. BERT-based cosine matching performs comparably, indicating frequent reliance on surface lexical cues.
- Reflection alignment: With explicit trace input, models reach $\sim$80% accuracy; without, accuracy drops to $\sim$30%, showing strong anchoring dependency.
- Trace attribution: Incremental gains from prior-trace context are small or negative (e.g., $\Delta=+0.008$ GPT-4o), indicating limited true adaptive reasoning.
- Role inference: Strict match accuracy is $\sim$30–40%, relaxed grouping $\sim$60–70%. Reasoning-enhanced models outperform general-purpose LLMs, but performance remains far from ceiling [2508.16072].

**$i$MIND (Neural Decoding):**  
Utilizing the NSD dataset (8 subjects, 10,000 images/subject, $69,\!566$ train, $7,674$ test), $i$MIND delivers:

| Method                | mAP   | AUC    | Hamming | Subject-ID ACC | Generalization (mAP) |
|-----------------------|-------|--------|---------|---------------|----------------------|
| Single-subject ViT/MLP| .24–.26| .82–.85| —       | —             | —                    |
| CLIP-MUSED            | .258  | .877   | —       | —             | —                    |
| $i$MIND (fMRI only)   | .310  | .913   | .027    | .999          | .784 (holdout)       |
| $i$MIND (fMRI+image)  | .784  | .984   | .012    | .999          | .790 (all)           |

This demonstrates state-of-the-art performance, elimination of scalability limits, and robust cross-subject generalization [2509.17313].

## 5. Key Insights and Theoretical Contributions

**InMind (LLM):**
- Standard LLMs default to lexical mimicry and shallow pattern recognition rather than temporally consistent, individualized reasoning. Reflection and role inference tasks cannot be satisfied unless models receive explicit round-level traces, i.e., temporal anchoring is not learned in the absence of direct cues.
- Reasoning-enhanced models, notably DeepSeek-R1, manifest partial style-sensitive reasoning: backward inference, hedging/certainty modulation, and context-consistent assignments improve, but absolute scores remain low. Performance gains under grouped scoring suggest that coarse-grained cognitive traits are more easily aligned than precise role attribution.

**$i$MIND (Neural Decoding):**
- Produces interpretable voxel–object activation fingerprints by tracing Grad-CAM attributions through ViT layers, yielding visualization of region-selective activations (e.g., consistent ventral stream responses to “horse”/“bird,” strong social stimulus activation for “person”).
- Reveals clear subject-specific attention dynamics during rapid (3 s) visual exposure: shared and residual attention maps highlight both universally salient object detection (e.g., “chair”) and idiosyncratic focal patterns (e.g., “cup” receiving elevated attention by specific subjects correlated with high recognition probability).
- Clustering voxels by mean and standard deviation of activation exposes functional specialization: “bystanders,” “discriminators,” and “supporters” for semantic decoding roles [2509.17313].

## 6. Limitations and Prospects

**LLM Evaluation:**
- Current LLMs lack the ability to internalize individualized reasoning styles without explicit temporal and strategic annotation. Dynamic adaptation remains shallow, with models frequently treating sequential rounds as independent.
- Future work in InMind aims to scale to diverse SDGs, automate strategy profile induction to mitigate annotation bias, and integrate memory/belief tracking modules to maintain cross-round coherence. Extension to cooperation and negotiation scenarios is proposed as crucial domains for contextually adaptive inference [2508.16072].

**Neural Decoding:**
- $i$MIND establishes a foundation for more interpretable and generalizable neural decoding, but practical constraints (e.g., requirement of large fMRI datasets, pre-defined object basis dimensionality) persist.
- A plausible implication is that learned subject–object disentanglement and interpretability frameworks could inform future low-shot or online neural decoding domains, and illuminate the neural basis of visual and attentional idiosyncrasy at population scale [2509.17313].

**References:**  
[2508.16072] InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles  
[2509.17313] $i$MIND: Insightful Multi-subject Invariant Neural Decoding

Source: https://www.emergentmind.com/topics/inmind