Mind: From Subjective Experience to Operational AI
- Mind is defined as the spectrum from subjective experience to AI computational models, linking the neural basis of qualia with decision support systems.
- It investigates the mapping X: B -> E, addressing the hard problem of relating neural activity to conscious phenomena through empirical and computational frameworks.
- Practical implementations span clinical diagnostics, multimodal reasoning, and embodied control, showcasing measurable performance gains in each domain.
Mind denotes, in one major research lineage, the private, personal, first-person domain of subjective experience, and, in another, a family of technical abstractions and systems used for decoding, reasoning, simulation, negotiation, and decision support. In Feldman’s formulation, mind is identified with subjective experience (SE), the “what it is like” aspect of seeing a color, feeling pain, or tasting sweetness, and the central unresolved problem is the mapping from bodily—especially neural—activity to SE (Feldman, 2018). In parallel, recent arXiv literature uses MIND as an acronym for methods and benchmarks in digital communications, multimodal reasoning, psychiatry, world models, humanoid control, and materials research, indicating that the term now spans both foundational and highly operational meanings (Jiang et al., 2019).
1. Subjective experience and the science of mind
Feldman’s “Science of Mind” is organized around four commitments: scientific realism of mind, agnostic mysterianism, careful attention to language, and concentration on the unknown mapping from neural activity to subjective experience (Feldman, 2018). Scientific realism here means treating subjective experience as ontologically real in the same sense that earlier science treated atoms and subatomic particles as theoretical posits later constrained by experiment. The target phenomenon is not generic “consciousness” in its many ordinary-language senses, but SE as phenomenology or “what it is like” to feel, see, and sense.
Within this framework, the Hard Problem is the explanatory gap between objective neural description and subjective character. Feldman’s proposal is neither reductionist dismissal nor non-physical invocation, but “agnostic mysterianism”: there really is an explanatory gap; science has demystified many previously intractable phenomena; and there is no reason either to deny the reality of SE or to assume the gap is permanently unbridgeable (Feldman, 2018). A plausible implication is that the framework treats ignorance as a research constraint rather than as a terminal philosophical verdict.
The paper therefore recommends a restricted technical vocabulary. “Mind” is private first-person subjective experience; “qualia” is better replaced by SE; and the Dehaene-style taxonomy C0/C1/C2 distinguishes non-experiential neural processes, report-accessible processes, and meta-cognitive monitoring. Feldman also introduces “actionability” as an internally computed measure of how well potential actions contribute to fitness, and (Chi) as the yet-unknown mapping from bodily states to SE. The central formal object is
where is a complete description of bodily activity and is the corresponding pattern of SE (Feldman, 2018).
Research is then organized around “touchstone” problems that any candidate must explain. The examples given are the broken-grid phenomenon, phi and apparent motion, feature binding, and the stable visual world despite saccades and interruptions in C0 processes. Feldman also points to experimental results “modulo ”: synesthesia as a reliable stimulus–experience mapping, postdictive TMS effects in which a later flash fills in an earlier scotoma, border ownership in V1/V2 with spike synchrony shaped by top-down grouping, and Chang and Tsao’s 50-dimensional feature space for face identity (Feldman, 2018). These are not solutions to the Hard Problem, but constrained empirical fragments of the brain-to-experience relation.
2. Formal models of cognition, concepts, and consciousness
A different line of work attempts to formalize mind mathematically. Perlovsky’s “Physics of the mind” argues that earlier logical approaches to intelligence ran into two obstacles: Gödel’s incompleteness and combinatorial complexity. Dynamic logic (DL) is proposed as an alternative in which mental models start vague and progressively become crisp as they fit bottom-up data (Perlovsky, 2010). The core quantities are the model–data similarity , normalized association variables , and a global similarity functional
which functions as a measure of knowledge.
In this account, perception and cognition are iterative self-organization rather than static logical classification. Vague top-down activations are gradually sharpened through recurrent interaction with sensory input; conscious awareness appears only after sufficient crispness. Perlovsky further defines the “knowledge instinct” as an inborn drive to maximize , with emotional signal proportional to changes in similarity over time, and extends the framework upward to concept formation, instincts, imagination, intuition, consciousness thresholds, and a dual hierarchy linking language and cognition (Perlovsky, 2010). The same review connects aesthetic emotions and music to cognitive integration, while also noting open questions at the top of the cognitive hierarchy and the absence of direct neurobiological demonstration of exact DL updates.
Panigrahy and Zhang propose a different abstraction: the mind as a place where a circuit grows from primitive components through repeated experience (Panigrahy et al., 2012). Concepts are represented as functions whose inputs and outputs are themselves concepts or percepts; the concept graph records invocation relations; and new concepts arise by composition, such as
0
Weights track frequency of invocation, and repeated experience drives bottom-up circuit accretion.
The guiding heuristic is compression. Each concept has an implementation length, and the system prefers compositions minimizing total description length as an upper bound on Kolmogorov complexity. This yields an explicitly compositional, human-readable alternative to distributed neural-network representations, but the authors also note limitations: no clear differentiable training algorithm, possible combinatorial explosion in composition search, and no formal theorems or proofs in the original paper (Panigrahy et al., 2012). Together, DL and growing-circuit models illustrate two recurrent strategies in formal mind research: vague-to-crisp dynamical refinement and incremental concept composition.
3. Theory of Mind, collective inference, and social interpretation
In contemporary AI, “mind” often appears in the narrower sense of Theory of Mind (ToM): inference over other agents’ hidden beliefs, desires, or intentions from observations and actions. Aru et al. formalize a ToM task as inferring hidden mental state 1 from observations 2, then predicting action 3, but argue that many claimed successes of deep-learning ToM systems can be explained by shortcuts induced by narrow task design (Aru et al., 2022). The paper catalogues failure modes in perspective-taking gridworlds, false-belief tasks, game-theoretic environments, and Hanabi, and recommends more complex open-ended environments together with interpretability tools such as feature visualization, linear probes, attribution methods, and ablations.
The move from individual ToM to group-level inference appears in the “Theory of Collective Mind” model. Instead of keeping separate ToM models for each partner, ToCM compresses all agents’ observations into a unified but plural collective mental state 4, trained with a variational free-energy objective and rolled forward in an imaginative latent space to support cooperation (Zhao et al., 2023). The model was evaluated in Multi-Agent Particle Environments and SMAC, where it sped up convergence by roughly 5–6, achieved 7 versus 8 in two-agent cooperative navigation, and exceeded MAPPO/QMIX by 9–0 percentage points on SMAC 3s_vs_3z win-rate curves. The same paper reports transfer from a ToCM pretrained on the hardest SMAC map to new maps with substantially faster adaptation.
A separate strand studies mind-state inference from nonverbal communication. Motion2Mind introduces a dataset of 1 clips covering 2 nonverbal cue types and 3 mind states, organized into beliefs, intentions, percepts, desires, knowledge, and emotions (Lee et al., 19 Nov 2025). Expert psychologists obtained 4 on Explanation and 5 on Prediction, whereas the best evaluated VLMs remained around 6–7 on detection and explanation tasks. The benchmark also reports a pronounced over-interpretation bias: false positives on invalid cues substantially outnumber false negatives, indicating that current systems often assign psychological meaning where none is warranted.
Negotiation dialogue introduces yet another operational meaning of mind. In travel planning, MIND models private willingness scores 8, infers opponents’ willingness from linguistic signals in a “Strategic Appraisal” phase, and conditions response strategies on relative stakes (Do et al., 23 Mar 2026). Across 9 inference instances, willingness decoding reached 0 accuracy, with 1 and Pearson 2. Relative to a traditional MAD baseline, the framework improved High-w Hit by 3, Debate Hit-Rate by 4, and achieved LLM-as-a-Judge wins in Rationality (5), Fluency (6), and overall win rate (7). This use of “mind” is explicitly strategic: a latent representation of hidden priorities required for consensus rather than factual correctness.
4. Clinical and therapeutic MIND systems
In clinical informatics, MIND is used for systems that organize heterogeneous patient data into decision-support structures. The “Multimodal data Integrated Narrative Dashboard” combines clinical notes, self-report surveys, and passive sensing streams such as sleep, steps, screen time, and location into a text-first dashboard for mental healthcare (Zou et al., 21 Jan 2026). Its pipeline comprises Session Recap, Guided Patient Data Insights, and Exploratory Patient Data Insights; an Analyzer with Inquirer, Planner, and Discoverer modules; a Synthesizer; and a Narrator that rewrites selected facts into a final narrative template. The design emerged from co-design sessions with five clinicians and was evaluated in a within-subject study of 8 against a FACT baseline. Reported gains include hidden insights 9 versus 0 (1) and decision support 2 versus 3 (4), with no significant workload differences on NASA-TLX.
Psychiatric consultation imposes a more stringent sequential decision problem. The RL-based MIND framework models consultation as an MDP with dialogue history 5, a compressed clinical retrieval state 6, question templates grounded in DSM/ICD criteria, and terminal diagnosis actions (Li et al., 4 Mar 2026). Its Criteria-Grounded Psychiatric Reasoning Bank stores triplets 7 of compressed consultation state, clinician-crafted support note, and reliability score; retrieval is gated by a similarity-and-quality score; and the policy is trained with rubric-based process rewards, retrieval shaping, information-gain reward, operational penalties, and value-aware trajectory rectification. On PsySim-Std, MIND-8B reached 8 diagnostic accuracy versus 9 for DoctorAgent-RL and 0 for Qwen3-8B1; on PsySim-Adapt it reached 2 versus 3 for DDO and 4 for DoctorAgent-RL. The paper also reports 5 average turns to diagnosis versus 6 for baselines, Empathy gains of 7 points, Naturalness gains of 8 points, and support faithfulness 9 versus 0 and 1.
A different therapeutic use appears in Multi-agent INner Dialogue for psychological healing. This MIND instantiates four role-specific agents—Trigger, Devil, Guide, and Strategist—plus a Human-Simulator, arranged in a loop that generates scenarios, cognitive distortions, guidance with memory updates, storyline planning, and user-comforting responses (Chen et al., 27 Feb 2025). The implementation uses prompt engineering rather than gradient-based fine-tuning, with temperature 2, and was evaluated on themes sampled from the C2D2 dataset. Gemini-2.0-flash was chosen as the Human-Simulator after a clinician role-playing evaluation. In paradigm comparison, MIND achieved approximately 3 on Immersion, Coherence, Engagement, Emotional Relief, Satisfaction, and Interest, including a 4 relative boost in Engagement over the best baseline and perfect 5 in Satisfaction and Interest. The paper’s ablations report drops greater than 6 points when Guide, Strategist, or memory is removed.
5. Decoding, control, and world-model evaluation
In digital communications, MIND first appears as Model Independent Neural Decoder, a meta-learned extension of neural convolutional and turbo decoders (Jiang et al., 2019). The core problem is fast adaptation under channel variation without explicit run-time channel models. MIND treats each channel as a task in Model-Agnostic Meta-Learning, learns a parameter initialization 7 from archetypal channels such as AWGN, additive 8-distributed noise, and radar/impulsive noise, and adapts to new channels with minimal pilot data. The reported implementation uses a 2-layer bidirectional GRU with 9 hidden units per direction per layer, block length 0, meta-batch size 1, and 2 meta-updates. At test time, MIND-1 uses 3, whereas full fine-tuning requires 4. On convolutional codes, MIND-1 cut BER nearly in half at SNR 5 dB, from 6 to 7, and remained within 8 dB of a specialized neural decoder; MIND-10 effectively matched specialized performance.
A second communications use is Maximum Mutual Information based Neural Decoder, which makes mutual information itself the decoding criterion (Tonello et al., 2022). Here the objective is to maximize
9
or equivalently minimize a-posteriori uncertainty. Because 0 is generally unknown, the method trains a discriminator to estimate the density ratio 1. The paper reports two implementations, one unsupervised over joint 2 pairs and one supervised over finite codebooks. On 4-PAM with non-uniform source over AWGN, the estimated source entropy, conditional entropy, and average MI were within 3 bits of the true values; on short-block codes over truncated Middleton noise, MIND nearly reached the genie bound and clearly outperformed classical ML decoding.
In embodied control, MIND denotes Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control (Li et al., 25 May 2026). The framework is fully end-to-end and introduces a latent “behavioral intent” representation between text and low-level actions. Its three jointly trained components are the Holistic Intent Predictor, Immediate Intent Predictor, and Action Diffusion Transformer; a 1D-causal-convolutional VAE encodes humanoid state windows into a latent space with 4 and 5. Using IsaacGym, a 24-joint SMPL-like humanoid, and HumanML3D motions retargeted through a PHC tracker, the method outperformed PDP, UniPhys, CLoSD, and Kimodo++. Reported gains include R-Precision@1 6 versus 7, FID 8 versus 9, MM-Dist 0 versus 1, foot-floating 2 mm versus 3 mm, and jerk 4 versus 5.
For generative world models, MIND is a benchmark rather than a controller. It introduces 6 videos at 7p and 8 FPS, with 9 first-person plus 00 third-person clips under a shared action space and 01 clips across varied action spaces in eight scene categories (Ye et al., 8 Feb 2026). The benchmark evaluates Long-Context Memory, Generated Scene Consistency, Action Accuracy via translational and rotational RPE, and Action-Space Generalization. It also provides MIND-World, an interactive Video-to-World baseline initialized from SkyReels-V2-I2V-1.3B. On mindM-First 50, video memory improved LCM from 02 to 03 and GSC from 04 to 05; on mindM-Third 50, MIND-World outperformed Matrix-Game 2.0 across LCM, ASG, aesthetic quality, imaging quality, and RPE metrics. The benchmark’s main claim is diagnostic rather than definitive: current world models still struggle with long-term memory consistency and action-space generalization.
6. Reasoning, distillation, and automated scientific inquiry
A major recent usage of MIND concerns active reasoning in language and multimodal models. Capability-aware Multi-Perspective CoT Distillation replaces single-rationale imitation with a Teaching Assistant network that scores multiple teacher reasoning styles, aligns them to the student’s evolving capacity through Feedback-Driven Inertia Calibration, and trains the student with preference-weighted SFT plus pairwise consistency regularization (Cui et al., 7 Jan 2026). The corpus contains 06 perspectives generated by Qwen3-235B and filtered to 07 high-quality samples. For a Qwen2.5-7B student, MIND achieved 08 on MATH500, 09 on GSM8K, 10 on SVAMP, 11 on CommonsenseQA, 12 on StrategyQA, and 13 on GPQA-Diamond, outperforming zero-shot CoT, SbS-KD, MCC-KD, MoDE-CoTD, EDIT, and a no-fusion variant. The latent-space analysis in the paper further reports that MIND-distilled students preserve multiple separable reasoning clusters rather than collapsing onto one or two dominant modes.
Multi-rationale INtegrated Discriminative reasoning extends the same general agenda to multimodal large models (Yu et al., 5 Dec 2025). Its RAD paradigm augments datasets with multiple positive and negative rationales; P2CL organizes training into positive learning followed by active logic discrimination and correction; and MCA enforces semantic aggregation of correct reasoning and boundary separation of incorrect reasoning. The paper reports ScienceQA 14 versus 15 for Multimodal-CoT on the base model, A-OKVQA 16 versus 17, and 18CoT 19 versus 20 on the base model and 21 versus 22 on the large model. Ablations on ScienceQA23 show 24 for the baseline, 25 with MCA only, 26 with P2CL only, and 27 for the full system.
Meta-learning for In-context Deduction uses MIND to denote a few-shot meta-learning fine-tuning strategy for selecting the unique minimal premise set that entails a syllogistic hypothesis from a knowledge base (Bertolazzi et al., 20 May 2025). Each episode consists of a KB, a support set of three study examples, and one query; the model performs one inner update on support examples and is optimized on the query. With LoRA on Qwen-2.5 models, MIND improved exact-match accuracy from 28 to 29 on Qwen-1.5B, from 30 to 31 on Qwen-3B, and from 32 to 33 on Qwen-7B. The paper also reports stronger length-OOD generalization than the non-meta baseline and notes that even the smallest MIND model exceeded GPT-4o and matched or surpassed o3-mini on the task.
Finally, MIND appears as an AI co-scientist for materials research. This system decomposes automated hypothesis validation into Pre-Experiment, Experiment, and Discussion stages within a LangGraph multi-agent pipeline, uses SevenNet-Omni for scalable in-silico experiments, and exposes the workflow through a Streamlit interface (Ahn et al., 15 Apr 2026). The benchmark consists of 34 real-world hypotheses across energetic, structural, and mechanical categories. Reported performance is 35 overall accuracy (36 correct), with 37 on energetic, 38 on structural, and 39 on mechanical cases, at an average runtime of 40 minutes and 41–42 speedup over human-in-the-loop SevenNet-Omni use. A 43-participant user study rated scientific validity 44, reasoning transparency 45, and research usefulness 46.
Across these literatures, “mind” does not denote a single object. It names subjective experience in foundational philosophy of neuroscience, dynamical and compositional models of cognition, latent mental-state inference in multi-agent systems, clinically grounded narrative and diagnostic infrastructures, and a proliferating class of domain-specific AI systems whose core design problem is hidden-state representation, adaptation, or reasoning. This suggests a persistent structural theme: whether the target is qualia, belief, intent, willingness, action control, or diagnosis, mind is repeatedly operationalized as a mapping from incomplete observables to consequential internal structure.