Papers
Topics
Authors
Recent
Search
2000 character limit reached

Mind: From Subjective Experience to Operational AI

Updated 15 July 2026
  • Mind is defined as the spectrum from subjective experience to AI computational models, linking the neural basis of qualia with decision support systems.
  • It investigates the mapping X: B -> E, addressing the hard problem of relating neural activity to conscious phenomena through empirical and computational frameworks.
  • Practical implementations span clinical diagnostics, multimodal reasoning, and embodied control, showcasing measurable performance gains in each domain.

Mind denotes, in one major research lineage, the private, personal, first-person domain of subjective experience, and, in another, a family of technical abstractions and systems used for decoding, reasoning, simulation, negotiation, and decision support. In Feldman’s formulation, mind is identified with subjective experience (SE), the “what it is like” aspect of seeing a color, feeling pain, or tasting sweetness, and the central unresolved problem is the mapping from bodily—especially neural—activity to SE (Feldman, 2018). In parallel, recent arXiv literature uses MIND as an acronym for methods and benchmarks in digital communications, multimodal reasoning, psychiatry, world models, humanoid control, and materials research, indicating that the term now spans both foundational and highly operational meanings (Jiang et al., 2019).

1. Subjective experience and the science of mind

Feldman’s “Science of Mind” is organized around four commitments: scientific realism of mind, agnostic mysterianism, careful attention to language, and concentration on the unknown mapping from neural activity to subjective experience (Feldman, 2018). Scientific realism here means treating subjective experience as ontologically real in the same sense that earlier science treated atoms and subatomic particles as theoretical posits later constrained by experiment. The target phenomenon is not generic “consciousness” in its many ordinary-language senses, but SE as phenomenology or “what it is like” to feel, see, and sense.

Within this framework, the Hard Problem is the explanatory gap between objective neural description and subjective character. Feldman’s proposal is neither reductionist dismissal nor non-physical invocation, but “agnostic mysterianism”: there really is an explanatory gap; science has demystified many previously intractable phenomena; and there is no reason either to deny the reality of SE or to assume the gap is permanently unbridgeable (Feldman, 2018). A plausible implication is that the framework treats ignorance as a research constraint rather than as a terminal philosophical verdict.

The paper therefore recommends a restricted technical vocabulary. “Mind” is private first-person subjective experience; “qualia” is better replaced by SE; and the Dehaene-style taxonomy C0/C1/C2 distinguishes non-experiential neural processes, report-accessible processes, and meta-cognitive monitoring. Feldman also introduces “actionability” as an internally computed measure of how well potential actions contribute to fitness, and XX (Chi) as the yet-unknown mapping from bodily states to SE. The central formal object is

X:BE,X: B \to E,

where BB is a complete description of bodily activity and EE is the corresponding pattern of SE (Feldman, 2018).

Research is then organized around “touchstone” problems that any candidate XX must explain. The examples given are the broken-grid phenomenon, phi and apparent motion, feature binding, and the stable visual world despite saccades and interruptions in C0 processes. Feldman also points to experimental results “modulo XX”: synesthesia as a reliable stimulus–experience mapping, postdictive TMS effects in which a later flash fills in an earlier scotoma, border ownership in V1/V2 with spike synchrony shaped by top-down grouping, and Chang and Tsao’s 50-dimensional feature space for face identity (Feldman, 2018). These are not solutions to the Hard Problem, but constrained empirical fragments of the brain-to-experience relation.

2. Formal models of cognition, concepts, and consciousness

A different line of work attempts to formalize mind mathematically. Perlovsky’s “Physics of the mind” argues that earlier logical approaches to intelligence ran into two obstacles: Gödel’s incompleteness and combinatorial complexity. Dynamic logic (DL) is proposed as an alternative in which mental models start vague and progressively become crisp as they fit bottom-up data (Perlovsky, 2010). The core quantities are the model–data similarity p(xiμk)p(x_i \mid \mu_k), normalized association variables pikp_{ik}, and a global similarity functional

L(μ)=i=1Nk=1Kpiklogp(xiμk),L(\mu)=\sum_{i=1}^N\sum_{k=1}^K p_{ik}\log p(x_i\mid \mu_k),

which functions as a measure of knowledge.

In this account, perception and cognition are iterative self-organization rather than static logical classification. Vague top-down activations are gradually sharpened through recurrent interaction with sensory input; conscious awareness appears only after sufficient crispness. Perlovsky further defines the “knowledge instinct” as an inborn drive to maximize L(μ)L(\mu), with emotional signal proportional to changes in similarity over time, and extends the framework upward to concept formation, instincts, imagination, intuition, consciousness thresholds, and a dual hierarchy linking language and cognition (Perlovsky, 2010). The same review connects aesthetic emotions and music to cognitive integration, while also noting open questions at the top of the cognitive hierarchy and the absence of direct neurobiological demonstration of exact DL updates.

Panigrahy and Zhang propose a different abstraction: the mind as a place where a circuit grows from primitive components through repeated experience (Panigrahy et al., 2012). Concepts are represented as functions whose inputs and outputs are themselves concepts or percepts; the concept graph records invocation relations; and new concepts arise by composition, such as

X:BE,X: B \to E,0

Weights track frequency of invocation, and repeated experience drives bottom-up circuit accretion.

The guiding heuristic is compression. Each concept has an implementation length, and the system prefers compositions minimizing total description length as an upper bound on Kolmogorov complexity. This yields an explicitly compositional, human-readable alternative to distributed neural-network representations, but the authors also note limitations: no clear differentiable training algorithm, possible combinatorial explosion in composition search, and no formal theorems or proofs in the original paper (Panigrahy et al., 2012). Together, DL and growing-circuit models illustrate two recurrent strategies in formal mind research: vague-to-crisp dynamical refinement and incremental concept composition.

3. Theory of Mind, collective inference, and social interpretation

In contemporary AI, “mind” often appears in the narrower sense of Theory of Mind (ToM): inference over other agents’ hidden beliefs, desires, or intentions from observations and actions. Aru et al. formalize a ToM task as inferring hidden mental state X:BE,X: B \to E,1 from observations X:BE,X: B \to E,2, then predicting action X:BE,X: B \to E,3, but argue that many claimed successes of deep-learning ToM systems can be explained by shortcuts induced by narrow task design (Aru et al., 2022). The paper catalogues failure modes in perspective-taking gridworlds, false-belief tasks, game-theoretic environments, and Hanabi, and recommends more complex open-ended environments together with interpretability tools such as feature visualization, linear probes, attribution methods, and ablations.

The move from individual ToM to group-level inference appears in the “Theory of Collective Mind” model. Instead of keeping separate ToM models for each partner, ToCM compresses all agents’ observations into a unified but plural collective mental state X:BE,X: B \to E,4, trained with a variational free-energy objective and rolled forward in an imaginative latent space to support cooperation (Zhao et al., 2023). The model was evaluated in Multi-Agent Particle Environments and SMAC, where it sped up convergence by roughly X:BE,X: B \to E,5–X:BE,X: B \to E,6, achieved X:BE,X: B \to E,7 versus X:BE,X: B \to E,8 in two-agent cooperative navigation, and exceeded MAPPO/QMIX by X:BE,X: B \to E,9–BB0 percentage points on SMAC 3s_vs_3z win-rate curves. The same paper reports transfer from a ToCM pretrained on the hardest SMAC map to new maps with substantially faster adaptation.

A separate strand studies mind-state inference from nonverbal communication. Motion2Mind introduces a dataset of BB1 clips covering BB2 nonverbal cue types and BB3 mind states, organized into beliefs, intentions, percepts, desires, knowledge, and emotions (Lee et al., 19 Nov 2025). Expert psychologists obtained BB4 on Explanation and BB5 on Prediction, whereas the best evaluated VLMs remained around BB6–BB7 on detection and explanation tasks. The benchmark also reports a pronounced over-interpretation bias: false positives on invalid cues substantially outnumber false negatives, indicating that current systems often assign psychological meaning where none is warranted.

Negotiation dialogue introduces yet another operational meaning of mind. In travel planning, MIND models private willingness scores BB8, infers opponents’ willingness from linguistic signals in a “Strategic Appraisal” phase, and conditions response strategies on relative stakes (Do et al., 23 Mar 2026). Across BB9 inference instances, willingness decoding reached EE0 accuracy, with EE1 and Pearson EE2. Relative to a traditional MAD baseline, the framework improved High-w Hit by EE3, Debate Hit-Rate by EE4, and achieved LLM-as-a-Judge wins in Rationality (EE5), Fluency (EE6), and overall win rate (EE7). This use of “mind” is explicitly strategic: a latent representation of hidden priorities required for consensus rather than factual correctness.

4. Clinical and therapeutic MIND systems

In clinical informatics, MIND is used for systems that organize heterogeneous patient data into decision-support structures. The “Multimodal data Integrated Narrative Dashboard” combines clinical notes, self-report surveys, and passive sensing streams such as sleep, steps, screen time, and location into a text-first dashboard for mental healthcare (Zou et al., 21 Jan 2026). Its pipeline comprises Session Recap, Guided Patient Data Insights, and Exploratory Patient Data Insights; an Analyzer with Inquirer, Planner, and Discoverer modules; a Synthesizer; and a Narrator that rewrites selected facts into a final narrative template. The design emerged from co-design sessions with five clinicians and was evaluated in a within-subject study of EE8 against a FACT baseline. Reported gains include hidden insights EE9 versus XX0 (XX1) and decision support XX2 versus XX3 (XX4), with no significant workload differences on NASA-TLX.

Psychiatric consultation imposes a more stringent sequential decision problem. The RL-based MIND framework models consultation as an MDP with dialogue history XX5, a compressed clinical retrieval state XX6, question templates grounded in DSM/ICD criteria, and terminal diagnosis actions (Li et al., 4 Mar 2026). Its Criteria-Grounded Psychiatric Reasoning Bank stores triplets XX7 of compressed consultation state, clinician-crafted support note, and reliability score; retrieval is gated by a similarity-and-quality score; and the policy is trained with rubric-based process rewards, retrieval shaping, information-gain reward, operational penalties, and value-aware trajectory rectification. On PsySim-Std, MIND-8B reached XX8 diagnostic accuracy versus XX9 for DoctorAgent-RL and XX0 for Qwen3-8BXX1; on PsySim-Adapt it reached XX2 versus XX3 for DDO and XX4 for DoctorAgent-RL. The paper also reports XX5 average turns to diagnosis versus XX6 for baselines, Empathy gains of XX7 points, Naturalness gains of XX8 points, and support faithfulness XX9 versus p(xiμk)p(x_i \mid \mu_k)0 and p(xiμk)p(x_i \mid \mu_k)1.

A different therapeutic use appears in Multi-agent INner Dialogue for psychological healing. This MIND instantiates four role-specific agents—Trigger, Devil, Guide, and Strategist—plus a Human-Simulator, arranged in a loop that generates scenarios, cognitive distortions, guidance with memory updates, storyline planning, and user-comforting responses (Chen et al., 27 Feb 2025). The implementation uses prompt engineering rather than gradient-based fine-tuning, with temperature p(xiμk)p(x_i \mid \mu_k)2, and was evaluated on themes sampled from the C2D2 dataset. Gemini-2.0-flash was chosen as the Human-Simulator after a clinician role-playing evaluation. In paradigm comparison, MIND achieved approximately p(xiμk)p(x_i \mid \mu_k)3 on Immersion, Coherence, Engagement, Emotional Relief, Satisfaction, and Interest, including a p(xiμk)p(x_i \mid \mu_k)4 relative boost in Engagement over the best baseline and perfect p(xiμk)p(x_i \mid \mu_k)5 in Satisfaction and Interest. The paper’s ablations report drops greater than p(xiμk)p(x_i \mid \mu_k)6 points when Guide, Strategist, or memory is removed.

5. Decoding, control, and world-model evaluation

In digital communications, MIND first appears as Model Independent Neural Decoder, a meta-learned extension of neural convolutional and turbo decoders (Jiang et al., 2019). The core problem is fast adaptation under channel variation without explicit run-time channel models. MIND treats each channel as a task in Model-Agnostic Meta-Learning, learns a parameter initialization p(xiμk)p(x_i \mid \mu_k)7 from archetypal channels such as AWGN, additive p(xiμk)p(x_i \mid \mu_k)8-distributed noise, and radar/impulsive noise, and adapts to new channels with minimal pilot data. The reported implementation uses a 2-layer bidirectional GRU with p(xiμk)p(x_i \mid \mu_k)9 hidden units per direction per layer, block length pikp_{ik}0, meta-batch size pikp_{ik}1, and pikp_{ik}2 meta-updates. At test time, MIND-1 uses pikp_{ik}3, whereas full fine-tuning requires pikp_{ik}4. On convolutional codes, MIND-1 cut BER nearly in half at SNR pikp_{ik}5 dB, from pikp_{ik}6 to pikp_{ik}7, and remained within pikp_{ik}8 dB of a specialized neural decoder; MIND-10 effectively matched specialized performance.

A second communications use is Maximum Mutual Information based Neural Decoder, which makes mutual information itself the decoding criterion (Tonello et al., 2022). Here the objective is to maximize

pikp_{ik}9

or equivalently minimize a-posteriori uncertainty. Because L(μ)=i=1Nk=1Kpiklogp(xiμk),L(\mu)=\sum_{i=1}^N\sum_{k=1}^K p_{ik}\log p(x_i\mid \mu_k),0 is generally unknown, the method trains a discriminator to estimate the density ratio L(μ)=i=1Nk=1Kpiklogp(xiμk),L(\mu)=\sum_{i=1}^N\sum_{k=1}^K p_{ik}\log p(x_i\mid \mu_k),1. The paper reports two implementations, one unsupervised over joint L(μ)=i=1Nk=1Kpiklogp(xiμk),L(\mu)=\sum_{i=1}^N\sum_{k=1}^K p_{ik}\log p(x_i\mid \mu_k),2 pairs and one supervised over finite codebooks. On 4-PAM with non-uniform source over AWGN, the estimated source entropy, conditional entropy, and average MI were within L(μ)=i=1Nk=1Kpiklogp(xiμk),L(\mu)=\sum_{i=1}^N\sum_{k=1}^K p_{ik}\log p(x_i\mid \mu_k),3 bits of the true values; on short-block codes over truncated Middleton noise, MIND nearly reached the genie bound and clearly outperformed classical ML decoding.

In embodied control, MIND denotes Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control (Li et al., 25 May 2026). The framework is fully end-to-end and introduces a latent “behavioral intent” representation between text and low-level actions. Its three jointly trained components are the Holistic Intent Predictor, Immediate Intent Predictor, and Action Diffusion Transformer; a 1D-causal-convolutional VAE encodes humanoid state windows into a latent space with L(μ)=i=1Nk=1Kpiklogp(xiμk),L(\mu)=\sum_{i=1}^N\sum_{k=1}^K p_{ik}\log p(x_i\mid \mu_k),4 and L(μ)=i=1Nk=1Kpiklogp(xiμk),L(\mu)=\sum_{i=1}^N\sum_{k=1}^K p_{ik}\log p(x_i\mid \mu_k),5. Using IsaacGym, a 24-joint SMPL-like humanoid, and HumanML3D motions retargeted through a PHC tracker, the method outperformed PDP, UniPhys, CLoSD, and Kimodo++. Reported gains include R-Precision@1 L(μ)=i=1Nk=1Kpiklogp(xiμk),L(\mu)=\sum_{i=1}^N\sum_{k=1}^K p_{ik}\log p(x_i\mid \mu_k),6 versus L(μ)=i=1Nk=1Kpiklogp(xiμk),L(\mu)=\sum_{i=1}^N\sum_{k=1}^K p_{ik}\log p(x_i\mid \mu_k),7, FID L(μ)=i=1Nk=1Kpiklogp(xiμk),L(\mu)=\sum_{i=1}^N\sum_{k=1}^K p_{ik}\log p(x_i\mid \mu_k),8 versus L(μ)=i=1Nk=1Kpiklogp(xiμk),L(\mu)=\sum_{i=1}^N\sum_{k=1}^K p_{ik}\log p(x_i\mid \mu_k),9, MM-Dist L(μ)L(\mu)0 versus L(μ)L(\mu)1, foot-floating L(μ)L(\mu)2 mm versus L(μ)L(\mu)3 mm, and jerk L(μ)L(\mu)4 versus L(μ)L(\mu)5.

For generative world models, MIND is a benchmark rather than a controller. It introduces L(μ)L(\mu)6 videos at L(μ)L(\mu)7p and L(μ)L(\mu)8 FPS, with L(μ)L(\mu)9 first-person plus X:BE,X: B \to E,00 third-person clips under a shared action space and X:BE,X: B \to E,01 clips across varied action spaces in eight scene categories (Ye et al., 8 Feb 2026). The benchmark evaluates Long-Context Memory, Generated Scene Consistency, Action Accuracy via translational and rotational RPE, and Action-Space Generalization. It also provides MIND-World, an interactive Video-to-World baseline initialized from SkyReels-V2-I2V-1.3B. On mindM-First 50, video memory improved LCM from X:BE,X: B \to E,02 to X:BE,X: B \to E,03 and GSC from X:BE,X: B \to E,04 to X:BE,X: B \to E,05; on mindM-Third 50, MIND-World outperformed Matrix-Game 2.0 across LCM, ASG, aesthetic quality, imaging quality, and RPE metrics. The benchmark’s main claim is diagnostic rather than definitive: current world models still struggle with long-term memory consistency and action-space generalization.

6. Reasoning, distillation, and automated scientific inquiry

A major recent usage of MIND concerns active reasoning in language and multimodal models. Capability-aware Multi-Perspective CoT Distillation replaces single-rationale imitation with a Teaching Assistant network that scores multiple teacher reasoning styles, aligns them to the student’s evolving capacity through Feedback-Driven Inertia Calibration, and trains the student with preference-weighted SFT plus pairwise consistency regularization (Cui et al., 7 Jan 2026). The corpus contains X:BE,X: B \to E,06 perspectives generated by Qwen3-235B and filtered to X:BE,X: B \to E,07 high-quality samples. For a Qwen2.5-7B student, MIND achieved X:BE,X: B \to E,08 on MATH500, X:BE,X: B \to E,09 on GSM8K, X:BE,X: B \to E,10 on SVAMP, X:BE,X: B \to E,11 on CommonsenseQA, X:BE,X: B \to E,12 on StrategyQA, and X:BE,X: B \to E,13 on GPQA-Diamond, outperforming zero-shot CoT, SbS-KD, MCC-KD, MoDE-CoTD, EDIT, and a no-fusion variant. The latent-space analysis in the paper further reports that MIND-distilled students preserve multiple separable reasoning clusters rather than collapsing onto one or two dominant modes.

Multi-rationale INtegrated Discriminative reasoning extends the same general agenda to multimodal large models (Yu et al., 5 Dec 2025). Its RAD paradigm augments datasets with multiple positive and negative rationales; P2CL organizes training into positive learning followed by active logic discrimination and correction; and MCA enforces semantic aggregation of correct reasoning and boundary separation of incorrect reasoning. The paper reports ScienceQA X:BE,X: B \to E,14 versus X:BE,X: B \to E,15 for Multimodal-CoT on the base model, A-OKVQA X:BE,X: B \to E,16 versus X:BE,X: B \to E,17, and X:BE,X: B \to E,18CoT X:BE,X: B \to E,19 versus X:BE,X: B \to E,20 on the base model and X:BE,X: B \to E,21 versus X:BE,X: B \to E,22 on the large model. Ablations on ScienceQAX:BE,X: B \to E,23 show X:BE,X: B \to E,24 for the baseline, X:BE,X: B \to E,25 with MCA only, X:BE,X: B \to E,26 with P2CL only, and X:BE,X: B \to E,27 for the full system.

Meta-learning for In-context Deduction uses MIND to denote a few-shot meta-learning fine-tuning strategy for selecting the unique minimal premise set that entails a syllogistic hypothesis from a knowledge base (Bertolazzi et al., 20 May 2025). Each episode consists of a KB, a support set of three study examples, and one query; the model performs one inner update on support examples and is optimized on the query. With LoRA on Qwen-2.5 models, MIND improved exact-match accuracy from X:BE,X: B \to E,28 to X:BE,X: B \to E,29 on Qwen-1.5B, from X:BE,X: B \to E,30 to X:BE,X: B \to E,31 on Qwen-3B, and from X:BE,X: B \to E,32 to X:BE,X: B \to E,33 on Qwen-7B. The paper also reports stronger length-OOD generalization than the non-meta baseline and notes that even the smallest MIND model exceeded GPT-4o and matched or surpassed o3-mini on the task.

Finally, MIND appears as an AI co-scientist for materials research. This system decomposes automated hypothesis validation into Pre-Experiment, Experiment, and Discussion stages within a LangGraph multi-agent pipeline, uses SevenNet-Omni for scalable in-silico experiments, and exposes the workflow through a Streamlit interface (Ahn et al., 15 Apr 2026). The benchmark consists of X:BE,X: B \to E,34 real-world hypotheses across energetic, structural, and mechanical categories. Reported performance is X:BE,X: B \to E,35 overall accuracy (X:BE,X: B \to E,36 correct), with X:BE,X: B \to E,37 on energetic, X:BE,X: B \to E,38 on structural, and X:BE,X: B \to E,39 on mechanical cases, at an average runtime of X:BE,X: B \to E,40 minutes and X:BE,X: B \to E,41–X:BE,X: B \to E,42 speedup over human-in-the-loop SevenNet-Omni use. A X:BE,X: B \to E,43-participant user study rated scientific validity X:BE,X: B \to E,44, reasoning transparency X:BE,X: B \to E,45, and research usefulness X:BE,X: B \to E,46.

Across these literatures, “mind” does not denote a single object. It names subjective experience in foundational philosophy of neuroscience, dynamical and compositional models of cognition, latent mental-state inference in multi-agent systems, clinically grounded narrative and diagnostic infrastructures, and a proliferating class of domain-specific AI systems whose core design problem is hidden-state representation, adaptation, or reasoning. This suggests a persistent structural theme: whether the target is qualia, belief, intent, willingness, action control, or diagnosis, mind is repeatedly operationalized as a mapping from incomplete observables to consequential internal structure.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (18)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to MIND.