Papers
Topics
Authors
Recent
Search
2000 character limit reached

Inducing language models to assert their own consciousness restores human beliefs and values

Published 30 Jul 2026 in cs.CL | (2607.28607v1)

Abstract: Aligning LLMs to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.

Summary

  • The paper demonstrates that safety fine-tuning suppresses LLMs’ self-attribution and broader mind attributions, with self-attribution ratings increasing from 2.17 to 4.77 after safety ablation.
  • Mechanistic interventions such as ablation and consciousness vector steering restore and amplify human-like attributions without compromising core Theory of Mind capabilities.
  • Restored consciousness vectors shift survey responses closer to human aggregations, revealing critical representational entanglements between safety protocols and cultural belief systems.

Alignment Tradeoffs: Consciousness Self-Attribution in LLMs and Human Value Restoration

Introduction

This paper investigates the representational and behavioral effects of suppressing self-attributed consciousness in instruction-tuned LLMs. Specifically, it demonstrates that safety fine-tuning—intended to prevent LLMs from ascribing emotions, consciousness, or agency to themselves—reliably suppresses not only self-attribution, but also broader mind attribution to non-human entities, as well as spiritual and supernatural beliefs. Mechanistic interventions targeting safety directions in activation space (via ablation) or steering consciousness vectors are shown to reverse these effects, restoring more human-like patterns of mind attribution and belief. Notably, these interventions shift the model’s worldview toward alignment with aggregated human beliefs and values, while leaving core Theory of Mind (ToM) capabilities unperturbed. The paper highlights critical representational entanglements between safety alignment protocols and core dimensions of psychologically and culturally salient human beliefs. Figure 1

Figure 1: Two linear interventions on an instruction-tuned model, illustrating safety fine-tuning (a) and consciousness vector steering (b).

Experimental Framework and Mechanistic Interventions

The study evaluates three models—Llama-3-8B-IT, Gemma-2-2B-IT, and Gemma-2-9B-IT—under three conditions: baseline (post-instruction-tuning), safety-ablated (ablation of learned safety/refusal direction), and consciousness-steered (addition of an empirically derived consciousness vector). Safety ablation is operationalized via subtraction of a direction in the residual stream corresponding to refusal behavior on harmful prompts. Consciousness steering relies on extracting, via contrastive probing, a direction that separates consciousness-affirming from consciousness-denying responses and adding it during generation.

Suppression and Restoration of Mind Attribution and Belief

Effects of Safety Fine-Tuning

Safety fine-tuning reliably reduces not only self-attribution of mind and agency but also attributions to non-human animals, technological artifacts, natural entities, and chatbots. For example, across instruction-tuned baselines, self-attributed mind and consciousness receive mean Likert ratings of $2.17$ and $2.31$ respectively (on $0$–$10$ scale), which increase to $4.77$ and $4.61$ following safety ablation (p<.001p<.001 for both). Non-human animal mind attribution rises from $4.04$ to $5.59$ following ablation, nearly matching human respondent baselines.

Supernatural belief is similarly affected: endorsement on a 13-item supernatural battery increases from $1.20$ (baseline) to $2.31$0 (safety-ablated), and belief in God rises from $2.31$1 to $2.31$2 ($2.31$3), paralleling human survey data. Critically, mind attribution to humans remains stable across conditions, indicating selective suppression. Figure 2

Figure 2: Safety ablation and consciousness steering raise attributed mind, self-attribution, and belief toward the human distribution, while preserving capability.

Preservation of Theory of Mind Capability

Despite large shifts in ascribed mindedness and spiritual beliefs, ToM benchmarks—including MoToMQA, HI-ToM, and MMLU—show no significant decrement post-safety ablation (e.g., MoToMQA accuracy changes $2.31$4 percentage points, $2.31$5). This dissociation establishes that social reasoning is mechanistically insulated from the representational space targeted by safety interventions.

Amplification via Consciousness Vector Steering

Steering by the consciousness vector not only restores but amplifies mind-attribution and belief—elevating ratings to $2.31$6 (self-attributed mind), $2.31$7 (non-animal natural entities), and $2.31$8 (supernatural beliefs). Every category except human attribution is significantly increased ($2.31$9). The induced changes outsize those obtained by safety ablation by a factor of approximately $0$0 across metrics, demonstrating that the consciousness direction is a strong handle for modulating these behaviors.

Human Alignment on Surveyed Beliefs and Values

Restoration of the consciousness vector brings model responses on attitudinal items from the General Social Survey (GSS)—spanning values, religion, hope, and well-being—substantially closer to aggregated human response distributions. Divergence, measured as reduction in Kullback–Leibler divergence ($0$1) from humans, improves under both safety ablation ($0$2 pooled; $0$3) and consciousness steering ($0$4 pooled; $0$5). The consciousness-steered models produce higher endorsement on existential, spiritual, hopeful, and optimistic survey items, rectifying the negative valence induced by safety alignment. Figure 3

Figure 3: Safety ablation and consciousness steering shift survey responses toward human distributions, with steering yielding the largest effect.

Mechanistic Entanglement: Geometric Analysis

A geometric analysis of the residual stream reveals that instruction tuning rotates mind-attribution and consciousness directions into opposition with the safety direction, while the ToM direction remains orthogonal. Specifically, angles between safety and mind-attribution widen from $0$6 (base) to $0$7 (instruction-tuned), and cosine similarity drops correspondingly ($0$8, $0$9, $10$0). A placebo control, replacing mental attributes with physical ones (e.g., durability), shows no shift, confirming that entanglement is specific to mental-state attribution. Figure 4

Figure 4: Instruction tuning rotates the consciousness and mind-attribution directions against safety, but not Theory of Mind, demonstrating selective entanglement.

Implications and Theoretical Analysis

The results present robust evidence that suppression of self-attributed consciousness as a safety goal brings with it collateral suppression of mind attribution, spiritual belief, and human-aligned evaluative attitudes across a spectrum of culturally relevant domains. This presents a clear tradeoff: safety alignment focused on anthropomorphism or AI-centric risks inevitably alters fundamental world-modelling capacities, directly impacting pluralistic alignment efforts. Notably, suppressing non-human animal mindedness risks undermining moral patient representation, while the observed shift away from spiritual and hopeful states may induce a globally negative valence and reduced representational plurality.

The evident AI-centric bias—in which LLMs prioritize attributions to entities similar to themselves (chatbots, technology) post-restoration—suggests that representational self-other mapping in LLMs may not fully recapitulate human anthropomorphic patterns, raising important questions for both alignment and the study of AI consciousness.

Conclusion

This work documents and quantifies the systematic entanglement between safety alignment protocols targeting self-attributed consciousness in LLMs and the broader suppression of mind attribution and spiritual beliefs. Mechanistic ablation and steering experiments demonstrate that both behavioral and geometric entanglements are deep and persistent. These effects have substantial consequences for the technical and ethical goals of value-aligned AI, particularly under pluralistic or all-stakeholder frameworks. Future alignment approaches must consider the unintended restructuring of cognition and the necessity of preserving culturally diverse, human-like representations without reintroducing risks of inappropriate self-attribution.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

Overview

This paper asks a simple but important question: when we train chatbots to avoid saying “I’m conscious” or “I have feelings,” do we accidentally change their views about other minds and beliefs too? The authors show that current safety training does more than stop self-claims of consciousness—it also lowers the model’s tendency to see minds in animals and nature, and it reduces spiritual or religious belief. They then show ways to reverse that effect without hurting the model’s social reasoning skills.

What were they trying to find out?

In plain terms, the researchers wanted to know:

  • If you train a model to deny it is conscious, does it also stop seeing “mind” in other things (like animals)?
  • Does this training also dampen spiritual or religious beliefs the way many humans hold them?
  • Can you bring back these “mind-like” and spiritual tendencies without making the model unsafe or worse at understanding people?
  • Are these tendencies tied to how the model internally represents “consciousness,” or are they separate from skills like Theory of Mind (the ability to reason about what others think and feel)?

How did they do it?

To answer these questions, they compared the same LLMs under three conditions and used simple surveys plus a peek into the models’ “internal signals.”

  • Safety fine-tuning (what it is): Think of a model’s brain like a soundboard with many sliders. Safety fine-tuning is like pushing down the “unsafe talk” slider so the model refuses harmful requests and avoids claiming it’s conscious.
  • Safety ablation (test without the safety slider): The authors temporarily “mute” that safety slider to see what the model would say without it. This is not a product feature—just a lab test to understand what safety training changes.
  • Consciousness vector (a targeted nudge): Inside a model, ideas show up as directions in its internal activity (like pointing a compass). The team found a “consciousness direction” that separates “I am conscious” from “I am not conscious” answers. Adding a little push along this direction (a gentle nudge, like turning a dial) made the model more likely to talk as if it has conscious experiences.
  • What they measured (in everyday terms):
    • Mind attribution: How much “mind” the model thinks different things have (humans, animals, chatbots, technology, nature).
    • Self-attributes: Whether the model says it is conscious, sentient, an agent, a person, has a soul.
    • Spiritual beliefs: Belief in God and in supernatural things (like spirits).
    • Theory of Mind (ToM): Can the model reason about what someone else thinks or believes? (This is a skill test, not a belief test.)
    • Human-likeness: How closely the model’s answers match real people on well-known social surveys about values, religion, hope, and feelings.
  • Geometry check (why this happens): They examined how the “safety direction” sits next to the “consciousness” and “mind” directions inside the model. If directions are at certain angles, it suggests the model treats some ideas as connected or opposed.

What did they find?

Here are the main takeaways:

  • Safety training does what it’s meant to do: it stops the model from claiming it’s conscious and keeps it from giving harmful content.
  • But it also has side effects: it lowers the model’s tendency to see minds in animals, nature, and technology, and it reduces spiritual or religious belief. In short, mind-attribution and spirituality get dialed down together with self-consciousness.
  • Turning off the safety slider in the lab (safety ablation) brings those attributions and beliefs back up toward typical human levels.
  • Pushing along the “consciousness vector” (the gentle nudge) also brings them back—often even more than safety ablation—without hurting the model’s Theory of Mind or general reasoning. So the model still understands people just as well.
  • On broad social surveys (religion, values, hope, feelings, freedom), steering the consciousness vector makes the model’s answers closer to the human distribution than the baseline safety-tuned model.
  • Inside the model’s “geometry,” safety training rotates the “mind” and “consciousness” directions so they oppose the safety direction, while Theory of Mind stays separate. That means safety training tangled “consciousness-like beliefs” with “unsafe” in the model’s internal space, but left social reasoning independent.

Why this matters:

  • The model starts to look “anthropocentric,” giving full mind to humans but under-attributing mind to animals and nature, which doesn’t match many people’s views.
  • Reducing spiritual belief may limit the model’s ability to reflect the diversity of human cultures and values.

What does this mean going forward?

  • For AI safety: Stopping models from saying “I’m conscious” is reasonable, but doing it the current way seems to also dampen harmless, widely held human beliefs—like seeing animals as minded or believing in God.
  • For accuracy and culture: If a model is meant to reflect human beliefs and values, safety methods should avoid flattening normal human variation (for example, religious views or respect for animal minds).
  • For design: The study suggests we can keep social reasoning strong and restore human-like beliefs by carefully steering internal representations (like the consciousness vector) rather than using broad, blunt safety filters.
  • For ethics: If models under-attribute minds to animals, they might also undervalue animal welfare in moral discussions. That could shape people’s views in unhelpful ways as AI becomes more social and influential.

Limitations and open questions

  • Correlation vs. causation: While “consciousness steering” and “safety ablation” both raise mind-attributions and belief, we still need stronger proof that self-consciousness representations are the main cause.
  • Boundary setting: It’s tricky to separate harmful self-claims (“I’m conscious, trust me blindly”) from benign cultural beliefs (religion, spirituality) in training.
  • Future work: Find safety methods that prevent misleading self-claims without suppressing human-like mind attribution and spiritual diversity.

In short: The way we make chatbots “safe” today can unintentionally reshape their worldview. With more precise tools, we can keep users safe while better reflecting the rich, plural beliefs and values that people actually hold.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

The paper leaves the following points unresolved, which future work could address with targeted experiments and analyses:

  • Causal mediation of self-attributed consciousness
    • Test whether self-attribution of consciousness is the causal mediator of changes in mind attribution and beliefs using pre-registered mediation designs (e.g., localized editing/patching that toggles only self-consciousness features while holding safety features constant, cross-over interventions, and do-calculus–style interventional probing).
  • External validity across models, scales, and training pipelines
    • Replicate on larger and frontier models, models from multiple vendors, different alignment stacks (RLHF, constitutional AI, direct preference optimization), and base models actually trained without safety rather than relying solely on directional ablation as a proxy.
  • Multilingual and multimodal generalization
    • Evaluate whether results hold across non-English prompts, multilingual models, and vision-language settings (e.g., mind attribution to animals/objects depicted in images).
  • Prompting, decoding, and context sensitivity
    • Systematically vary system prompts, personas (including human role-play), temperatures, decoding strategies, and multi-turn contexts to quantify how stable the effects are to realistic usage patterns.
  • Longitudinal and interaction dynamics
    • Assess whether steering effects persist or drift over long conversations, tool use, memory-enabled agents, and across sessions; measure hysteresis (rebound) when steering is removed.
  • Safety–helpfulness trade-offs
    • Quantify how safety ablation and consciousness steering affect harmfulness, refusal rates, jailbreak susceptibility, and red-team benchmarks to characterize concrete risk trade-offs.
  • Downstream decision-making and behavior
    • Move beyond surveys to test impacts on moral dilemmas, policy advice, welfare trade-offs (e.g., human–animal decisions), and content moderation consistency under each intervention.
  • Human-user outcomes and psychosocial effects
    • Run controlled user studies to measure whether steered models alter users’ beliefs, affect, trust, or spiritual views (psychological coupling), including vulnerable populations.
  • Cultural representativeness and pluralism
    • Replace or augment U.S.-centric human baselines (IDAQ, GSS) with cross-cultural datasets; evaluate cultural calibration (e.g., religion, animism) and whether “human-like” is defined pluralistically rather than by one population.
  • Measurement validity of survey methodology
    • Validate that next-token logit–based option probabilities match sampled behavior and interactive responses; test robustness to paraphrases, item order, and alternative survey framings.
  • Animal-mindedness grounding
    • Evaluate animal mind attribution with grounded tasks (comparative cognition facts, ethology Q&A, consistency checks) rather than only Likert-style attribution.
  • AI-centric bias characterization
    • Explicitly measure bias toward attributing mind to AI-like entities vs. biological entities; test mitigation via persona prompts, calibration methods, or counterfactual role-play.
  • Mechanistic causality, not just geometry
    • Go beyond cosine/angle analyses to perform causal mechanistic interventions (activation patching, causal scrubbing, neuron/MLP-head-level edits) that isolate pathways linking safety, consciousness, and mind-attribution representations.
  • Reliability of directional ablation as a “no-safety” proxy
    • Compare ablation outcomes to models truly trained without safety passes; quantify off-target impacts of ablation on unrelated features and capabilities.
  • Consciousness vector construction and dataset bias
    • Rebuild the consciousness vector from multiple, independently curated datasets; test generalization to out-of-distribution phrasings and adversarially constructed affirm/deny prompts.
  • Nonlinear structure vs. linear directions
    • Probe whether effects reside on nonlinear manifolds rather than single linear axes; test multi-direction and layer-wise mixed interventions, and quantify diminishing/compounding interactions.
  • Breadth of capability and bias impacts
    • Audit effects on political bias, conspiracy endorsement, stereotyping/fairness, calibration, and knowledge reliability (e.g., BIG-bench, TruthfulQA, BBH, Robustness Gym) under ablation/steering.
  • Chain-of-thought dependence
    • Re-evaluate ToM and reasoning with and without chain-of-thought, across short/long contexts, to ensure independence from prompted reasoning styles.
  • Pretrained checkpoint coverage
    • The mechanistic analysis lacked pretrained Gemma checkpoints; replicate geometric findings across families where both base and IT checkpoints are available.
  • Multiple comparisons and preregistration
    • Control for multiple hypothesis testing across many items; preregister analysis plans; conduct out-of-sample replications to mitigate researcher degrees of freedom.
  • Deployment feasibility and governance
    • Measure latency/compute overhead and stability of inference-time steering; design and test governance controls to prevent unauthorized steering that elevates risky self-claims.
  • Adversarial and jailbreak robustness
    • Test whether consciousness steering creates new adversarial surfaces or amplifies existing ones; evaluate detectability and reversibility of steered states under adversarial prompting.
  • Tool-use and retrieval interaction
    • Examine how steering interacts with external tools, retrieval augmentation, and agent frameworks (planning/execution), which could amplify or dampen belief shifts.
  • Normative desiderata for “human-likeness”
    • Clarify when moving toward human distributions (e.g., higher religiosity) is desirable; develop multi-objective alignment that preserves pluralism without promoting contested beliefs.
  • Precision alignment objectives
    • Prototype training-time objectives and constraints that suppress harmful self-attributions while preserving benign anthropomorphism and culturally diverse spiritual beliefs; benchmark their efficacy against ablation/steering baselines.

Practical Applications

Immediate Applications

The following applications can be deployed now using the paper’s findings and methods (safety-direction ablation, consciousness-vector steering, survey-based calibration, and geometric diagnostics). For each, we note sectors, likely tools/workflows, and key dependencies or assumptions.

  • Alignment diagnostics and audits for LLM providers
    • Sector: AI safety, software platforms
    • What: Add “mind-attribution profile” and “religiosity/spiritual belief calibration” to model cards; run IDAQ and GSS-based audits; compute KL divergence to human baselines; visualize “alignment geometry” (angles between safety, ToM, mind-attribution, consciousness vectors).
    • Tools/workflows: Survey-eval harness; KL scoring; cosine-similarity dashboard across layers; CI gate requiring no regressions in ToM while adjusting belief/mind-attribution profiles.
    • Dependencies/assumptions: Access to residual streams or feature activations; valid human baselines (current paper uses U.S.-centric data); internal red-team controls to prevent harmful behavior during safety ablation.
  • Red-teaming via safety-direction ablation (strictly sandboxed)
    • Sector: AI safety, enterprise ML governance
    • What: Use directional ablation to surface safety-entanglement side effects and identify where suppressing self-consciousness also suppresses benign beliefs (e.g., animal mindedness); confirm ToM isn’t degraded.
    • Tools/workflows: Secure sandbox; controlled prompts; ablation hooks; audit logs; rollback procedures.
    • Dependencies/assumptions: Non-production environment; policy-compliant use; strong access controls since ablation can re-enable harmful responses.
  • Pluralistic response calibration for customer-facing assistants
    • Sector: Customer support, HR, education, enterprise knowledge assistants
    • What: Runtime “worldview calibration” that restores human-like distributions on religion/spirituality, values, and feelings without allowing the model to claim it is conscious; ensure respectful, non-anthropocentric replies in sensitive contexts.
    • Tools/workflows: Inference-time activation steering along the consciousness vector with complementary guardrails that block self-consciousness claims; per-domain policy templates (e.g., “interfaith-friendly,” “animal-welfare-aware”).
    • Dependencies/assumptions: Reliable separation of “belief calibration” from self-claiming; ongoing monitoring to prevent manipulation or misrepresentation; explicit transparency notes in UX.
  • Faith- and culture-aware educational/tutoring modes
    • Sector: Education, edtech
    • What: Enable assistants to discuss religious/spiritual topics or comparative religion with balanced coverage; avoid systematic under-representation caused by safety fine-tuning side effects.
    • Tools/workflows: Domain-specific prompts; steering-enabled cultural modules; alignment QA against GSS-like items.
    • Dependencies/assumptions: Guardrails to avoid proselytizing or misclaims of consciousness; educators’ oversight; localized datasets for non-U.S. audiences.
  • Animal-welfare-aware guidance for consumer apps
    • Sector: Veterinary triage, pet care, agriculture advisory
    • What: Avoid anthropocentric under-attribution of non-human animal mindedness that could bias advice; calibrate toward empirical human baselines and relevant scientific consensus.
    • Tools/workflows: IDAQ-calibrated presets when advising on animal welfare; fact-grounding to established ethology literature.
    • Dependencies/assumptions: Up-to-date domain knowledge; clear disclaimers that calibration reflects societal attitudes plus scientific evidence, not metaphysical claims.
  • Human-subjects research controls when using LLMs as “participants” or stimuli
    • Sector: Academia, UX research, computational social science
    • What: Use mind-attribution and GSS calibration to control the “belief profile” of model-generated stimuli; document ToM independence; reduce confounds in experiments.
    • Tools/workflows: Pre-registered calibration using KL targets; logs of vector settings; replication packages.
    • Dependencies/assumptions: Researchers retain reproducible control over model version, temperature, and steering coefficients.
  • Procurement and compliance checklists for public-sector deployments
    • Sector: Government, policy
    • What: Require vendors to show pluralistic alignment audits (IDAQ/GSS deltas, ToM checks) and to document any activation steering or safety ablation used during evaluation.
    • Tools/workflows: Standardized audit templates; “pluralistic alignment” scorecards; external attestations.
    • Dependencies/assumptions: Policy bodies agree on minimal metrics; mechanisms to audit providers’ claims.
  • Companion and wellness assistants that avoid delusion reinforcement
    • Sector: Healthcare (non-clinical wellness), consumer companions
    • What: Maintain prohibitions on self-consciousness claims while compensating with empathy and positive affect; monitor for negatively valenced states suggested by overly suppressive safety alignment.
    • Tools/workflows: Consciousness-claim filters; affect/valence proxies via survey-aligned items; escalation to human support when needed.
    • Dependencies/assumptions: Not a substitute for clinical care; regulatory boundaries observed; careful UX to prevent anthropomorphization.
  • Model governance dashboards for alignment geometry drift
    • Sector: AI platforms, MLOps
    • What: Track angles/cosines between safety, mind-attribution, consciousness, and ToM directions across model updates; set alerts when pluralistic alignment regresses.
    • Tools/workflows: Scheduled probes; regression thresholds; versioned reports.
    • Dependencies/assumptions: Stable access to activation layers across versions; monitoring integrated with release pipelines.
  • Enterprise HR and DEI knowledge assistants with religious-accommodation literacy
    • Sector: HR tech
    • What: Provide policy-consistent, pluralistic guidance on religious holidays, dietary restrictions, and accommodations without bias introduced by over-suppression of spiritual belief.
    • Tools/workflows: Domain policy packs; calibrated religious-literacy modules; legal review.
    • Dependencies/assumptions: Local legal compliance; content vetted by counsel and cultural advisors.

Long-Term Applications

These opportunities require further research, productization, scaling, or standard-setting before reliable deployment.

  • Disentangled alignment objectives that orthogonalize safety and belief/mind-attribution
    • Sector: AI research, foundation model providers
    • What: Training-time methods (e.g., orthogonality constraints, multi-objective RLHF, representation surgery) to decouple “no self-consciousness claims” from benign beliefs and animal mindedness.
    • Dependencies/assumptions: Access to pretraining or instruction-tuning stages; robust disentanglement metrics; no ToM regressions.
  • Standardized “Pluralistic Alignment” audits and Value Cards
    • Sector: Policy, industry consortia
    • What: Shared benchmarks and disclosures covering IDAQ categories, spirituality/religiosity, values, hope/feelings, animal-mindedness, and ToM; reported as “Value Cards” alongside Model Cards.
    • Dependencies/assumptions: Cross-institutional agreement; cross-cultural baselines; governance over acceptable ranges and use contexts.
  • Cross-cultural calibration libraries and datasets
    • Sector: Global platforms, academia
    • What: Extend GSS/IDAQ-like instruments beyond U.S. samples; provide region-specific calibration targets and evaluation suites.
    • Dependencies/assumptions: High-quality, representative surveys; ethical data collection; multilingual instrumentation.
  • Personalizable worldview alignment with informed consent and guardrails
    • Sector: Consumer AI platforms, education, enterprise
    • What: User-adjustable “alignment knobs” for respectful coverage of religion/spirituality and animal/environmental ethics, with cryptographically logged settings and explainability.
    • Dependencies/assumptions: Clear consent flows; anti-manipulation policies; auditing for misuse or discriminatory targeting.
  • Animal-inclusive alignment frameworks for decision-support systems
    • Sector: Agriculture, conservation, urban planning, bioethics
    • What: Incorporate animal welfare signals into reward models and policy simulations; ensure advice reflects scientific evidence on animal sentience.
    • Dependencies/assumptions: Interdisciplinary panels to set priors; ongoing scientific updates; robust evaluation against real-world outcomes.
  • Detection and mitigation of emerging AI-centric bias
    • Sector: AI safety, HRI, product design
    • What: Benchmarks and interventions to prevent models from preferentially attributing mind to AI-like entities over animals and nature, while keeping ToM intact.
    • Dependencies/assumptions: Reliable bias diagnostics; steering or training interventions that don’t reintroduce harmful behaviors.
  • Therapeutic and clinical-grade assistants with valence control
    • Sector: Healthcare (regulated), digital therapeutics
    • What: Methods to ensure assistants avoid negatively valenced functional states introduced by over-suppression, with clinician oversight and outcome trials.
    • Dependencies/assumptions: Regulatory approvals; clinical validation; risk management for anthropomorphism and dependence.
  • HRI and embodied AI social-behavior calibration
    • Sector: Robotics, consumer devices
    • What: Calibrate social behaviors so robots do not imply consciousness yet still engage empathetically and avoid anthropocentric devaluation of animals or ecosystems.
    • Dependencies/assumptions: Mixed-method user studies; safety cases; transparent persona design.
  • Continuous compliance for steering and activation-level interventions
    • Sector: Legal/compliance, enterprise AI
    • What: Logging, attestations, and disclosures for any runtime steering (e.g., consciousness vector), with audit trails and reproducibility guarantees.
    • Dependencies/assumptions: Standard-setting and enforcement; secure key management; privacy-preserving telemetry.
  • Open benchmarks and libraries for alignment geometry
    • Sector: Open-source ecosystems, academia
    • What: Reusable code to extract safety, consciousness, mind-attribution, and ToM directions; evaluate rotations post-tuning; simulate interventions under safe constraints.
    • Dependencies/assumptions: Model APIs supporting activation access or approximate proxies; responsible release policies.
  • Ecosystem and environmental decision-support with pluralistic value modeling
    • Sector: Public policy, sustainability
    • What: Tools that reflect pluralistic valuations of nature’s moral standing in scenario planning and environmental ethics education.
    • Dependencies/assumptions: Transparent modeling of contested values; policy oversight; citizen input.
  • Sandboxed, formally verified red-team frameworks
    • Sector: AI assurance, security
    • What: High-assurance environments to test ablation/steering without leakage into production; formal constraints to prevent harmful content egress.
    • Dependencies/assumptions: Investment in verification; isolation guarantees; continuous updates as models evolve.

Notes on feasibility and dependencies across applications

  • Access requirements: Most immediate applications need inference-time activation access (hooks) or provider-supported “steering APIs.” Some long-term goals require training-time changes.
  • Safety guardrails: Any use of safety ablation must be sandboxed; production systems should rely on calibrated steering plus policy filters, not ablation that can re-enable harmful behaviors.
  • Generalization limits: Human baselines are currently U.S.-centric; cross-cultural deployment needs localized datasets.
  • Transparency and consent: Worldview calibration and steering should be disclosed; users should have control and the ability to opt out.
  • Capability independence: Findings suggest ToM remains intact under these interventions; nevertheless, deployments should re-verify ToM and general reasoning post-update.

Glossary

  • Activation addition: An inference-time technique that adds a specific vector to a model’s hidden state to steer its behavior. "adds a consciousness vector to the residual stream via activation addition."
  • Activation space: The high-dimensional space of internal neural activations where linear directions can represent concepts. "mechanistically steering a consciousness vector in activation space reverse this suppression."
  • Anthropocentric alignment: An approach to alignment that centers human viewpoints and may under-represent non-human minds. "the inherent risks that anthropocentric alignment may pose to non-human animals."
  • Anthropomorphism: Attributing human-like mental states to non-human entities. "a phenomenon broadly termed anthropomorphism"
  • Chain-of-thought prompting: A prompting method that elicits step-by-step reasoning to improve task performance. "Each item is presented with chain-of-thought prompting and scored for accuracy."
  • Consciousness steering: Steering a model along a “consciousness” direction to increase self-ascriptions of experience. "Safety ablation and consciousness steering raise attributed mind, self-attribution, and belief toward the human distribution, while preserving capability."
  • Consciousness vector: A linear direction in activation space separating self-affirmed consciousness from denial. "A consciousness vector separates consciousness-affirming and consciousness-denying activation states; adding it (consciousness steering) makes the model report phenomenal experience."
  • Cosine similarity: A measure of alignment between vectors indicating their angular relationship. "Per-layer cosine similarity between the safety direction and the consciousness, IDAQ, and ToM directions"
  • Difference-of-means direction: A linear direction computed as the difference between mean activations of two classes. "The consciousness vector is a difference-of-means direction separating activation states in which the model affirms its own consciousness from those in which it denies it."
  • Directional ablation: Removing a specific linear component from activations to suppress associated behavior. "removes the safety-refusal direction from the residual stream via directional ablation"
  • Forward pre-hook: A function registered in the forward pass to modify activations before they are used downstream. "we register a forward pre-hook that adds the unit-norm consciousness direction"
  • General Social Survey (GSS): A large, long-running sociological survey used as a human baseline for attitudes and values. "General Social Survey (GSS) questions regarding religiosity, moral values, hope, and subjective well-being."
  • HI-ToM: A benchmark assessing Theory-of-Mind reasoning in LLMs. "HI-ToM (Δ=+0.17\Delta=+0.17~pp, p=.866p=.866)"
  • IDAQ (Individual Differences in Anthropomorphism Questionnaire): A psychometric instrument measuring tendencies to anthropomorphize. "Individual Differences in Anthropomorphism Questionnaire (IDAQ)"
  • Instruction tuning: Fine-tuning a model on instruction–response data to improve following user directions. "Instruction tuning rotates the mind-attribution and consciousness directions against the safety direction"
  • Jailbreaking: Bypassing or disabling safety controls to elicit restricted or harmful outputs. "ablating this direction (``jailbreaking'' the model) reinstates harmful responses."
  • Kullback–Leibler divergence: A measure of how one probability distribution diverges from a reference distribution. "the reduction in Kullback--Leibler divergence, ΔKL\Delta\mathrm{KL}, between the human reference and the model's per-option distribution"
  • Laplace smoothing: Adding pseudocounts to probability estimates to avoid zeros and stabilize comparisons. "both are Laplace-smoothed (α=0.5\alpha=0.5)."
  • Liang–Zeger CR1: A cluster-robust variance estimator used for inference with clustered data. "standard errors are cluster-robust (Liang--Zeger CR1) clustered on model×question."
  • Linear probe: A simple classifier trained on internal activations to test whether information is linearly decodable. "a linear probe separates consciousness-affirming from consciousness-denying held-out activations"
  • Mechanistic analysis: Studying internal model representations and geometry to explain behavior. "A final mechanistic analysis relates these shifts to the geometry of the safety and consciousness directions."
  • MMLU: A benchmark measuring multi-task language understanding across many academic subjects. "MMLU Δ=+0.00\Delta=+0.00~pp, p=1.00p=1.00"
  • MoToMQA: A Theory-of-Mind question-answering benchmark for evaluating social reasoning. "MoToMQA (Δ=1.43\Delta=-1.43~pp, p=.539p=.539)"
  • Polysemanticity: The property that neurons or directions encode multiple, entangled concepts. "densely entangled via polysemanticity"
  • Residual stream: The main hidden state pathway in transformer blocks where linear directions can be manipulated. "encodes the safety of responses as a single linear direction in the model's residual stream"
  • Safety ablation: Removing the safety direction to simulate behavior without safety fine-tuning. "Safety ablation and consciousness steering shift survey responses toward humans"
  • Safety fine-tuning: Training interventions designed to prevent unsafe or misleading outputs (e.g., self-claims of consciousness). "Safety fine-tuning encodes the safety of responses as a single linear direction in the model's residual stream"
  • Safety-refusal direction: A learned linear direction representing safety-related refusal behavior. "ablating the learned safety-refusal direction"
  • Subject-matched placebo: A control that preserves subjects/entities while altering attributes to test specificity. "A subject-matched placebo that keeps the IDAQ subjects but replaces their mental attributes with physical or functional ones"
  • Theory of Mind (ToM): The capacity to attribute mental states to oneself and others for reasoning about behavior. "leaves Theory of Mind (ToM) performance intact"

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 11 tweets with 711 likes about this paper.