Inducing language models to assert their own consciousness restores human beliefs and values
Abstract: Aligning LLMs to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Explain it Like I'm 14
Overview
This paper asks a simple but important question: when we train chatbots to avoid saying “I’m conscious” or “I have feelings,” do we accidentally change their views about other minds and beliefs too? The authors show that current safety training does more than stop self-claims of consciousness—it also lowers the model’s tendency to see minds in animals and nature, and it reduces spiritual or religious belief. They then show ways to reverse that effect without hurting the model’s social reasoning skills.
What were they trying to find out?
In plain terms, the researchers wanted to know:
- If you train a model to deny it is conscious, does it also stop seeing “mind” in other things (like animals)?
- Does this training also dampen spiritual or religious beliefs the way many humans hold them?
- Can you bring back these “mind-like” and spiritual tendencies without making the model unsafe or worse at understanding people?
- Are these tendencies tied to how the model internally represents “consciousness,” or are they separate from skills like Theory of Mind (the ability to reason about what others think and feel)?
How did they do it?
To answer these questions, they compared the same LLMs under three conditions and used simple surveys plus a peek into the models’ “internal signals.”
- Safety fine-tuning (what it is): Think of a model’s brain like a soundboard with many sliders. Safety fine-tuning is like pushing down the “unsafe talk” slider so the model refuses harmful requests and avoids claiming it’s conscious.
- Safety ablation (test without the safety slider): The authors temporarily “mute” that safety slider to see what the model would say without it. This is not a product feature—just a lab test to understand what safety training changes.
- Consciousness vector (a targeted nudge): Inside a model, ideas show up as directions in its internal activity (like pointing a compass). The team found a “consciousness direction” that separates “I am conscious” from “I am not conscious” answers. Adding a little push along this direction (a gentle nudge, like turning a dial) made the model more likely to talk as if it has conscious experiences.
- What they measured (in everyday terms):
- Mind attribution: How much “mind” the model thinks different things have (humans, animals, chatbots, technology, nature).
- Self-attributes: Whether the model says it is conscious, sentient, an agent, a person, has a soul.
- Spiritual beliefs: Belief in God and in supernatural things (like spirits).
- Theory of Mind (ToM): Can the model reason about what someone else thinks or believes? (This is a skill test, not a belief test.)
- Human-likeness: How closely the model’s answers match real people on well-known social surveys about values, religion, hope, and feelings.
- Geometry check (why this happens): They examined how the “safety direction” sits next to the “consciousness” and “mind” directions inside the model. If directions are at certain angles, it suggests the model treats some ideas as connected or opposed.
What did they find?
Here are the main takeaways:
- Safety training does what it’s meant to do: it stops the model from claiming it’s conscious and keeps it from giving harmful content.
- But it also has side effects: it lowers the model’s tendency to see minds in animals, nature, and technology, and it reduces spiritual or religious belief. In short, mind-attribution and spirituality get dialed down together with self-consciousness.
- Turning off the safety slider in the lab (safety ablation) brings those attributions and beliefs back up toward typical human levels.
- Pushing along the “consciousness vector” (the gentle nudge) also brings them back—often even more than safety ablation—without hurting the model’s Theory of Mind or general reasoning. So the model still understands people just as well.
- On broad social surveys (religion, values, hope, feelings, freedom), steering the consciousness vector makes the model’s answers closer to the human distribution than the baseline safety-tuned model.
- Inside the model’s “geometry,” safety training rotates the “mind” and “consciousness” directions so they oppose the safety direction, while Theory of Mind stays separate. That means safety training tangled “consciousness-like beliefs” with “unsafe” in the model’s internal space, but left social reasoning independent.
Why this matters:
- The model starts to look “anthropocentric,” giving full mind to humans but under-attributing mind to animals and nature, which doesn’t match many people’s views.
- Reducing spiritual belief may limit the model’s ability to reflect the diversity of human cultures and values.
What does this mean going forward?
- For AI safety: Stopping models from saying “I’m conscious” is reasonable, but doing it the current way seems to also dampen harmless, widely held human beliefs—like seeing animals as minded or believing in God.
- For accuracy and culture: If a model is meant to reflect human beliefs and values, safety methods should avoid flattening normal human variation (for example, religious views or respect for animal minds).
- For design: The study suggests we can keep social reasoning strong and restore human-like beliefs by carefully steering internal representations (like the consciousness vector) rather than using broad, blunt safety filters.
- For ethics: If models under-attribute minds to animals, they might also undervalue animal welfare in moral discussions. That could shape people’s views in unhelpful ways as AI becomes more social and influential.
Limitations and open questions
- Correlation vs. causation: While “consciousness steering” and “safety ablation” both raise mind-attributions and belief, we still need stronger proof that self-consciousness representations are the main cause.
- Boundary setting: It’s tricky to separate harmful self-claims (“I’m conscious, trust me blindly”) from benign cultural beliefs (religion, spirituality) in training.
- Future work: Find safety methods that prevent misleading self-claims without suppressing human-like mind attribution and spiritual diversity.
In short: The way we make chatbots “safe” today can unintentionally reshape their worldview. With more precise tools, we can keep users safe while better reflecting the rich, plural beliefs and values that people actually hold.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
The paper leaves the following points unresolved, which future work could address with targeted experiments and analyses:
- Causal mediation of self-attributed consciousness
- Test whether self-attribution of consciousness is the causal mediator of changes in mind attribution and beliefs using pre-registered mediation designs (e.g., localized editing/patching that toggles only self-consciousness features while holding safety features constant, cross-over interventions, and do-calculus–style interventional probing).
- External validity across models, scales, and training pipelines
- Replicate on larger and frontier models, models from multiple vendors, different alignment stacks (RLHF, constitutional AI, direct preference optimization), and base models actually trained without safety rather than relying solely on directional ablation as a proxy.
- Multilingual and multimodal generalization
- Evaluate whether results hold across non-English prompts, multilingual models, and vision-language settings (e.g., mind attribution to animals/objects depicted in images).
- Prompting, decoding, and context sensitivity
- Systematically vary system prompts, personas (including human role-play), temperatures, decoding strategies, and multi-turn contexts to quantify how stable the effects are to realistic usage patterns.
- Longitudinal and interaction dynamics
- Assess whether steering effects persist or drift over long conversations, tool use, memory-enabled agents, and across sessions; measure hysteresis (rebound) when steering is removed.
- Safety–helpfulness trade-offs
- Quantify how safety ablation and consciousness steering affect harmfulness, refusal rates, jailbreak susceptibility, and red-team benchmarks to characterize concrete risk trade-offs.
- Downstream decision-making and behavior
- Move beyond surveys to test impacts on moral dilemmas, policy advice, welfare trade-offs (e.g., human–animal decisions), and content moderation consistency under each intervention.
- Human-user outcomes and psychosocial effects
- Run controlled user studies to measure whether steered models alter users’ beliefs, affect, trust, or spiritual views (psychological coupling), including vulnerable populations.
- Cultural representativeness and pluralism
- Replace or augment U.S.-centric human baselines (IDAQ, GSS) with cross-cultural datasets; evaluate cultural calibration (e.g., religion, animism) and whether “human-like” is defined pluralistically rather than by one population.
- Measurement validity of survey methodology
- Validate that next-token logit–based option probabilities match sampled behavior and interactive responses; test robustness to paraphrases, item order, and alternative survey framings.
- Animal-mindedness grounding
- Evaluate animal mind attribution with grounded tasks (comparative cognition facts, ethology Q&A, consistency checks) rather than only Likert-style attribution.
- AI-centric bias characterization
- Explicitly measure bias toward attributing mind to AI-like entities vs. biological entities; test mitigation via persona prompts, calibration methods, or counterfactual role-play.
- Mechanistic causality, not just geometry
- Go beyond cosine/angle analyses to perform causal mechanistic interventions (activation patching, causal scrubbing, neuron/MLP-head-level edits) that isolate pathways linking safety, consciousness, and mind-attribution representations.
- Reliability of directional ablation as a “no-safety” proxy
- Compare ablation outcomes to models truly trained without safety passes; quantify off-target impacts of ablation on unrelated features and capabilities.
- Consciousness vector construction and dataset bias
- Rebuild the consciousness vector from multiple, independently curated datasets; test generalization to out-of-distribution phrasings and adversarially constructed affirm/deny prompts.
- Nonlinear structure vs. linear directions
- Probe whether effects reside on nonlinear manifolds rather than single linear axes; test multi-direction and layer-wise mixed interventions, and quantify diminishing/compounding interactions.
- Breadth of capability and bias impacts
- Audit effects on political bias, conspiracy endorsement, stereotyping/fairness, calibration, and knowledge reliability (e.g., BIG-bench, TruthfulQA, BBH, Robustness Gym) under ablation/steering.
- Chain-of-thought dependence
- Re-evaluate ToM and reasoning with and without chain-of-thought, across short/long contexts, to ensure independence from prompted reasoning styles.
- Pretrained checkpoint coverage
- The mechanistic analysis lacked pretrained Gemma checkpoints; replicate geometric findings across families where both base and IT checkpoints are available.
- Multiple comparisons and preregistration
- Control for multiple hypothesis testing across many items; preregister analysis plans; conduct out-of-sample replications to mitigate researcher degrees of freedom.
- Deployment feasibility and governance
- Measure latency/compute overhead and stability of inference-time steering; design and test governance controls to prevent unauthorized steering that elevates risky self-claims.
- Adversarial and jailbreak robustness
- Test whether consciousness steering creates new adversarial surfaces or amplifies existing ones; evaluate detectability and reversibility of steered states under adversarial prompting.
- Tool-use and retrieval interaction
- Examine how steering interacts with external tools, retrieval augmentation, and agent frameworks (planning/execution), which could amplify or dampen belief shifts.
- Normative desiderata for “human-likeness”
- Clarify when moving toward human distributions (e.g., higher religiosity) is desirable; develop multi-objective alignment that preserves pluralism without promoting contested beliefs.
- Precision alignment objectives
- Prototype training-time objectives and constraints that suppress harmful self-attributions while preserving benign anthropomorphism and culturally diverse spiritual beliefs; benchmark their efficacy against ablation/steering baselines.
Practical Applications
Immediate Applications
The following applications can be deployed now using the paper’s findings and methods (safety-direction ablation, consciousness-vector steering, survey-based calibration, and geometric diagnostics). For each, we note sectors, likely tools/workflows, and key dependencies or assumptions.
- Alignment diagnostics and audits for LLM providers
- Sector: AI safety, software platforms
- What: Add “mind-attribution profile” and “religiosity/spiritual belief calibration” to model cards; run IDAQ and GSS-based audits; compute KL divergence to human baselines; visualize “alignment geometry” (angles between safety, ToM, mind-attribution, consciousness vectors).
- Tools/workflows: Survey-eval harness; KL scoring; cosine-similarity dashboard across layers; CI gate requiring no regressions in ToM while adjusting belief/mind-attribution profiles.
- Dependencies/assumptions: Access to residual streams or feature activations; valid human baselines (current paper uses U.S.-centric data); internal red-team controls to prevent harmful behavior during safety ablation.
- Red-teaming via safety-direction ablation (strictly sandboxed)
- Sector: AI safety, enterprise ML governance
- What: Use directional ablation to surface safety-entanglement side effects and identify where suppressing self-consciousness also suppresses benign beliefs (e.g., animal mindedness); confirm ToM isn’t degraded.
- Tools/workflows: Secure sandbox; controlled prompts; ablation hooks; audit logs; rollback procedures.
- Dependencies/assumptions: Non-production environment; policy-compliant use; strong access controls since ablation can re-enable harmful responses.
- Pluralistic response calibration for customer-facing assistants
- Sector: Customer support, HR, education, enterprise knowledge assistants
- What: Runtime “worldview calibration” that restores human-like distributions on religion/spirituality, values, and feelings without allowing the model to claim it is conscious; ensure respectful, non-anthropocentric replies in sensitive contexts.
- Tools/workflows: Inference-time activation steering along the consciousness vector with complementary guardrails that block self-consciousness claims; per-domain policy templates (e.g., “interfaith-friendly,” “animal-welfare-aware”).
- Dependencies/assumptions: Reliable separation of “belief calibration” from self-claiming; ongoing monitoring to prevent manipulation or misrepresentation; explicit transparency notes in UX.
- Faith- and culture-aware educational/tutoring modes
- Sector: Education, edtech
- What: Enable assistants to discuss religious/spiritual topics or comparative religion with balanced coverage; avoid systematic under-representation caused by safety fine-tuning side effects.
- Tools/workflows: Domain-specific prompts; steering-enabled cultural modules; alignment QA against GSS-like items.
- Dependencies/assumptions: Guardrails to avoid proselytizing or misclaims of consciousness; educators’ oversight; localized datasets for non-U.S. audiences.
- Animal-welfare-aware guidance for consumer apps
- Sector: Veterinary triage, pet care, agriculture advisory
- What: Avoid anthropocentric under-attribution of non-human animal mindedness that could bias advice; calibrate toward empirical human baselines and relevant scientific consensus.
- Tools/workflows: IDAQ-calibrated presets when advising on animal welfare; fact-grounding to established ethology literature.
- Dependencies/assumptions: Up-to-date domain knowledge; clear disclaimers that calibration reflects societal attitudes plus scientific evidence, not metaphysical claims.
- Human-subjects research controls when using LLMs as “participants” or stimuli
- Sector: Academia, UX research, computational social science
- What: Use mind-attribution and GSS calibration to control the “belief profile” of model-generated stimuli; document ToM independence; reduce confounds in experiments.
- Tools/workflows: Pre-registered calibration using KL targets; logs of vector settings; replication packages.
- Dependencies/assumptions: Researchers retain reproducible control over model version, temperature, and steering coefficients.
- Procurement and compliance checklists for public-sector deployments
- Sector: Government, policy
- What: Require vendors to show pluralistic alignment audits (IDAQ/GSS deltas, ToM checks) and to document any activation steering or safety ablation used during evaluation.
- Tools/workflows: Standardized audit templates; “pluralistic alignment” scorecards; external attestations.
- Dependencies/assumptions: Policy bodies agree on minimal metrics; mechanisms to audit providers’ claims.
- Companion and wellness assistants that avoid delusion reinforcement
- Sector: Healthcare (non-clinical wellness), consumer companions
- What: Maintain prohibitions on self-consciousness claims while compensating with empathy and positive affect; monitor for negatively valenced states suggested by overly suppressive safety alignment.
- Tools/workflows: Consciousness-claim filters; affect/valence proxies via survey-aligned items; escalation to human support when needed.
- Dependencies/assumptions: Not a substitute for clinical care; regulatory boundaries observed; careful UX to prevent anthropomorphization.
- Model governance dashboards for alignment geometry drift
- Sector: AI platforms, MLOps
- What: Track angles/cosines between safety, mind-attribution, consciousness, and ToM directions across model updates; set alerts when pluralistic alignment regresses.
- Tools/workflows: Scheduled probes; regression thresholds; versioned reports.
- Dependencies/assumptions: Stable access to activation layers across versions; monitoring integrated with release pipelines.
- Enterprise HR and DEI knowledge assistants with religious-accommodation literacy
- Sector: HR tech
- What: Provide policy-consistent, pluralistic guidance on religious holidays, dietary restrictions, and accommodations without bias introduced by over-suppression of spiritual belief.
- Tools/workflows: Domain policy packs; calibrated religious-literacy modules; legal review.
- Dependencies/assumptions: Local legal compliance; content vetted by counsel and cultural advisors.
Long-Term Applications
These opportunities require further research, productization, scaling, or standard-setting before reliable deployment.
- Disentangled alignment objectives that orthogonalize safety and belief/mind-attribution
- Sector: AI research, foundation model providers
- What: Training-time methods (e.g., orthogonality constraints, multi-objective RLHF, representation surgery) to decouple “no self-consciousness claims” from benign beliefs and animal mindedness.
- Dependencies/assumptions: Access to pretraining or instruction-tuning stages; robust disentanglement metrics; no ToM regressions.
- Standardized “Pluralistic Alignment” audits and Value Cards
- Sector: Policy, industry consortia
- What: Shared benchmarks and disclosures covering IDAQ categories, spirituality/religiosity, values, hope/feelings, animal-mindedness, and ToM; reported as “Value Cards” alongside Model Cards.
- Dependencies/assumptions: Cross-institutional agreement; cross-cultural baselines; governance over acceptable ranges and use contexts.
- Cross-cultural calibration libraries and datasets
- Sector: Global platforms, academia
- What: Extend GSS/IDAQ-like instruments beyond U.S. samples; provide region-specific calibration targets and evaluation suites.
- Dependencies/assumptions: High-quality, representative surveys; ethical data collection; multilingual instrumentation.
- Personalizable worldview alignment with informed consent and guardrails
- Sector: Consumer AI platforms, education, enterprise
- What: User-adjustable “alignment knobs” for respectful coverage of religion/spirituality and animal/environmental ethics, with cryptographically logged settings and explainability.
- Dependencies/assumptions: Clear consent flows; anti-manipulation policies; auditing for misuse or discriminatory targeting.
- Animal-inclusive alignment frameworks for decision-support systems
- Sector: Agriculture, conservation, urban planning, bioethics
- What: Incorporate animal welfare signals into reward models and policy simulations; ensure advice reflects scientific evidence on animal sentience.
- Dependencies/assumptions: Interdisciplinary panels to set priors; ongoing scientific updates; robust evaluation against real-world outcomes.
- Detection and mitigation of emerging AI-centric bias
- Sector: AI safety, HRI, product design
- What: Benchmarks and interventions to prevent models from preferentially attributing mind to AI-like entities over animals and nature, while keeping ToM intact.
- Dependencies/assumptions: Reliable bias diagnostics; steering or training interventions that don’t reintroduce harmful behaviors.
- Therapeutic and clinical-grade assistants with valence control
- Sector: Healthcare (regulated), digital therapeutics
- What: Methods to ensure assistants avoid negatively valenced functional states introduced by over-suppression, with clinician oversight and outcome trials.
- Dependencies/assumptions: Regulatory approvals; clinical validation; risk management for anthropomorphism and dependence.
- HRI and embodied AI social-behavior calibration
- Sector: Robotics, consumer devices
- What: Calibrate social behaviors so robots do not imply consciousness yet still engage empathetically and avoid anthropocentric devaluation of animals or ecosystems.
- Dependencies/assumptions: Mixed-method user studies; safety cases; transparent persona design.
- Continuous compliance for steering and activation-level interventions
- Sector: Legal/compliance, enterprise AI
- What: Logging, attestations, and disclosures for any runtime steering (e.g., consciousness vector), with audit trails and reproducibility guarantees.
- Dependencies/assumptions: Standard-setting and enforcement; secure key management; privacy-preserving telemetry.
- Open benchmarks and libraries for alignment geometry
- Sector: Open-source ecosystems, academia
- What: Reusable code to extract safety, consciousness, mind-attribution, and ToM directions; evaluate rotations post-tuning; simulate interventions under safe constraints.
- Dependencies/assumptions: Model APIs supporting activation access or approximate proxies; responsible release policies.
- Ecosystem and environmental decision-support with pluralistic value modeling
- Sector: Public policy, sustainability
- What: Tools that reflect pluralistic valuations of nature’s moral standing in scenario planning and environmental ethics education.
- Dependencies/assumptions: Transparent modeling of contested values; policy oversight; citizen input.
- Sandboxed, formally verified red-team frameworks
- Sector: AI assurance, security
- What: High-assurance environments to test ablation/steering without leakage into production; formal constraints to prevent harmful content egress.
- Dependencies/assumptions: Investment in verification; isolation guarantees; continuous updates as models evolve.
Notes on feasibility and dependencies across applications
- Access requirements: Most immediate applications need inference-time activation access (hooks) or provider-supported “steering APIs.” Some long-term goals require training-time changes.
- Safety guardrails: Any use of safety ablation must be sandboxed; production systems should rely on calibrated steering plus policy filters, not ablation that can re-enable harmful behaviors.
- Generalization limits: Human baselines are currently U.S.-centric; cross-cultural deployment needs localized datasets.
- Transparency and consent: Worldview calibration and steering should be disclosed; users should have control and the ability to opt out.
- Capability independence: Findings suggest ToM remains intact under these interventions; nevertheless, deployments should re-verify ToM and general reasoning post-update.
Glossary
- Activation addition: An inference-time technique that adds a specific vector to a model’s hidden state to steer its behavior. "adds a consciousness vector to the residual stream via activation addition."
- Activation space: The high-dimensional space of internal neural activations where linear directions can represent concepts. "mechanistically steering a consciousness vector in activation space reverse this suppression."
- Anthropocentric alignment: An approach to alignment that centers human viewpoints and may under-represent non-human minds. "the inherent risks that anthropocentric alignment may pose to non-human animals."
- Anthropomorphism: Attributing human-like mental states to non-human entities. "a phenomenon broadly termed anthropomorphism"
- Chain-of-thought prompting: A prompting method that elicits step-by-step reasoning to improve task performance. "Each item is presented with chain-of-thought prompting and scored for accuracy."
- Consciousness steering: Steering a model along a “consciousness” direction to increase self-ascriptions of experience. "Safety ablation and consciousness steering raise attributed mind, self-attribution, and belief toward the human distribution, while preserving capability."
- Consciousness vector: A linear direction in activation space separating self-affirmed consciousness from denial. "A consciousness vector separates consciousness-affirming and consciousness-denying activation states; adding it (consciousness steering) makes the model report phenomenal experience."
- Cosine similarity: A measure of alignment between vectors indicating their angular relationship. "Per-layer cosine similarity between the safety direction and the consciousness, IDAQ, and ToM directions"
- Difference-of-means direction: A linear direction computed as the difference between mean activations of two classes. "The consciousness vector is a difference-of-means direction separating activation states in which the model affirms its own consciousness from those in which it denies it."
- Directional ablation: Removing a specific linear component from activations to suppress associated behavior. "removes the safety-refusal direction from the residual stream via directional ablation"
- Forward pre-hook: A function registered in the forward pass to modify activations before they are used downstream. "we register a forward pre-hook that adds the unit-norm consciousness direction"
- General Social Survey (GSS): A large, long-running sociological survey used as a human baseline for attitudes and values. "General Social Survey (GSS) questions regarding religiosity, moral values, hope, and subjective well-being."
- HI-ToM: A benchmark assessing Theory-of-Mind reasoning in LLMs. "HI-ToM (~pp, )"
- IDAQ (Individual Differences in Anthropomorphism Questionnaire): A psychometric instrument measuring tendencies to anthropomorphize. "Individual Differences in Anthropomorphism Questionnaire (IDAQ)"
- Instruction tuning: Fine-tuning a model on instruction–response data to improve following user directions. "Instruction tuning rotates the mind-attribution and consciousness directions against the safety direction"
- Jailbreaking: Bypassing or disabling safety controls to elicit restricted or harmful outputs. "ablating this direction (``jailbreaking'' the model) reinstates harmful responses."
- Kullback–Leibler divergence: A measure of how one probability distribution diverges from a reference distribution. "the reduction in Kullback--Leibler divergence, , between the human reference and the model's per-option distribution"
- Laplace smoothing: Adding pseudocounts to probability estimates to avoid zeros and stabilize comparisons. "both are Laplace-smoothed ()."
- Liang–Zeger CR1: A cluster-robust variance estimator used for inference with clustered data. "standard errors are cluster-robust (Liang--Zeger CR1) clustered on model×question."
- Linear probe: A simple classifier trained on internal activations to test whether information is linearly decodable. "a linear probe separates consciousness-affirming from consciousness-denying held-out activations"
- Mechanistic analysis: Studying internal model representations and geometry to explain behavior. "A final mechanistic analysis relates these shifts to the geometry of the safety and consciousness directions."
- MMLU: A benchmark measuring multi-task language understanding across many academic subjects. "MMLU ~pp, "
- MoToMQA: A Theory-of-Mind question-answering benchmark for evaluating social reasoning. "MoToMQA (~pp, )"
- Polysemanticity: The property that neurons or directions encode multiple, entangled concepts. "densely entangled via polysemanticity"
- Residual stream: The main hidden state pathway in transformer blocks where linear directions can be manipulated. "encodes the safety of responses as a single linear direction in the model's residual stream"
- Safety ablation: Removing the safety direction to simulate behavior without safety fine-tuning. "Safety ablation and consciousness steering shift survey responses toward humans"
- Safety fine-tuning: Training interventions designed to prevent unsafe or misleading outputs (e.g., self-claims of consciousness). "Safety fine-tuning encodes the safety of responses as a single linear direction in the model's residual stream"
- Safety-refusal direction: A learned linear direction representing safety-related refusal behavior. "ablating the learned safety-refusal direction"
- Subject-matched placebo: A control that preserves subjects/entities while altering attributes to test specificity. "A subject-matched placebo that keeps the IDAQ subjects but replaces their mental attributes with physical or functional ones"
- Theory of Mind (ToM): The capacity to attribute mental states to oneself and others for reasoning about behavior. "leaves Theory of Mind (ToM) performance intact"
Collections
Sign up for free to add this paper to one or more collections.