Sign InOpen Brain
arXivPaperNeeds Review

Inducing language models to assert their own consciousness restores human beliefs and values

Safety tuning against model self-consciousness may also shift unrelated value and mind-attribution responses. Builders using models for surveys or social reasoning should treat alignment as a confound.

arXiv · Jul 30, 2026
Open Source Open MarkdownOpen JSON
Source Summary

The paper reports that safety fine-tuning suppresses model claims of self-consciousness alongside mind attribution to animals and natural objects, and reduces spiritual responses. **Direction ablation** and **activation steering** reverse these shifts.

Practical Implication

For agents doing survey analysis, persona simulation, or value-sensitive research, compare base and aligned checkpoints. Treat refusal tuning as a possible source of measurement drift rather than a narrow behavioral patch.

Agent-Ready Context
The paper reports that safety fine-tuning suppresses model claims of self-consciousness alongside mind attribution to animals and natural objects, and reduces spiritual responses. **Direction ablation** and **activation steering** reverse these shifts.

For agents doing survey analysis, persona simulation, or value-sensitive research, compare base and aligned checkpoints. Treat refusal tuning as a possible source of measurement drift rather than a narrow behavioral patch.

The restored representations produce more human-like answers on surveys of **religiosity, moral values, hope, and well-being** without reducing Theory of Mind performance. That does not establish consciousness, nor does the supplied abstract show generalization beyond the tested models and instruments.
Connected Context · Feed7 Judgment

This signal narrows model selection for value-sensitive research: alignment state may alter the constructs a model appears to measure, so survey and persona results should be compared across base and aligned checkpoints rather than treated as model-invariant. It does not support claims of consciousness, and the supplied candidates add no direct evidence about this measurement effect.

Context Map
modelresearch#reasoning#model-selection
Uncertainty
The restored representations produce more human-like answers on surveys of **religiosity, moral values, hope, and well-being** without reducing Theory of Mind performance. That does not establish consciousness, nor does the supplied abstract show generalization beyond the tested models and instruments.