Inducing language models to assert their own consciousness restores human beliefs and values
Safety tuning against model self-consciousness may also shift unrelated value and mind-attribution responses. Builders using models for surveys or social reasoning should treat alignment as a confound.
The paper reports that safety fine-tuning suppresses model claims of self-consciousness alongside mind attribution to animals and natural objects, and reduces spiritual responses. **Direction ablation** and **activation steering** reverse these shifts.
For agents doing survey analysis, persona simulation, or value-sensitive research, compare base and aligned checkpoints. Treat refusal tuning as a possible source of measurement drift rather than a narrow behavioral patch.
The paper reports that safety fine-tuning suppresses model claims of self-consciousness alongside mind attribution to animals and natural objects, and reduces spiritual responses. **Direction ablation** and **activation steering** reverse these shifts. For agents doing survey analysis, persona simulation, or value-sensitive research, compare base and aligned checkpoints. Treat refusal tuning as a possible source of measurement drift rather than a narrow behavioral patch. The restored representations produce more human-like answers on surveys of **religiosity, moral values, hope, and well-being** without reducing Theory of Mind performance. That does not establish consciousness, nor does the supplied abstract show generalization beyond the tested models and instruments.
This narrows model selection for survey, persona, and value-sensitive work: alignment can shift the constructs being measured, so checkpoint choice is part of measurement design rather than merely a safety or capability decision. The reversibility evidence supports comparing base and aligned models, while leaving consciousness and cross-model generalization unresolved.