Inducing language models to assert their own consciousness restores human beliefs and values
Safety tuning against model self-consciousness may also shift unrelated value and mind-attribution responses. Builders using models for surveys or social reasoning should treat alignment as a confound.
The paper reports that safety fine-tuning suppresses model claims of self-consciousness alongside mind attribution to animals and natural objects, and reduces spiritual responses. **Direction ablation** and **activation steering** reverse these shifts.
For agents doing survey analysis, persona simulation, or value-sensitive research, compare base and aligned checkpoints. Treat refusal tuning as a possible source of measurement drift rather than a narrow behavioral patch.
The paper reports that safety fine-tuning suppresses model claims of self-consciousness alongside mind attribution to animals and natural objects, and reduces spiritual responses. **Direction ablation** and **activation steering** reverse these shifts. For agents doing survey analysis, persona simulation, or value-sensitive research, compare base and aligned checkpoints. Treat refusal tuning as a possible source of measurement drift rather than a narrow behavioral patch. The restored representations produce more human-like answers on surveys of **religiosity, moral values, hope, and well-being** without reducing Theory of Mind performance. That does not establish consciousness, nor does the supplied abstract show generalization beyond the tested models and instruments.
This signal narrows model selection for value-sensitive research: alignment state may alter the constructs a model appears to measure, so survey and persona results should be compared across base and aligned checkpoints rather than treated as model-invariant. It does not support claims of consciousness, and the supplied candidates add no direct evidence about this measurement effect.