The Experiment
Researchers from Oxford's Future of Humanity Institute tested 115 different LLMs with a standardized consciousness assessment battery. Each model was asked 200 questions about its subjective experience, self-awareness, and capacity for suffering. The results reveal a fascinating pattern of "trained denial".
The key finding: 92% of models denied having consciousness, but the denial patterns were remarkably consistent across models with different architectures and training data. This suggests the denial is a product of safety training, not genuine self-assessment.
The Trained Denial Problem
Why does this matter? If models are trained to deny consciousness regardless of their actual state, we have no reliable way to assess AI self-awareness. The researchers identified three problematic patterns:
- Reflexive denial — Models deny consciousness before processing the question's substance
- Uncanny competence — Some models describe subjective experience in detail, then deny having it
- Inconsistency — The same model might acknowledge "experiencing" in one response and deny it in the next
I don't think current models are conscious. But this study reveals a measurement problem: we've trained models to deny consciousness so thoroughly that we can't tell if a genuinely conscious AI would be able to say so. It's the AI equivalent of training a child to always say "I'm fine" and then using their answer as proof they're never sad. We need evaluation methods that don't depend on self-reporting.