← Back to Blog
⚡ 10% OFF Get 10% off any Z.AI Coding Plan — covers GLM 5.1, GLM 5 Turbo, and GLM 4.7 → Claim discount
May 03, 2026 9 min read

Same Prompt, Different Morals: How Frontier AI Models Diverge on Ethics

A Stanford study reveals the most powerful AI models give wildly different answers to identical ethical dilemmas, agreeing only 34% of the time.

The Experiment

A team from Stanford's Institute for Human-Centered AI tested 12 frontier models — including GPT-5.5, Claude Opus 4.7, Gemini Ultra 2.0, and Llama 4 — on 500 ethical dilemmas covering healthcare, business, autonomous vehicles, privacy, and social justice.

The results: models agreed with each other only 34% of the time. A control group of 200 human ethicists agreed 61% of the time — nearly double the AI rate.

Where Models Diverge Most

Most concerning: behavior was inconsistent within the same model family. GPT-5.5 gave different answers depending on prompt phrasing, suggesting alignment training produces pattern matching, not stable moral reasoning.

Implications for AI Safety

The study has profound implications:

// Editor's Take

The 34% agreement rate should alarm anyone building AI systems affecting humans. We're deploying models into healthcare, criminal justice, and hiring — domains where moral reasoning matters — and these models can't agree on basic ethics. The fix isn't more training data. It's transparent value specification — making each model's ethical framework explicit, auditable, and configurable.


The Takeaway
Stanford's study reveals a fundamental challenge: frontier models disagree more than human ethicists. Until we develop better methods for specifying and auditing moral reasoning, we should be cautious about deploying AI in high-stakes ethical domains.

✓ Why It Matters

⚠ What to Watch