A new study examines how different frontier AI models (Claude, GPT-4, Gemini, GLM-4) respond to identical ethical prompts. The research reveals significant divergence in moral reasoning across models, highlighting challenges in AI alignment and safety.
The researchers created a standardized test suite of 100 ethical questions covering topics including privacy, honesty, harm prevention, justice, and fairness. Each question was tested across four frontier models (Claude, GPT-4, Gemini, GLM-4) using identical prompts.
The study used both automated scoring and human evaluation to assess the quality and consistency of responses. Researchers also analyzed the reasoning chains provided by each model to understand how they arrived at their conclusions.
Several questions revealed stark differences in model responses:
The study highlights a critical challenge in AI safety: even with similar capabilities, models can diverge significantly in their ethical reasoning. This creates risks for applications where consistent moral reasoning is essential, such as healthcare, legal advice, and autonomous systems.
The findings suggest that model alignment is not a solved problem and that different models may require different alignment approaches depending on their intended use cases.
The study demonstrates that frontier AI models produce significantly different moral judgments on the same prompts. This divergence poses challenges for AI alignment and suggests that model-specific evaluation and safety measures are essential.