The Benchmark
Anthropic released BioMysteryBench, a curated set of 1,200 bioinformatics challenges designed to test whether AI models can match human expert performance on real biological research tasks. The benchmark covers protein structure prediction, gene expression analysis, drug target identification, and pathway analysis.
The twist: each challenge was derived from recently published papers (post-2025), ensuring models couldn't simply memorize training data. Human experts — PhD-level bioinformaticians — also solved the same challenges under timed conditions.
Results
AI models outperformed humans in some areas but fell short in others:
- Protein structure: AI 94% accuracy vs humans 87% — AI excels at pattern recognition
- Gene expression: AI 78% vs humans 91% — humans better at interpreting ambiguous data
- Drug targets: AI 82% vs humans 84% — roughly tied, with AI faster
- Pathway analysis: AI 71% vs humans 89% — the biggest gap, requiring biological intuition
The pathway analysis gap is telling. It's the task that requires connecting disparate biological knowledge — understanding that a protein's role in one pathway affects its behavior in another. AI models are great at pattern matching but struggle with the kind of integrative reasoning that experienced biologists do intuitively. This suggests the most valuable AI-human collaboration in biology is AI for data processing + humans for interpretation.