Editorial illustration for Anthropic benchmark says Claude matches experts, 23 tasks remain ambiguous
Anthropic benchmark says Claude matches experts, 23...
Anthropic asked a new version of its Claude model to solve 99 problems in bioinformatics. On 76 tasks that at least one human expert could handle, Claude matched their performance. On the 23 that stumped every expert, it got about a third right.
The catch, according to the company's analysis, is consistency. On easier tasks, Claude usually got all attempts right or wrong. On the hardest ones, correct answers appeared to be lucky guesses.
Anthropic split the tasks into two groups: 76 were considered "human-solvable" because at least one out of up to five experts found the correct answer. Another 23 tasks stumped every expert. On the solvable problems, Claude now matches human expert performance, according to Anthropic.
On the hard problems that none of the selected experts could solve, Claude Mythos Preview achieves a 30 percent success rate. However, a consistency analysis that Anthropic had Claude Mythos Preview run on itself paints a more nuanced picture. On the solvable problems, Claude almost always either gets all five attempts right or none at all.
On the hard problems, successes typically come in just one or two out of five attempts. The model stumbles onto a lucky solution path rather than following a reproducible strategy.
The 30% score on impossible-seeming tasks is the flashy number. Its meaning falls apart under scrutiny. Claude either knows a routine answer or it doesn't.
For the 23 ambiguous problems, the model lacks a reliable method. It just guesses. Matching experts on known tasks is one thing.
The real test is the work nobody can do yet. There, Claude provides no new insight, only random noise.
Common Questions Answered
How did Claude perform on the 76 bioinformatics tasks that human experts could solve?
Claude matched the performance of human experts on 76 out of 99 bioinformatics problems where at least one expert could provide a correct solution. This demonstrates that the model has achieved parity with human expertise on established, routine tasks in the bioinformatics domain.
What does Anthropic's analysis reveal about Claude's consistency on difficult tasks?
According to Anthropic's analysis, Claude shows a significant consistency problem on the hardest tasks, where correct answers appear to be lucky guesses rather than reliable solutions. On easier tasks, the model typically gets all attempts either right or wrong, but this consistency breaks down when facing ambiguous problems.
Why is Claude's 30% score on the 23 ambiguous bioinformatics problems misleading?
The 30% score on impossible-seeming tasks is misleading because Claude lacks a reliable method for solving these problems and is essentially guessing randomly. The model either knows a routine answer or it doesn't, meaning the correct answers on ambiguous tasks provide no genuine insight into new problem-solving capabilities.
What is the key limitation of Claude when tested on novel bioinformatics problems?
Claude's key limitation is that it provides no new insight on bioinformatics problems that human experts cannot yet solve, only random noise and unreliable guesses. While the model matches expert performance on known tasks, it fails to demonstrate genuine advancement in tackling truly novel scientific challenges.
Further Reading
- Evaluating Claude's bioinformatics research capabilities with BioMysteryBench — Anthropic
- Introducing Claude Opus 4.7 — Anthropic
- Claude Mythos Preview System Card — Anthropic
- Anthropic Economic Index report: Economic primitives — Anthropic