Anthropic’s Claude Opus 5 Shatters Reasoning Benchmark Record, But Gains Raise Questions About Breadth
Anthropic’s latest model scores 30.2% on ARC-AGI-3, nearly quadrupling the previous best, but researchers caution that the improvement may be narrower on unfamiliar tasks.
Read more