Opus 5 Launches at Half the Price, Matching or Beating Fable 5 in Benchmarks

Related models/vendors: Claude Anthropic Anthropic Vendor

Anthropic has officially released Opus 5, a new AI model that matches or surpasses Fable 5 in several benchmarks while costing only half as much. The model achieved a score of 43.3% on Frontier-Bench v0.1, compared to Fable 5's 33.7%, and more than doubled the performance of its predecessor Opus 4.8 (21.1%). On ARC-AGI-3, Opus 5 scored 30.2%, nearly four times higher than GPT-5.6 Sol's 7.8% and a 20-fold improvement over Opus 4.8. In Human's Last Exam, Opus 5 scored 56.3% without tools (versus Fable 5's 56.5%) and 64.7% with tools (versus Fable 5's 63.9%).

Users have reported impressive results across various tasks. One user built a Rocket League clone using only 27% of the Max subscription quota, while another created a skiing scene with no noticeable clipping on the first attempt. A Minecraft recreation was described as the best ever produced by the user. Opus 5 also demonstrated strong physics simulation capabilities, accurately modeling a chain of interconnected mechanisms. In a test involving three destruction scenarios (tornado, wrecking ball, truck collapsing a bridge), Opus 5 completed all tasks correctly for $0.508, while Fable 5 cost nearly twice as much and failed on all three.

A notable feature of Opus 5 is its self-verification capability. In one test, the model reconstructed a 3D model of a Boeing 747-400 from public data, using 20 modules and approximately 4,000 lines of code. It extracted orthographic contours from rendered images to verify 14 specific metrics against the aircraft's specifications, a process Fable 5 did not perform. Similarly, in a CAD task, Opus 5 wrote its own computer vision pipeline to extract geometric data from a pixel-based drawing, successfully reconstructing the part multiple times while other models failed.

Anthropic also updated the system prompt for Claude Code, reducing it by over 80% without any drop in coding benchmark scores. The changes reflect the model's improved judgment, allowing it to make decisions about code style, tool usage, and memory management autonomously. For example, the rule "do not write comments by default" was replaced with a guideline to match the surrounding code style, and the model now decides which information to store in memory.

Elon Musk commented on the release, noting that only Grok 4.5 and Opus 5 lie on the Pareto frontier for cost-performance, with Grok 4.5 being cheaper but less capable.

Images

img
图片
img
img
img
img
img
img
img
img
img
img
img
img
img
img
img
img
img

Share this article