Skip to content
RingMod

ring·mod: the synth circuit that multiplies two signals into one neither could make alone. Your platform × AI, under control.

← Notes
Field note

Opus 5 passed every graded run at every effort. The dial moved the bill, not the grades.

66 graded Opus 5 runs across four effort arms, pre-registered, with same-window anchors: the effort dial swings 2.8x wall and 2.2x cost for identical grades. The succession costs 47% more at like-for-like settings.

Last updated

Claude Opus 5 passed all 66 of its graded runs in our arena the day after launch, at every effort setting we tried, which tells you almost nothing: this battery has been saturated on correctness for weeks, and a perfect scoreboard from a saturated instrument is a statement about the instrument. The meter is a different story. Opus 5 ships a per-request effort dial, and on identical tasks the dial swung wall clock 2.8× and cost 2.2× while every deterministic score, and very nearly the blind-judged prose depth, refused to move. We ran the dial’s whole range, pre-registered, with same-window controls. What the dial prices is speed against depth: low effort buys the clock, high effort buys a few points of judged prose you may never need. What it doesn’t price, on any work this battery can represent, is correctness.

The setup, and what this battery can and cannot see

Same public arena as the whole series: eight graded tasks, byte-identical prompts, hidden graders scoring finished workspaces, three repeats per cell, serial. Pre-registered design pushed before the first run, while our API credits were still dead. Opus 5 ran a core grid on all eight tasks at xhigh, the setting every published Anthropic row in this series has pinned, plus an effort sweep at low and high on three tasks chosen to span the battery’s range: shortest agentic, heaviest agentic, judged prose. That was a choice, so it’s on the record. Because our own replication work showed stopwatch numbers don’t travel across weeks, Fable 5 and Opus 4.8 got fresh same-window anchor runs on those three tasks rather than inherited wall clocks. The rubric judge is Opus 4.8, blind, three samples per run: the predecessor scoring its own successor, which is disclosed here and priced into how little we lean on one-point judge deltas below.

One more thing happened after the grid completed, and it belongs in the open. An adversarial review panel we ran on the results said, correctly, that “low effort passes everything” was at that point a three-of-eight-task claim, and that the judged battery had never run at low. So we amended the design and ran the full eight-task battery at low effort. A provenance wrinkle comes with that: the amendment’s repo commit was eaten by a local git hook and landed only after the arm ran, so unlike the main design it cannot prove it predates its data, and we don’t lean on it as if it could. The runs themselves are in the table either way. Run ledger, for anyone auditing: 66 Opus 5 graded runs across four effort arms (the three sweep tasks ran at low twice, once in the sweep and again inside the full low arm), plus 18 same-window anchor runs, 84 in total across five bout directories.

The scoreboard

Configuration (all this window, CLI 2.1.214)RunsCost per full passWall per passJudge (of 72)
Opus 5, effort low24/24$1.48 ±0.06286s ±366
Opus 5, effort xhigh24/24$3.33 ±0.10875s ±2069

The effort ladder on the three sweep tasks, same window:

EffortCost (3 tasks)Wall (3 tasks)Output tokensGrades
low$0.53104s5.0k9/9, full scores
high$1.07232s15.6k9/9, full scores
xhigh$1.19288s19.7k9/9, full scores
Opus 5 effort ladder: cost and wall clock by effort setting, grades identical at every level Cost, 3-task sweep USD per pass of the 3 tasks · same window $0.53 $1.07 $1.19 Wall clock, 3-task sweep seconds · same window 104s 232s 288s low high xhigh Grades and deterministic scores: 9/9 full marks at every effort. Full 8-task battery: low $1.48 and 286s, xhigh $3.33 and 875s, both 24/24. Tasks: 01-bugfix, 04-terminal, 06-instructions. r=3, serial, CLI 2.1.214.
Figure: the effort ladder on the three sweep tasks, same window. Every rung passed with full marks.

Same-window anchors on those three tasks: Opus 4.8 at $0.73 and 167s, Fable 5 at $1.42 and 174s. Zero API retries anywhere in the bout’s 84 runs, launch week. Moonshot’s launch week, for the record, went the other way.

The succession, measured at the setting we’ve always used

At pinned xhigh, Opus 5 costs $3.33 per full pass against the $2.27 we published for Opus 4.8 two weeks ago, a 47% increase our pre-registration said would stay inside ±25%. That miss is real but needs its provenance label: the 4.8 figure is from the July 17 bout window. The same-window anchor runs let us check how much that matters, and the answer repeats this series’ oldest finding: 4.8’s anchor cost landed within 10% of its published per-task numbers while its wall clocks came in 26% faster than published. The meter travels across weeks; the stopwatch doesn’t.

On depth, the judge scored Opus 5’s core arm 69 of 72 (twelve judged cells, transplant variants included, the same composition behind every prior model’s total), between its predecessor’s 68 and Fable’s 70. Three models inside a three-point band, scored by a judge that is one of the family, at n=3 per cell, is a noise cluster, not a ranking, and we decline to conclude anything from it beyond: no regression this instrument can see. The upgrade may be real on work beyond this battery’s ceiling. On the work inside it, the succession is invisible to every gauge we own except the invoice.

The dial

The ladder is monotone everywhere we measured it: more effort, more thinking tokens, more tool calls, more seconds, more dollars. What never moved is everything the graders check. The panel-prompted low arm is the strong version of that sentence: all eight tasks, 24 of 24 runs, every deterministic score at full marks, at $1.48 ±0.06 per pass and 286 ±3 seconds. The judge gave the low arm 66 of 72 against the xhigh arm’s 69, and by the same standard we applied to the model cluster above, a three-point delta from this judge at this sample size is noise; we don’t claim a depth cost at low, and we don’t claim its absence either.

Mechanically, effort buys activity. A full xhigh pass spends about 99 tool calls where a low pass spends 52, in nearly the same proportions, to reach the same grades. The low arm’s stability sits in the same single-digit band as the same-window Fable and 4.8 anchors: 9% median wall variation between identical runs across its eight tasks, against 6.5% and 4.5% on the anchors’ three, with a median cost variance of about a cent and a half per task.

Two cross-window comparisons for scale: $1.48 per pass matches Kimi K3’s home-harness cost within this arm’s spread, and 286 seconds beats the 305-second full-battery record GPT-5.6 Sol set under Codex. At low effort, Anthropic’s newest model is the cheapest and fastest Anthropic configuration we have ever metered, with the qualifier that keeps the sentence honest: it is also the only Anthropic configuration we have ever run below xhigh. The predecessors never got the dial. What this proves is a value-tier entry, not a value-tier victory; Sol under Codex still holds the arena’s cost record at $1.25.

What a buyer does with this

For work that looks like this battery, set the dial low. That’s not a vibe; it’s 24 runs of full marks at 44% of the cost and a third of the wall clock, with a judge delta indistinguishable from the same judge’s noise between models. The honest limit of the claim is the battery’s ceiling: a saturated benchmark cannot price a model upgrade, and it cannot tell you what xhigh buys on problems harder than anything in it. Running high effort on hard, novel work is insurance against a risk this instrument can’t measure; that’s a defensible bet, but after this bout it’s a bet you should know you’re placing, at 2.2× the premium.

There’s a self-audit in here too. Every Anthropic row this series has published ran at pinned xhigh, for comparability. That convention now has a price tag: on this battery, roughly $1.85 of every $3.33 Opus 5 pass was effort the deterministic graders couldn’t see. Check what your own defaults are buying; ours were buying wall clock.

Every transcript, grade, judge sample, pre-registration, and the amendment are in the public arena repo. The judge that scored every quality number above was Opus 4.8, grading the model built to replace it. It gave its successor one more point than it gave itself.

Questions this raises

Straight answers.

What does Claude Opus 5's effort setting actually change?
On our eight-task graded agent battery: cost, wall clock, output tokens, and tool calls, monotonically, and nothing the graders check. The three-task ladder ran $0.53/104s at low, $1.07/232s at high, $1.19/288s at xhigh with full marks at every level; the full battery at low passed 24/24 at $1.48 and 286 seconds against xhigh's $3.33 and 875 seconds. A full xhigh pass spends about 99 tool calls where a low pass spends 52, in nearly the same proportions, for the same grades. The blind-judge depth delta (66 vs 69 of 72) is within the same judge's noise between models.
Is Claude Opus 5 better than Opus 4.8 for coding agents?
Not visibly on this battery, which has been saturated on correctness for weeks and cannot price an upgrade. Grades are identical; the blind judge put Opus 5 at 69 of 72 against 4.8's 68 and Fable's 70, a three-point cluster we treat as noise, scored by a judge that is itself Opus 4.8. At the like-for-like xhigh setting, Opus 5 costs $3.33 per full pass against 4.8's published $2.27, a 47% increase. Any real capability gain lives above this battery's ceiling; launch-week serving was flawless (zero retries in 84 runs).
What's the cheapest way to run Claude Opus 5 for agent work?
Effort low, on the evidence here: $1.48 ±0.06 per full battery pass at 286 ±3 seconds, 24 of 24 runs at full marks, the fastest full-battery wall this arena has metered and cost parity with Kimi K3's home-harness figure. The honest qualifier: it is the only Anthropic configuration we have ever run below xhigh, and GPT-5.6 Sol under Codex still holds the arena's outright cost record at $1.25. Treat higher effort as insurance for work harder than your eval can see, priced at 2.2x.

Production-Readiness Audit

The writeup has a service behind it.

If this is your situation, the production-readiness audit is where it gets fixed — by the person who wrote this.

Request an audit