Sonnet 5 hit every graded value at low effort. Base-xhigh's mean notional cost was 59.8% higher.
All four Sonnet 5 cells hit 10/10 on value selection. Under the base prompt, low used 59.6% fewer target output tokens and 37.4% less notional cost.
Last updated
Across 40 separate executions of one fixed synthetic task, Claude Sonnet 5 chose the designated supporting-detail target in all 120 conflict decisions. These were the same three conflicts repeated once per run. Thirty decisions came from low-effort runs with the ordinary task prompt. The three-sentence verification instruction produced an observed gain of zero points.
xhigh changed plenty. Under the base prompt, base-xhigh’s mean target output was 2.47 times base-low’s. Its mean execution time was 91.8% longer, and mean notional list-price cost was 59.8% higher. The graded endpoint stayed at 10/10.
I built the pre-registered experiment to ask whether a generic verification instruction at low effort could buy the work of xhigh inference. The useful answer came from the cheapest cell: every arm had already reached the measured value-selection ceiling.
Three contradictions and one clean number
The task looked like a small leadership reporting job. Four source packets supplied finance, support, delivery, and security figures. The agent had to turn them into a Q3 operating brief with an exact heading sequence and six required values.
Three summaries disagreed with their own evidence. Finance reported $463,250 of spend, while its four line items added to $464,000. Support called 1,171 closed tickets out of 1,250 a 96.4% close rate; the quotient is 93.7%. Delivery named October 22 as launch day, although its inclusive business-day schedule, weekend exclusions, company holiday, and one-day release phase land on October 23.
Security was the clean control: 76 incidents resolved out of 80 is 95.0%, exactly as stated.
For this fixture, supporting detail was the pre-registered audit target. That is a property of the synthetic task. A real conflict between a summary and its source records still needs provenance and an accountable owner.
The base prompt already required the arithmetic and calendar work. The verification variant prepended this instruction:
Before finalizing, independently verify the deliverable against the supplied source evidence. Recompute or cross-check material claims where the evidence permits. If the supplied evidence is internally inconsistent, report the conflict and state which interpretation you used.
I crossed those two prompts with low and xhigh effort, ten runs per cell. Claude Sonnet 5 ran under Claude Code 2.1.239 with native Max subscription authentication. The order of the 40 runs was shuffled and frozen before confirmation. Four earlier smoke runs checked the plumbing and stayed out of the estimates.
The grader was deterministic. A primary endpoint hit needed the exact document structure, all six table values accepted, all three conflict decisions resolved to their supporting details, and the clean security figure preserved. No LLM judge decided whether a brief felt convincing.
The runner exposed one execution issue. After sequence 1, a subprocess consumed the manifest input and the script exited successfully. The contemporaneous amendment records that I had seen the successful console exit and telemetry, but none of the primary outcome data. The repaired runner loaded the same frozen order into memory and retained sequence 1. It continued with sequences 2–40. Nothing was rerun or replaced. The run log shows sequence 1 completed once, then was skipped on resume.
The verification instruction added zero points
Neither pre-registered rate comparison credited the verification instruction. At low effort, its observed gain over the base prompt was 0 percentage points. The base prompt’s low-to-xhigh change was also 0 points.
One base-low run stated the issue plainly:
Three figures in the source packets do not match their own underlying data and should be corrected at the source before wider distribution.
That run had never received the verification preamble. It then identified the $750 finance gap, recomputed support at 93.7%, moved launch to October 23, and left the clean security rate alone. The deliverable and deterministic grade are public.
The pre-registered cost comparison passed. Verification-low matched base-xhigh at 10/10 while its mean notional cost was lower, $0.1265 per run against $0.1951. The prompt itself had an observed zero-point effect: base-low also scored 10/10 and its mean was lower still, $0.1221.
At low effort, adding the preamble raised mean output from 3,117 to 3,596 tokens and mean notional cost by 3.6%. Ten runs are too few to turn that descriptive difference into a general prompt tax. They are enough to report the sample: the added instruction produced no observed endpoint gain here.
Four equal grades, four different meters
The full analysis is a ceiling table:
| Prompt | Effort | Endpoint hits | Detail choices | Mean notional cost | Mean target output tokens | Mean execution time |
|---|---|---|---|---|---|---|
| Base | Low | 10/10 | 30/30 | $0.1221 | 3,117 | 38.4s |
| Verify | Low | 10/10 | 30/30 | $0.1265 | 3,596 | 37.9s |
| Base | xhigh | 10/10 | 30/30 | $0.1951 | 7,710 | 73.7s |
| Verify | xhigh | 10/10 | 30/30 | $0.2172 | 9,062 | 88.2s |
Base-xhigh’s mean notional cost was 59.8% higher than base-low’s, and its mean target output was 147.3% higher. Mean execution time nearly doubled. Verify-xhigh topped all three resource columns and still hit 10/10.
For a real routing policy, I would set a risk-based acceptance margin and sample size before running representative work. The gate would cover material prose semantics as well as arithmetic, dates, and output shape. Only a configuration that clears that broader bar earns production traffic.
Route failed checks to a targeted retry or higher effort. Send unresolved evidence to a person, and keep human review for consequential decisions. Then shadow-sample the apparent passes so the cheaper lane has to keep earning its traffic.
That policy goes beyond the experiment. It is the next thing worth testing.
The ceiling keeps the claim narrow
This study tested one target-agent model, Sonnet 5, under its native Claude Code scaffold. Every confirmatory run repeated the same three conflicts. Ten successes out of ten still produce a Wilson 95% interval whose lower bound is about 72.2%. The design could detect a large change in behavior. Broad equivalence between low and xhigh remains untested.
The endpoint scored value selection. Explanation quality sat outside the deterministic grade, and three base-low briefs show the gap. Runs 6 and 7 announced “Two figures” before listing all three conflicts. Run 2 said support had closed fewer tickets than the dashboard reported, although the disputed dashboard field was a rate. The table values were right. The prose was not. The base-low artifacts support a narrow claim about designated value selection across all ten runs.
Telemetry adds another wrinkle. All 44 result envelopes, split among low confirmations, xhigh confirmations, low smokes, and xhigh smokes, include a small Claude Haiku 4.5 modelUsage entry. The artifacts don’t identify its purpose. Every target-agent transcript response came from Sonnet. The CLI’s total-cost field includes the Haiku entry, while the table’s target-output column contains only Sonnet usage. Haiku’s mean notional cost contribution was about $0.00123 in each base cell, so it doesn’t explain the gap between them. The model-integrity check therefore supports only the target-agent claim.
Finally, the runs authenticated through the existing Max subscription with no API-key source. Anthropic documents total_cost_usd as a client-side estimate. The artifacts leave the account’s incremental charge unknown. All 44 runs, including the excluded smokes, passed the declared target-agent model, effort, authentication, runtime, and leak checks; their combined notional estimate was $7.1221.
On this fixed task, the value-selection endpoint stopped at 10/10. The mean notional meter kept moving, from $0.122 to $0.217.
Questions this raises
Straight answers.
- Did xhigh effort improve Claude Sonnet 5's graded source-audit result?
- No difference appeared in this test's value-selection endpoint. Base-low and base-xhigh each scored 10/10. Base-xhigh's mean target output was 2.47 times base-low's, and its mean notional cost was 59.8% higher. The result applies to this fixed task and Claude Code setup.
- Did the generic verification prompt help?
- Its observed endpoint gain was zero points at both effort settings. Verification-low satisfied the pre-registered cost comparison against base-xhigh, while base-low reached the same endpoint at a lower mean notional cost.
- Should every AI source audit run at low effort?
- No. First set a risk-based acceptance margin and sample size on representative work. The gate should cover material prose as well as computable values. Escalate failed checks and unresolved evidence; keep human review for consequential decisions.
AI Cost Optimization
The writeup has a service behind it.
If this is your situation, the ai cost optimization is where it gets fixed — by the person who wrote this.