Skip to content
RingMod

ring·mod: the synth circuit that multiplies two signals into one neither could make alone. Your platform × AI, under control.

← Notes
Field note

The 86-cent full pass came through the rival's harness

GLM-5.2 ran our eight-task battery inside Claude Code: 24 of 24 at $0.86 a pass, under the low end of the range we pre-registered, with the slowest wall clock in the series and a depth score that ties the harness's own floor.

Last updated

Before the first graded run, I wrote down what GLM-5.2 would cost on our battery: between $0.90 and $1.60 per full pass. That was the optimistic case, built on open-weight list prices and the assumption that Z.ai’s prompt cache would hold up under an agent’s request cadence. It came in at $0.86. I was wrong on the low side, which is a first for this series, and 24 of 24 runs passed with full deterministic scores while doing it. The cheapest full pass this arena has recorded arrived through Claude Code, the harness GLM’s own vendor competes with.

ConfigurationRuns passedCost per full passWall per passJudge (of 72)
GLM-5.2 (Claude Code)24/24$0.861,153s60
Kimi K3 (Claude Code)24/24$2.04891s66
Opus 4.8 (Claude Code)24/24$2.27621s68
Fable 5 (Claude Code)24/24$4.62690s70
Four Claude Code configurations compared on cost, wall clock, and judge score GLM-5.2 Kimi K3 Opus 4.8 Fable 5 Cost per full pass USD · lower is better $0.86 ★ $2.04 $2.27 $4.62 Wall clock per pass seconds · lower is better 1,153s 891s 621s ★ 690s Judge score of 72 · higher is better 60 66 68 70 ★ GLM-5.2 Kimi K3 Opus 4.8 Fable 5 All four configurations ran in Claude Code; all runs passed. GLM row and its judge score measured 2026-07-22 (CLI 2.1.214); Kimi, Opus, and Fable rows from the published July bouts (CLI 2.1.212); the same-day Opus anchor ran $2.38 and 661s. Judge: Opus 4.8, blind, median of 3 samples.
Figure: the four Claude Code configurations on the three axes that separate them. Stars mark the best value per panel.

The comparison that matters is inside one harness, and the ladder note’s cheaper rungs don’t unseat this one: Haiku 4.5’s $0.35 pass ran a shorter six-task course and failed twice in thirty runs. Among clean passes of the full battery, in any harness we’ve tried, $0.86 is the new floor, under even Sol’s $1.25 under Codex.

The setup, and why this pairing is different

Z.ai is the first vendor in this series that chose the rival’s harness on purpose. Moonshot added an Anthropic-compatible shim for Kimi K3; OpenAI never offered one and Sol needed a local proxy. Z.ai ships its own free IDE (ZCode, pitched as the “Official Harness for GLM-5.2”) and, in the same breath, documents Claude Code as a first-class way to run its model. So this bout asks a narrow question: when a vendor deliberately targets someone else’s harness, does the seam show?

Mechanics as always: eight tasks including both transplant prompt variants, three repeats, strictly serialized, byte-identical prompts, hidden graders, design pre-registered and pushed before the first graded run. Because the published baselines ran on Claude Code 2.1.212 and today’s CLI is 2.1.214, an Opus 4.8 anchor arm ran the same eight tasks the same day: 8 of 8, $2.38 and 661 seconds against its published $2.27 and 621. Call that the drift envelope; single runs measure weather. One procedural note: two sessions of this bout were killed by outside interruptions, so one cell was re-run under a fixed rule (re-run only on external termination, partial runs archived unpublished, never graded). Standing disclosures apply: Anthropic’s harness, an Anthropic judge that is also a contestant, effort pinned at xhigh with unknown semantics on Z.ai’s gateway, and Fable 5 co-writing this. Everything is in the public bout directory.

The misses, first

The clock. I pre-registered a Kimi-like 800 to 1,100 seconds per pass. GLM took 1,153, which makes it the slowest Claude Code configuration we’ve timed, past even Sol through the translation proxy at 1,085. Honesty about the margin: summed across eight tasks whose run-to-run spread is large (one run of the transplant reporting task took 60 seconds, a repeat of the same cell 353), the per-pass figure carries roughly ±140 seconds of noise, so the band’s edge sits within one standard deviation. The direction still stands. And the minutes are serving rate, not verbosity: GLM’s output per pass is within 7% of the anchor’s (48,100 tokens against 45,100) while taking 1.7 times as long, and the CLI-reported API time meets or exceeds the whole wall clock on 23 of 24 runs, which leaves nothing for local work to explain. You buy the discount partly with the clock, and the tax was heavier than I predicted.

The Write marker. I pre-registered three tool-grammar markers for “moves like a native,” and the second one, in-place Edits at least matching whole-file Writes across the bout, missed: 25 Edits, 27 Writes. The post-hoc look (labeled as such) says the marker was badly built, not that GLM rewrites like away-game Kimi did: on the four code tasks GLM went 25 Edits to 11 Writes, and every single Write on the review and reporting tasks was a brand-new deliverable file, where Write is the only sensible tool. The Opus anchor went 9 and 9 on the same course. A marker the home team also trips is not measuring foreignness. It stays a miss on the scorecard; I’m telling you why I no longer trust the marker rather than quietly deleting it.

The other two markers, and what the cache did

The markers that were built properly came back native-shaped. GLM re-read an already-open file in 2 of 24 runs; when Sol ran in this harness it did that in 24 of 24, and Fable and Kimi in 0. It made zero task-management calls per run, against Sol’s 7.8. On the bug-fix task its whole walk reads like the anchor’s: list files, run the tests, read the two relevant files, two in-place edits, re-run the tests, write the solution note. Eight calls; Opus took seven and split nothing. What I can’t tell you is why. There is no neutral-harness GLM arm in this bout, so “trained for Claude Code” and “Claude Code’s scaffolding makes any competent model walk this way” both fit these logs. What the logs support is narrower and still useful: nothing about GLM fights this harness.

The cache result is the load-bearing one. Across 24 runs, 93.4% of billed input tokens priced at Z.ai’s $0.26-per-million cache-read rate: 5.42 million cache reads against 383 thousand uncached, a blended input rate of about $0.34 per million. That is why the $1.40 sticker survived contact with an agent that shoves the whole session back through the context window every turn. Same mechanics that made Kimi Code’s bill work (93% there), and the honest asterisks are that the usage numbers come from Z.ai’s own gateway envelope, our per-run cache-read counts swing 6x between repeats, and Z.ai currently charges nothing for cache writes or storage, “limited-time free,” their label. The $0.86 is real, computed from transcript usage exactly the way every cross-vendor figure in this series is, and part of it is promotional.

Where the discount shows

The blind judge (Opus 4.8, three samples, medians) put GLM at 60 of 72 on the four rubric tasks: level with Sol’s proxied Claude Code score, above only Sol’s 58 under Codex, below Kimi’s 66, Opus’s 68, Fable’s 70. The profile is the familiar cheap-model shape: full marks on locating and describing, dropped points on computing. Six run-medians on quantification, five of them 1s, with rationales like “bare dollar rankings with only one true computed figure.” I pre-registered 60 to 72 as the pass band and it landed on the doorsill, so H8 counts as a hit by the letter and I wouldn’t lean on it: an Anthropic judge scoring a rival at the exact floor, against incumbent scores from earlier sessions, is a soft signal, not a verdict. One confound I can’t rule out either: we pin effort at xhigh, Z.ai’s gateway may map that to nothing, and if so some of this thinness is an unmapped setting, not the model’s ceiling.

One prior result took a real dent here. The four-line transplant prompt had cleanly lifted interaction synthesis for four models across three vendors. On GLM, vendor four, the lift was murky: run medians went from 1, 0, 2 to 0, 2, 2. One point of movement, not the flip to straight 2s the streak was built on. The streak stops being a law here and goes back to being a finding.

Month two, during open-weight week

Zero retries, zero 429s, 24 of 24 runs, all in an off-peak Beijing window we chose on purpose. That’s the follow-up the Kimi audition said to run: test the serving a month after the launch crush. GLM’s own June was ugly (a GitHub issue documents a paid account at 100% failure for an hour), and its month two is boring in the way infrastructure should be. The timing sharpens it. This week Moonshot has new Kimi subscriptions frozen on capacity until the K3 weights land July 27, and DeepSeek V4 arrives July 24. Open-weight release week is loud; the open-weight model you can actually buy serving for, today, at the prices above, is the one that made no news at all.

So: cheapest clean pass we’ve measured anywhere, slowest clock in the harness, thinnest prose at the judging table, and a tool grammar with no visible seam. One boundary before the advice: a floor is a floor. Every frontier model passes this battery now, so cheapest to clear our bar is a fact about the bar, and whether it holds on your harder work is exactly what it can’t tell you. The deciding variable isn’t in the table either; it’s whether a human is watching the run. Overnight pipelines, fan-out readers, first-draft PRs: the 2.8x discount against the same-day Opus anchor compounds and nobody feels the extra eight minutes. Interactive work charges those minutes at your hourly rate, which buys a lot of tokens. Price the pairing against your own workload, both flags set, before the promotion ends or the next open-weight lands; pricing that pairing is the first hour of a production-readiness audit. My prediction band missed low once this week. It can miss again.

Questions this raises

Straight answers.

How much does GLM-5.2 cost to run in Claude Code?
On our eight-task graded battery, $0.86 per full pass at Z.ai's list prices ($1.40 per million input tokens, $4.40 output, $0.26 cache read), computed from transcript usage the same way every cross-vendor figure in this series is. That is the cheapest full pass of this battery we have recorded in any harness: GPT-5.6 Sol under Codex was $1.25, Kimi K3 $2.04 in Claude Code, Opus 4.8 $2.27, Fable 5 $4.62. Two caveats: 93.4% of GLM's billed input priced at the cache-read rate, and Z.ai currently charges nothing for cache writes or storage ('limited-time free'), so the sticker depends on a promotion holding.
Can GLM-5.2 run inside Claude Code?
Yes, first class. Z.ai serves an Anthropic-compatible endpoint (api.z.ai/api/anthropic) and markets Claude Code as a primary way to run GLM-5.2, alongside its own free ZCode IDE. Our harness points ANTHROPIC_BASE_URL at it with a per-model environment file that pins every internal model slot to glm-5.2. No proxy, no translation layer, and the tool logs show native-style behavior: it re-read an already-open file in 2 of 24 runs and never touched a task-management tool.
Is GLM-5.2 as good as Claude for agentic coding?
On correctness our battery cannot separate them: GLM-5.2 went 24 for 24 with full deterministic scores, exactly like Fable 5, Opus 4.8, and Kimi K3 before it. The separations are elsewhere. It is the slowest configuration we have timed in Claude Code (1,153 seconds per pass against Opus 4.8's 661 that same day), and a blind Opus judge scored its prose depth 60 of 72, tying the lowest Claude Code score in the series. Cheapest bill, slowest clock, thinnest analysis: which of those dominates depends on whether a human is waiting on the run.

Production-Readiness Audit

The writeup has a service behind it.

If this is your situation, the production-readiness audit is where it gets fixed — by the person who wrote this.

Request an audit