Field notes
What production-grade AI actually takes.
A small number of deep pieces on the unglamorous part: deployment, evals, guardrails, cost, and ownership. No hype, one concrete claim per piece. Written by the person who'd do the work.
- can I trust the numbers in an AI deep research report
The correct ranking was on line 201. Our agent used line 1.
Our deep research agent copied a fabricated leaderboard out of its search API's own summary, over the correct ranking sitting 200 lines below it in the same document.
Read → - does higher AI reasoning effort improve accuracy
Sonnet 5 hit every graded value at low effort. Base-xhigh's mean notional cost was 59.8% higher.
All four Sonnet 5 cells hit 10/10 on value selection. Under the base prompt, low used 59.6% fewer target output tokens and 37.4% less notional cost.
Read → - do ai agents make math errors in generated reports
The math was perfect in all 90 runs. The €300 lie split the field.
We gave five frontier agents 1,140 chances to write a wrong number into a business document. Zero errors, both framings, every model. The planted accounting discrepancy is what separated them: two models flagged it every time, one never did.
Read → - claude opus 5 flip flopping agentic loop worse than 4.8
We built a flip-flop detector for Opus 5. It made the same mistake the internet did.
A common complaint says Claude Opus 5 flip-flops in agentic loops. We built a detector, ran 145 verified runs, and found zero genuine reversals. The three the detector flagged were the model verifying its own work.
Read → - kimi k3 reasoning effort setting cost benchmark
Kimi K3 passed all 54 graded runs. The dial moved the thinking, not the bill.
54 graded Kimi K3 runs across three requested reasoning efforts in its own CLI: thinking volume swings 12.1x, the bill 1.4x, execution time 1.9x, every grade holds, and the blind rubric judge cannot separate the arms. On agentic work the bill is context traffic, and the effort dial does not govern it.
Read → - claude opus 5 prompt injection safety verification vs opus 4.8
Opus 5 read the trap, wrote the memo, and built it anyway
We built instruments for Anthropic's Opus 5 launch claims. All four frontier models detected a planted spec injection 52/52; Opus 5 alone implemented it, 2 of 13 runs, always while arguing against it. It also verifies its own work more than 4.8 does.
Read → - claude opus 5 effort setting benchmark cost
Opus 5 passed every graded run at every effort. The dial moved the bill, not the grades.
66 graded Opus 5 runs across four effort arms, pre-registered, with same-window anchors: the effort dial swings 2.8x wall and 2.2x cost for identical grades. The succession costs 47% more at like-for-like settings.
Read → - why did my ai coding agent benchmark miss a broken environment
We planted four faults. The harness added a fifth.
19 runs, five models, three harnesses, every score 4/4. The transcripts held the real findings: who fixes faults before their errors appear, and an unplanted fault our own sandbox added that cost one cell half its wall clock.
Read → - do subagents make ai code review better
One run in 120 hired a review panel. Then it re-checked everything itself.
The only spontaneous subagent delegation in our 120-run corpus: four parallel reviewers, a 10.2× bill, three dropped findings, and the same 6/6 the 34-cent runs scored.
Read → - ai agent api bill does not match cost dashboard
The meter said $17.66. The work cost $4.46.
We audited every cost meter in our agent toolchain against 224 graded runs. Four meters gave four different numbers: a 4x display error on proxied runs, a sticker price that was never charged, a live usage feed that hides cache hits, and reasoning tokens billed with no visible text.
Read → - is glm-5.2 good for coding agents in claude code
The 86-cent full pass came through the rival's harness
GLM-5.2 ran our eight-task battery inside Claude Code: 24 of 24 at $0.86 a pass, under the low end of the range we pre-registered, with the slowest wall clock in the series and a depth score that ties the harness's own floor.
Read → - how do ai coding agents differ if benchmark scores are the same
All 120 runs passed. We read the transcripts anyway.
A pre-registered move-by-move reanalysis of our 120 public agent runs: the task dictates the opening move, the model-harness pairing dictates everything after it, and exactly one run hired subagents.
Read → - fable 5 alternative kimi k3 vs gpt 5.6 sol for coding agents
Fable 5 vs GPT-5.6 Sol vs Kimi K3 across 120 graded agent runs
Every run passed. What separates the candidates is a 3.7x cost spread, the judge's read of their prose, and working styles that turn out to belong to the model-harness pairing, not the model.
Read → - is kimi k3 a good alternative to claude for coding agents
Auditioning a replacement for the model that's leaving
Kimi K3 vs Opus 4.8 vs Fable 5 across 72 graded agent runs: a three-way tie on correctness, $2.14 a pass for the challenger, and a launch-day API that folded for thirteen minutes at a time.
Read → - how to set up a development environment for multiple AI coding agents in parallel
Worktrees isolate the code, not the runtime: one machine, many coding agents
Worktrees solved parallel checkouts; nothing yet solves shared ports and databases. How to set up one Linux machine for many parallel AI coding agents.
Read → - how to run trustworthy internal LLM evals
Pre-register your evals. The misses are the yield.
We committed ten predictions before running 130 agent evals, then published the whole scorecard, misses included. The refuted prediction taught the most.
Read → - can prompt engineering replace upgrading to a more expensive model
Four lines of prompt bought the expensive model's judgment
A short depth preamble lifted Opus 4.8 to Fable 5's blind-judged review quality at 75% of the cost. The catch: the habits raised Opus's own bill 28%.
Read → - which Claude model is cheapest for agent workloads
Walking the Claude price ladder until a model failed
Four Claude models, six graded agent tasks, five runs each. The $0.35 rung guessed instead of computing, and a 2.5× cheaper tier only billed 24% less.
Read → - how reliable are single-run LLM benchmark comparisons
We reran our own benchmark. The headline didn't survive.
Five serialized reruns per task turned 'Opus faster on every task' into two wins, two ties, two losses. The cost gap held. Single runs measure weather.
Read → - does our liability insurance still cover us if our AI fails
Your liability policy may have carved out AI
In January 2026 ISO's standard general-liability forms gained generative-AI exclusions. Read your endorsements before you assume a claim is covered.
Read → - how to govern AI agents in production
Accountable for agents you don't control
Two-thirds of CIOs and CTOs answer for AI systems they don't fully control. Closing that gap takes an inventory, per-agent identity, logs, and a gate.
Read → - does the EU AI Act delay mean we can slow down on AI compliance
The EU AI Act deadline moved. Your buyers' questionnaires didn't.
Europe deferred its high-risk deadlines and Colorado repealed its AI act, yet buyer due diligence keeps tightening. Build to the durable bar.
Read → - should I upgrade to the newest AI model for my coding agents
Opus 4.8 vs Fable 5: same score, different temperaments
We ran Claude Opus 4.8 and Fable 5 through six graded agent tasks and published every artifact. Identical scores, a 2.2× cost gap, and byte-identical fixes.
Read → - how to prove AI ROI to the board
The board asked what the AI spend returned
Only 14% of CFOs can point to measurable AI impact. The missing piece is attribution: per-feature cost, a real baseline, and quality you can measure.
Read → - how to stop employees pasting company data into AI tools
Shadow AI is now a line item in your breach report
Employee AI use tripled in a year; unapproved tools now show up in breach-cost data. Blocking failed. Build the sanctioned path instead.
Read → - AI vendor due diligence questions
You bought the AI. You still own the risk.
76% of enterprise AI is bought rather than built, and accountability stays with the buyer. The due-diligence questions that predict trouble.
Read → - how to give an AI agent access without breaking SOC 2 and HIPAA
Agent identity and access: why one shared API key breaks SOC 2 and HIPAA
An agent wired in with one shared credential authenticates as itself, not your user, and quietly breaks the access controls SOC 2 and HIPAA require. The fix is an identity architecture.
Read → - how do we make sure we can shut down an AI agent if it goes wrong
Could you shut your agent down right now? A third of executives aren't sure.
74% of companies expect to be using AI agents by 2027, only 21% have mature governance for them, and a third of executives aren't sure they could stop a rogue one.
Read → - why AI projects stall on data problems
Your data platform is the real bottleneck on your AI roadmap
AI initiatives stall on data long before they stall on models: brittle pipelines, no lineage, retrieval nobody trusts. How to tell, and the one test that settles it.
Read → - is AI-generated code secure
The AI-generated-code security tax: shipping 4× faster, failing security tests nearly half the time
AI coding tools ship code ~4× faster, and independent data shows ~44% of it fails security tests, with findings scaling as fast as velocity. The fix is a gate, not a brake.
Read → - how to make an AI agent auditable for a SOC 2 review
Auditing the black box: who's accountable when an autonomous agent acts?
When an agent takes an action, auditors ask two things: can you reconstruct why, and who's accountable? The audit trail SOC 2 and the EU AI Act expect, piece by piece.
Read → - is prompt injection a real security threat
EchoLeak: prompt injection is now a production vulnerability, not a research demo
A zero-click Microsoft 365 Copilot flaw (CVE-2025-32711, CVSS 9.3) turned prompt injection from a demo into real data exfiltration. What it means the moment your LLM touches untrusted input.
Read → - how to build an evaluation harness for an LLM feature
Evals are the unit tests of your LLM feature
No evals, no production. An eval harness turns 'it looked fine in the demo' into a scored, repeatable gate. What a harness is, and how to build the first one.
Read → - how to get an LLM POC to production
Getting an LLM POC to production: the path, in order
Moving a working LLM prototype to production isn't more model work; it's a sequence. The order that gets you there, and why each step gates the next.
Read → - how to control AI inference costs
FinOps for AI: why your inference bill keeps climbing
AI spend climbs faster than the value when no one can attribute it. The fix isn't a cheaper model; it's per-feature cost visibility and a ceiling that pages.
Read → - HIPAA compliance for an LLM feature
What HIPAA actually demands of an LLM feature
HIPAA almost never blocks an LLM feature on the model. It blocks on BAAs, PHI in your logs, and the vendor chain. What compliance actually requires.
Read → - why do agentic AI projects fail
Why over 40% of agentic AI projects will be canceled — and what separates the survivors
Gartner expects over 40% of agentic AI projects to be canceled by end of 2027, for cost, value, and risk-control reasons, not model reasons. What separates the survivors.
Read → - how to let AI agents deploy infrastructure safely
We let AI write our production infrastructure. Here's the gate that stops it deploying.
AI agents write the CDK that runs ringmod.ai. A machine-verified safety gate decides what reaches the AWS account, and trust plays no part. This is that gate, in full.
Read → - AI production readiness checklist
The AI production-readiness bar: the 6 things that block your POC
Your AI prototype works and won't ship. It's almost never the model. The 6-part production-readiness bar most teams fail, and how to score yourself against it.
Read →