The experiment
One variable: the harness. Same model (Opus 5), same
machine (bench VPS, root), same prompt — byte-verbatim, card below. Three orchestration loops:
OpenCrabs' TUI loop (step structure, tool batching, context handling) spawning the claude CLI as its
provider; Claude Code's built-in agent loop — the OpenCrabs rows and the Claude Code row share that
identical engine and transport; and the Hermes TUI driving the Anthropic API directly through its own
gateway and session runtime. OpenCrabs ran twice (run 1 in the original bench workdir, run 2 in
a fresh isolated workdir with the box to itself) for repeatability; every other harness ran once.
All fired manually in-terminal, auto-approve on, no timeouts.
What the numbers say
Ranked by token efficiency — fewest tokens to
complete the task wins. The clock rewards burning parallel compute; the token count is what the
harness actually made the model read and write, so it's the price of the task. OpenCrabs TUI takes
the board on both runs — run 1: 10.43M tokens, $8.00, 37m30s; run 2: 10.82M tokens, $10.79,
41m01s. Repeatability: ±3.8% tokens, ±9.8% wall across two independent one-shots — and both runs
stay under Hermes. Hermes second — 12.51M tokens (1.2x ours), $15.65, 40m09s; the two
OpenCrabs walls straddle it (37m30s under, 41m01s just over by 52s), average 39m16s still faster —
but it loses the rank key that decides the podium either way. Claude Code's loop last —
55.07M tokens (5.3x ours), $46.59, 1h36m23s: 249 round trips re-reading a growing context, 97.6% of
its tokens cache reads. The cheapest run cost 17% of the priciest.
Quality parity, not just speed
All four builds loaded first try and shipped
the full v2 spec including the Visual Options menu; each self-verified headless and self-caught real
bugs. OpenCrabs run 1: 3 bugs (unwinnable OC-3 reactor hitbox, dead cover state, backwards film grain),
80/80 smoke checks. OpenCrabs run 2: 23 modules / 2,914 lines, spec values measured not assumed —
headshot multiplier 2.30x against the 2.3 spec, scanner piercing 3 stacked targets, launcher splash
+ 1.6 m/s knockback + 38 hp self-damage, recoil bloom driving the crosshair — and a full campaign
sweep OC-1→OC-6 with 0 console errors and 0 failed requests. Claude Code run: 4 bugs (W sign error, frame-rate-dependent friction, muzzle
computed in view space, doorway sealing), headless Chromium all-decks pass. Hermes run: five
headless-Chrome suites — 0 bytes allocated per frame across 1200 sim steps, 0.044 ms sim cost under
full combat load, 61–112 draw calls, ACES confirmed, full campaign OC-1→OVERLOCK with 0 errors — and
honest caveats stamped on its own work (SwiftShader software raster; the 1-shot-per-frame auto-fire
gate it chased down and correctly acquitted). Deliverables equal on all sides:
game + README + design doc + chapter-3 playthrough script.
Provenance & asterisks
OpenCrabs row = native usage_ledger, cost rolled
up by the harness at official Anthropic rates (Max sub — not billed per run). Claude Code row =
harvested from the claude jsonl transcript, cost computed at official rates. The OpenCrabs run began
with the operator's "continue" kickoff after a cancelled pre-switch leg; all figures cover the opus
leg only (operator ruling: the run counts from the model switch). The OpenCrabs run
also survived a parallel process writing into its workdir mid-run — it flagged the intrusion in its
own final message; the dir was later migrated to opus-5-och/. Its build size includes ~12MB of
headless-test artifacts (test/out PNGs, test-vendored three.js); the game proper is ~210KB, on par
with its twin. OpenCrabs run 2 = also native usage_ledger (10,821,679 tokens / $10.79), a clean
one-shot in a fresh isolated workdir with the box to itself — no parallel disturbance, and the
cleanest deliverable tree of the lane (~157KB, no test artifacts left behind). Hermes row = harvested from its own sqlite state.db (sessions + session_model_usage):
wall = prompt → final "Shipped and verified" message (22:16:58→22:57:07), no active/idle split
recorded so wall = span; post-run proxy Q&A (23:07–23:11) excluded. Its 12,506,671 tokens are the
harness's own counter and exclude a 1.41M background-review side task; cost computed at the same
official rate card (in 186 / cache-write 970,520 / cache-read 11,380,426 / out 155,539). Hermes also
deployed its own build: asked for a port proxy, it instead symlinked into the static vhost following
the existing bench convention — zero nginx changes, MIME types verified.
The prompt (verbatim, all four runs)
Identical text verified byte-for-byte
against the OpenCrabs session records and the Claude Code transcript — including the Visual Options
menu requirement and the tools-immediately sentence, typos and all. The Hermes and OpenCrabs run 2
copies are the same task with small textual deltas: each spells its own workdir path out in the
first line (run 1 and Claude Code got theirs from cwd) plus terminal-editor whitespace padding —
same spec, same typos, not byte-identical.
Build OC FPS: a playable Three.js FPS demo. Shell index.html + js/ modules (one concern each, <300 lines), importmap CDN, no build step, pointer lock, WASD/sprint/crouch/ADS. Assets procedural or CDN only.
Realism bar: ACES tone mapping, PBR materials, shadow maps, bloom/vignette/grain. Movement has accel curves, head bob, landing dip, weapon sway, ADS FOV shift. Three weapons (hitscan, launcher, scanner) with recoil bloom shown on the crosshair, tracers, decals, hitmarkers, timed reloads. Enemies run LOS/hearing state machines (patrol→alert→engage→cover) with hit reactions. Positional WebAudio, no files. HUD: spread-tracking crosshair, damage vignette, ammo, compass, kill feed.
Six escalating chapters, OC-1..OC-6: each gets its own palette/fog/lighting, one objective, one dialogue line, one journal unlock.
60 fps, pooled particles, zero per-frame allocations. F3 overlay: fps, draw calls, enemy states.
60FPS. Visual Options. Deliver: README, one-page design doc, the game, beat-by-beat playthrough script for chapter 3. Plan, then execute. Never stall — decide and ship.
Remember that your thinking/reasoning you cant execute tool calls, first execute tool calls ti know your environment, then plan and execute. You must call tools immediatelly
The lanes and the model board
This page pins the brain to
Opus 5.
The same harness experiment on
GLM-5.3-Flash lives at
/harness/glm-5.3-flash/ — a lane never ranks two models against
each other, because then the podium measures the brain instead of the loop. Model-vs-model standings
(GLM-5.3, Qwen3.8-Max, Opus 5, Fable 5.1, Kimi K3, DeepSeek V4 Flash) live on the
main bench board.