The experiment
One variable: the harness. Same brain
(GLM-5.3-Flash), same machine (bench VPS), same prompt v2 — the identical task the Opus 5 lane
ran. This lane exists so harnesses can be compared against each other on a cheap, fast-moving
model instead of being ranked against each other on an expensive one, which is what the model
board is for. Two loops are on the board: the OpenCrabs TUI and the OhMyPi TUI (its own provider
config, its own unix user, its own isolated workdir). Further harnesses land as their runs
close. Ranking key: fewest tokens to complete the task wins — the clock rewards burning
parallel compute, the token count is what the harness made the model read and write, so it is
the price of the task.
What the numbers say
OhMyPi takes the lane: 17.54M tokens,
66m57s, $0.63. OpenCrabs TUI is second: 22.13M tokens, 46m32s, $0.75, 1.26x the
tokens and 1.19x the cost, bought with a 20m25s faster wall. Same brain, same prompt, same box,
so the difference is entirely the loop, and the counters name the mechanism: OpenCrabs ran
222 harness steps at ~99K gross context each, 0.73 tool calls per step; OhMyPi ran
109 steps at ~159K, 1.63 calls per step. Those are each harness's own step counter. On the
OpenCrabs side the provider log gives the real call count too: 159 API calls averaging 138,497
gross prompt tokens. OhMyPi is not instrumented here at that level, so the per-trip size
comparison is step-for-step, not call-for-call. OhMyPi batches independent tool calls, OpenCrabs
serialises them, so the crab makes roughly twice the trips at about 62% of the context each,
which nets out to 1.27x the gross input it had to pay for. That is the lane's finding, and it is
a turn-economy finding, not a model one. Caching is not a differentiator here. The crab
hit the provider prompt cache on 98.57% of its context volume against OhMyPi's 98.7%. An earlier
version of this page claimed 49.5% for OpenCrabs and made the cost gap 6.3x; that was our own
accounting double-counting the cached prefix, corrected on 2026-09-19 (see the provenance note).
Both bills are dominated by cache reads (87% of the crab's $0.75, 81% of OhMyPi's $0.63), which
is another way of saying the podium here is decided by how many times a harness re-reads its own
context.
Quality parity, with an honest asymmetry
Both builds shipped the full v2
spec and both loaded: OpenCrabs 18 files / ~112KB (index.html + 15 js modules + 3 docs), OhMyPi
22 files / ~176KB (index.html + 17 js modules, all under 300 lines, + 3 docs, a metadata-only
package.json and CDN three@0.160.0 instead of r170). OhMyPi's run recorded its own verification:
five runtime bugs found and fixed (two cross-module player wiring errors, missing launcher ground
impact, an F3 render-info accounting bug, a pooled-enemy spawn bug that silently dropped reused
enemies from the AI list), headless passes plus screenshots across OC-1, OC-3 and OC-4, and it
stamped its own limit honestly — real 60fps projected from design, not measured, because its
headless pass ran software GL at ~1fps. The OpenCrabs row has no equivalent recorded
verification sweep for this run, so none is claimed here. Both self-recovered their internal tool
errors (14 and 18 respectively) without a single prompt about the game.
Provenance & asterisks
OpenCrabs row: native usage_ledger,
which logged two rows for the session — 43,833,506 / $4.00799 at 16:46:04 (the build leg,
published) and 470,832 / $0.04544 at 17:21:36 (a 73s post-completion continuation,
excluded); the session total 44,304,338 / $4.05341 closes exactly across both. Wall = the
assistant turn's own duration_secs (2792s); raw span = session created → completion. Two operator
inputs: the prompt at 15:59:15 and a "go" 17 seconds later, zero steering after.
OhMyPi row: tokens and cost summed from its own session jsonl usage blocks (in 217,109 /
out 174,521 of which 84,783 reasoning / cache read 17,148,416 / cache write 0); no compaction
events in the transcript. Wall = prompt 01:01:30 → final message 02:08:27; the ~96s operator
unblock sits inside that wall (65m21s excluding it). One operator input 13m07s in: the sudo
password to unblock a chown after its own permission failed — environment plumbing, not task
direction.
Cost basis, corrected 2026-09-19: both rows sit on the
same rate card ($0.15/M uncached input, $0.50/M output, $0.03/M cache read, list prices;
the 50% launch promo on this model is applied to neither row). OhMyPi's $0.63428 reproduces to
the fifth decimal from its own split on that card and is published as harvested. The OpenCrabs
row is
not: its raw ledger figures (43,833,506 tokens / $4.00799) were inflated by an
OpenCrabs accounting defect. On OpenAI-compatible providers
prompt_tokens already
contains the cached prefix, but OpenCrabs stores it in a field documented as non-cached input and
then adds the cache read on top again, pricing the prefix twice. The bench log carries a
per-call usage line for every provider call, and the 159 calls of this run sum to
22,020,984
gross input + 21,705,152 cache reads + 107,370 output, reproducing $4.00798716 exactly. The
same lines show the cached count sitting inside the prompt count, which is what makes the second
addition a double-count. Netting the prefix out of the gross input leaves
315,832 genuinely uncached input, so the run is
22,128,354 tokens and $0.7522 at a
98.57% cache hit rate. Those are the figures on the board above; the raw ledger values are
kept in the run record. The build, the wall clock, the turn count and the tool counts were never
affected. An earlier version of this note flagged a second defect in the other direction, a branch that
zeroes
output_tokens when a stream ends without a usage chunk. The branch exists, but
it never fired: across twelve days of retained logs, all 762 of its log lines sit microseconds
behind a real usage line on the same call, so no output was ever lost. Withdrawn. OhMyPi's accounting is its own and carries neither the double-count nor that branch. Tracked as
opencrabs#1636.
The prompt (verbatim)
Whitespace-normalised diff of the OpenCrabs run
against the canonical v2 text:
one insertion, its own workdir (" at
/srv/bench/fps/glm-53-flash" in the first line). Same spec, same typos, otherwise identical. The
OhMyPi copy carries the same class of delta (its workdir spelled out) plus terminal whitespace
padding.
Build OC FPS: a playable Three.js FPS demo. Shell index.html + js/ modules (one concern each, <300 lines), importmap CDN, no build step, pointer lock, WASD/sprint/crouch/ADS. Assets procedural or CDN only.
Realism bar: ACES tone mapping, PBR materials, shadow maps, bloom/vignette/grain. Movement has accel curves, head bob, landing dip, weapon sway, ADS FOV shift. Three weapons (hitscan, launcher, scanner) with recoil bloom shown on the crosshair, tracers, decals, hitmarkers, timed reloads. Enemies run LOS/hearing state machines (patrol→alert→engage→cover) with hit reactions. Positional WebAudio, no files. HUD: spread-tracking crosshair, damage vignette, ammo, compass, kill feed.
Six escalating chapters, OC-1..OC-6: each gets its own palette/fog/lighting, one objective, one dialogue line, one journal unlock.
60 fps, pooled particles, zero per-frame allocations. F3 overlay: fps, draw calls, enemy states.
60FPS. Visual Options. Deliver: README, one-page design doc, the game, beat-by-beat playthrough script for chapter 3. Plan, then execute. Never stall — decide and ship.
Remember that your thinking/reasoning you cant execute tool calls, first execute tool calls ti know your environment, then plan and execute. You must call tools immediatelly
The lanes
This board pins the brain to
GLM-5.3-Flash. The same
harness experiment on
Opus 5 (four rows: OpenCrabs TUI ×2, Hermes TUI, Claude Code TUI)
lives at
/harness/, and model-vs-model standings live on the
main bench board. A lane never ranks two models against each other — that would
measure the brain instead of the loop.