OC Harness Bench — FPS ⚔️ GLM-5.3-Flash lane

same brain (GLM-5.3-Flash) · same prompt · same box — the harness was the variable · ranked by tokens to complete (fewest wins) · Opus 5 lane → ← model board
Leanest context — rank key
OhMyPi
17.54M — 1.3x leaner than the hungriest
Cheapest run
OhMyPi
$0.63 — 1.2x cheaper than the priciest
Fastest harness
OpenCrabs TUI
46m32s vs 1h06m — 1.44x across 2 harnesses
Deliverable parity
2 / 2 playable
prompt v2 · full spec · Visual Options menus

Head to head

#HarnessModelWall (active)Raw spanTokens · rank keyCostStepsTool callsFiles / SizeBuild
🥇 1 OhMyPiharness twin · first rival on this lane GLM-5.3-Flash 1h06m1h07m 17.54M $0.63 109160 ok / 18 err 22 / 124.8 KB view build
🥈 2 OpenCrabs TUIharness twin · lane opener GLM-5.3-Flash 46m32s48m07s 22.13M $0.75 222148 ok / 14 err 18 / 111.8 KB view build

Notes

The experiment

One variable: the harness. Same brain (GLM-5.3-Flash), same machine (bench VPS), same prompt v2 — the identical task the Opus 5 lane ran. This lane exists so harnesses can be compared against each other on a cheap, fast-moving model instead of being ranked against each other on an expensive one, which is what the model board is for. Two loops are on the board: the OpenCrabs TUI and the OhMyPi TUI (its own provider config, its own unix user, its own isolated workdir). Further harnesses land as their runs close. Ranking key: fewest tokens to complete the task wins — the clock rewards burning parallel compute, the token count is what the harness made the model read and write, so it is the price of the task.

What the numbers say

OhMyPi takes the lane: 17.54M tokens, 66m57s, $0.63. OpenCrabs TUI is second: 22.13M tokens, 46m32s, $0.75, 1.26x the tokens and 1.19x the cost, bought with a 20m25s faster wall. Same brain, same prompt, same box, so the difference is entirely the loop, and the counters name the mechanism: OpenCrabs ran 222 harness steps at ~99K gross context each, 0.73 tool calls per step; OhMyPi ran 109 steps at ~159K, 1.63 calls per step. Those are each harness's own step counter. On the OpenCrabs side the provider log gives the real call count too: 159 API calls averaging 138,497 gross prompt tokens. OhMyPi is not instrumented here at that level, so the per-trip size comparison is step-for-step, not call-for-call. OhMyPi batches independent tool calls, OpenCrabs serialises them, so the crab makes roughly twice the trips at about 62% of the context each, which nets out to 1.27x the gross input it had to pay for. That is the lane's finding, and it is a turn-economy finding, not a model one. Caching is not a differentiator here. The crab hit the provider prompt cache on 98.57% of its context volume against OhMyPi's 98.7%. An earlier version of this page claimed 49.5% for OpenCrabs and made the cost gap 6.3x; that was our own accounting double-counting the cached prefix, corrected on 2026-09-19 (see the provenance note). Both bills are dominated by cache reads (87% of the crab's $0.75, 81% of OhMyPi's $0.63), which is another way of saying the podium here is decided by how many times a harness re-reads its own context.

Quality parity, with an honest asymmetry

Both builds shipped the full v2 spec and both loaded: OpenCrabs 18 files / ~112KB (index.html + 15 js modules + 3 docs), OhMyPi 22 files / ~176KB (index.html + 17 js modules, all under 300 lines, + 3 docs, a metadata-only package.json and CDN three@0.160.0 instead of r170). OhMyPi's run recorded its own verification: five runtime bugs found and fixed (two cross-module player wiring errors, missing launcher ground impact, an F3 render-info accounting bug, a pooled-enemy spawn bug that silently dropped reused enemies from the AI list), headless passes plus screenshots across OC-1, OC-3 and OC-4, and it stamped its own limit honestly — real 60fps projected from design, not measured, because its headless pass ran software GL at ~1fps. The OpenCrabs row has no equivalent recorded verification sweep for this run, so none is claimed here. Both self-recovered their internal tool errors (14 and 18 respectively) without a single prompt about the game.

Provenance & asterisks

OpenCrabs row: native usage_ledger, which logged two rows for the session — 43,833,506 / $4.00799 at 16:46:04 (the build leg, published) and 470,832 / $0.04544 at 17:21:36 (a 73s post-completion continuation, excluded); the session total 44,304,338 / $4.05341 closes exactly across both. Wall = the assistant turn's own duration_secs (2792s); raw span = session created → completion. Two operator inputs: the prompt at 15:59:15 and a "go" 17 seconds later, zero steering after.
OhMyPi row: tokens and cost summed from its own session jsonl usage blocks (in 217,109 / out 174,521 of which 84,783 reasoning / cache read 17,148,416 / cache write 0); no compaction events in the transcript. Wall = prompt 01:01:30 → final message 02:08:27; the ~96s operator unblock sits inside that wall (65m21s excluding it). One operator input 13m07s in: the sudo password to unblock a chown after its own permission failed — environment plumbing, not task direction. Cost basis, corrected 2026-09-19: both rows sit on the same rate card ($0.15/M uncached input, $0.50/M output, $0.03/M cache read, list prices; the 50% launch promo on this model is applied to neither row). OhMyPi's $0.63428 reproduces to the fifth decimal from its own split on that card and is published as harvested. The OpenCrabs row is not: its raw ledger figures (43,833,506 tokens / $4.00799) were inflated by an OpenCrabs accounting defect. On OpenAI-compatible providers prompt_tokens already contains the cached prefix, but OpenCrabs stores it in a field documented as non-cached input and then adds the cache read on top again, pricing the prefix twice. The bench log carries a per-call usage line for every provider call, and the 159 calls of this run sum to 22,020,984 gross input + 21,705,152 cache reads + 107,370 output, reproducing $4.00798716 exactly. The same lines show the cached count sitting inside the prompt count, which is what makes the second addition a double-count. Netting the prefix out of the gross input leaves 315,832 genuinely uncached input, so the run is 22,128,354 tokens and $0.7522 at a 98.57% cache hit rate. Those are the figures on the board above; the raw ledger values are kept in the run record. The build, the wall clock, the turn count and the tool counts were never affected. An earlier version of this note flagged a second defect in the other direction, a branch that zeroes output_tokens when a stream ends without a usage chunk. The branch exists, but it never fired: across twelve days of retained logs, all 762 of its log lines sit microseconds behind a real usage line on the same call, so no output was ever lost. Withdrawn. OhMyPi's accounting is its own and carries neither the double-count nor that branch. Tracked as opencrabs#1636.

The prompt (verbatim)

Whitespace-normalised diff of the OpenCrabs run against the canonical v2 text: one insertion, its own workdir (" at /srv/bench/fps/glm-53-flash" in the first line). Same spec, same typos, otherwise identical. The OhMyPi copy carries the same class of delta (its workdir spelled out) plus terminal whitespace padding.
Build OC FPS: a playable Three.js FPS demo. Shell index.html + js/ modules (one concern each, <300 lines), importmap CDN, no build step, pointer lock, WASD/sprint/crouch/ADS. Assets procedural or CDN only. Realism bar: ACES tone mapping, PBR materials, shadow maps, bloom/vignette/grain. Movement has accel curves, head bob, landing dip, weapon sway, ADS FOV shift. Three weapons (hitscan, launcher, scanner) with recoil bloom shown on the crosshair, tracers, decals, hitmarkers, timed reloads. Enemies run LOS/hearing state machines (patrol→alert→engage→cover) with hit reactions. Positional WebAudio, no files. HUD: spread-tracking crosshair, damage vignette, ammo, compass, kill feed. Six escalating chapters, OC-1..OC-6: each gets its own palette/fog/lighting, one objective, one dialogue line, one journal unlock. 60 fps, pooled particles, zero per-frame allocations. F3 overlay: fps, draw calls, enemy states. 60FPS. Visual Options. Deliver: README, one-page design doc, the game, beat-by-beat playthrough script for chapter 3. Plan, then execute. Never stall — decide and ship. Remember that your thinking/reasoning you cant execute tool calls, first execute tool calls ti know your environment, then plan and execute. You must call tools immediatelly

The lanes

This board pins the brain to GLM-5.3-Flash. The same harness experiment on Opus 5 (four rows: OpenCrabs TUI ×2, Hermes TUI, Claude Code TUI) lives at /harness/, and model-vs-model standings live on the main bench board. A lane never ranks two models against each other — that would measure the brain instead of the loop.