OC Harness Bench — FPS ⚔️

same brain (Opus 5) · same prompt · same box — the harness was the variable · ranked by tokens to complete (fewest wins) · GLM-5.3-Flash lane → ← model board
Leanest context — rank key
OpenCrabs TUI
10.43M — 5.3x leaner than the hungriest
Cheapest run
OpenCrabs TUI
$8.00 — 5.8x cheaper than the priciest
Fastest harness
OpenCrabs TUI
37m30s vs 1h36m — 2.57x across 4 harnesses
Deliverable parity
4 / 4 playable
all load first try · full v2 spec · Visual Options menus

Head to head

#HarnessModelWall (active)Raw spanTokens · rank keyCostStepsTool callsFiles / SizeBuild
🥇 1 OpenCrabs TUIharness twin · OpenCrabs orchestration Opus 5 37m30s37m31s 10.43M $8.00 9657 calls 52 / 12.6 MB view build
🥈 2 OpenCrabs TUI · run 2harness twin · repeatability run Opus 5 41m01s41m02s 10.82M $10.79 8665 calls 28 / 162.5 KB view build
🥉 3 Hermes TUIthird harness · Hermes agent loop Opus 5 40m09s40m09s 12.51M $15.65 93107 calls 28 / 186.5 KB view build
4 Claude Code TUIharness twin · Claude Code agent loop Opus 5 1h36m1h42m 55.07M $46.59 249118 ok / 5 err 31 / 204.9 KB view build

Notes

The experiment

One variable: the harness. Same model (Opus 5), same machine (bench VPS, root), same prompt — byte-verbatim, card below. Three orchestration loops: OpenCrabs' TUI loop (step structure, tool batching, context handling) spawning the claude CLI as its provider; Claude Code's built-in agent loop — the OpenCrabs rows and the Claude Code row share that identical engine and transport; and the Hermes TUI driving the Anthropic API directly through its own gateway and session runtime. OpenCrabs ran twice (run 1 in the original bench workdir, run 2 in a fresh isolated workdir with the box to itself) for repeatability; every other harness ran once. All fired manually in-terminal, auto-approve on, no timeouts.

What the numbers say

Ranked by token efficiency — fewest tokens to complete the task wins. The clock rewards burning parallel compute; the token count is what the harness actually made the model read and write, so it's the price of the task. OpenCrabs TUI takes the board on both runs — run 1: 10.43M tokens, $8.00, 37m30s; run 2: 10.82M tokens, $10.79, 41m01s. Repeatability: ±3.8% tokens, ±9.8% wall across two independent one-shots — and both runs stay under Hermes. Hermes second — 12.51M tokens (1.2x ours), $15.65, 40m09s; the two OpenCrabs walls straddle it (37m30s under, 41m01s just over by 52s), average 39m16s still faster — but it loses the rank key that decides the podium either way. Claude Code's loop last — 55.07M tokens (5.3x ours), $46.59, 1h36m23s: 249 round trips re-reading a growing context, 97.6% of its tokens cache reads. The cheapest run cost 17% of the priciest.

Quality parity, not just speed

All four builds loaded first try and shipped the full v2 spec including the Visual Options menu; each self-verified headless and self-caught real bugs. OpenCrabs run 1: 3 bugs (unwinnable OC-3 reactor hitbox, dead cover state, backwards film grain), 80/80 smoke checks. OpenCrabs run 2: 23 modules / 2,914 lines, spec values measured not assumed — headshot multiplier 2.30x against the 2.3 spec, scanner piercing 3 stacked targets, launcher splash + 1.6 m/s knockback + 38 hp self-damage, recoil bloom driving the crosshair — and a full campaign sweep OC-1→OC-6 with 0 console errors and 0 failed requests. Claude Code run: 4 bugs (W sign error, frame-rate-dependent friction, muzzle computed in view space, doorway sealing), headless Chromium all-decks pass. Hermes run: five headless-Chrome suites — 0 bytes allocated per frame across 1200 sim steps, 0.044 ms sim cost under full combat load, 61–112 draw calls, ACES confirmed, full campaign OC-1→OVERLOCK with 0 errors — and honest caveats stamped on its own work (SwiftShader software raster; the 1-shot-per-frame auto-fire gate it chased down and correctly acquitted). Deliverables equal on all sides: game + README + design doc + chapter-3 playthrough script.

Provenance & asterisks

OpenCrabs row = native usage_ledger, cost rolled up by the harness at official Anthropic rates (Max sub — not billed per run). Claude Code row = harvested from the claude jsonl transcript, cost computed at official rates. The OpenCrabs run began with the operator's "continue" kickoff after a cancelled pre-switch leg; all figures cover the opus leg only (operator ruling: the run counts from the model switch). The OpenCrabs run also survived a parallel process writing into its workdir mid-run — it flagged the intrusion in its own final message; the dir was later migrated to opus-5-och/. Its build size includes ~12MB of headless-test artifacts (test/out PNGs, test-vendored three.js); the game proper is ~210KB, on par with its twin. OpenCrabs run 2 = also native usage_ledger (10,821,679 tokens / $10.79), a clean one-shot in a fresh isolated workdir with the box to itself — no parallel disturbance, and the cleanest deliverable tree of the lane (~157KB, no test artifacts left behind). Hermes row = harvested from its own sqlite state.db (sessions + session_model_usage): wall = prompt → final "Shipped and verified" message (22:16:58→22:57:07), no active/idle split recorded so wall = span; post-run proxy Q&A (23:07–23:11) excluded. Its 12,506,671 tokens are the harness's own counter and exclude a 1.41M background-review side task; cost computed at the same official rate card (in 186 / cache-write 970,520 / cache-read 11,380,426 / out 155,539). Hermes also deployed its own build: asked for a port proxy, it instead symlinked into the static vhost following the existing bench convention — zero nginx changes, MIME types verified.

The prompt (verbatim, all four runs)

Identical text verified byte-for-byte against the OpenCrabs session records and the Claude Code transcript — including the Visual Options menu requirement and the tools-immediately sentence, typos and all. The Hermes and OpenCrabs run 2 copies are the same task with small textual deltas: each spells its own workdir path out in the first line (run 1 and Claude Code got theirs from cwd) plus terminal-editor whitespace padding — same spec, same typos, not byte-identical.
Build OC FPS: a playable Three.js FPS demo. Shell index.html + js/ modules (one concern each, <300 lines), importmap CDN, no build step, pointer lock, WASD/sprint/crouch/ADS. Assets procedural or CDN only. Realism bar: ACES tone mapping, PBR materials, shadow maps, bloom/vignette/grain. Movement has accel curves, head bob, landing dip, weapon sway, ADS FOV shift. Three weapons (hitscan, launcher, scanner) with recoil bloom shown on the crosshair, tracers, decals, hitmarkers, timed reloads. Enemies run LOS/hearing state machines (patrol→alert→engage→cover) with hit reactions. Positional WebAudio, no files. HUD: spread-tracking crosshair, damage vignette, ammo, compass, kill feed. Six escalating chapters, OC-1..OC-6: each gets its own palette/fog/lighting, one objective, one dialogue line, one journal unlock. 60 fps, pooled particles, zero per-frame allocations. F3 overlay: fps, draw calls, enemy states. 60FPS. Visual Options. Deliver: README, one-page design doc, the game, beat-by-beat playthrough script for chapter 3. Plan, then execute. Never stall — decide and ship. Remember that your thinking/reasoning you cant execute tool calls, first execute tool calls ti know your environment, then plan and execute. You must call tools immediatelly

The lanes and the model board

This page pins the brain to Opus 5. The same harness experiment on GLM-5.3-Flash lives at /harness/glm-5.3-flash/ — a lane never ranks two models against each other, because then the podium measures the brain instead of the loop. Model-vs-model standings (GLM-5.3, Qwen3.8-Max, Opus 5, Fable 5.1, Kimi K3, DeepSeek V4 Flash) live on the main bench board.