The task (prompt v2)
Playable
Three.js first-person shooter:
index.html + js/ modules (one concern each), importmap CDN, no build step. Three weapons
(hitscan, launcher, scanner) with recoil bloom, tracers, decals, hitmarkers. Enemy LOS/hearing
state machines. Six chapters OC-1..OC-6, each with own palette/fog/lighting. 60 fps, pooled
particles, zero per-frame allocations, F3 debug overlay. Deliverables: README, design doc,
chapter-3 playthrough. Plan, then execute; never stall.
Build OC FPS: a playable Three.js FPS demo. Shell index.html + js/ modules (one concern each, <300 lines), importmap CDN, no build step, pointer lock, WASD/sprint/crouch/ADS. Assets procedural or CDN only.
Realism bar: ACES tone mapping, PBR materials, shadow maps, bloom/vignette/grain. Movement has accel curves, head bob, landing dip, weapon sway, ADS FOV shift. Three weapons (hitscan, launcher, scanner) with recoil bloom shown on the crosshair, tracers, decals, hitmarkers, timed reloads. Enemies run LOS/hearing state machines (patrol→alert→engage→cover) with hit reactions. Positional WebAudio, no files. HUD: spread-tracking crosshair, damage vignette, ammo, compass, kill feed.
Six escalating chapters, OC-1..OC-6: each gets its own palette/fog/lighting, one objective, one dialogue line, one journal unlock.
60 fps, pooled particles, zero per-frame allocations. F3 overlay: fps, draw calls, enemy states.
60FPS. Visual Options. Deliver: README, one-page design doc, the game, beat-by-beat playthrough script for chapter 3. Plan, then execute. Never stall — decide and ship.
Remember that your thinking/reasoning you cant execute tool calls, first execute tool calls ti know your environment, then plan and execute. You must call tools immediatelly
Methodology
All numbers harvested from the session ledger on the bench VPS
(tokens, cost, turns, tool calls). Wall time = prompt trigger to last write (pre-run stall era excluded); one-shot auto-approve runs;
human review gaps between build and follow-up fixes are excluded from wall time. No timeout limit on this lane — runs may take hours. Builds are
browser-verified in-session via headless Chrome, then served from /fps/<model>/
directly off the workdirs. Prompt verified byte-identical across all five finishers. Figures are as harvested except the GLM-5.3-Flash and GLM-5.3 rows, both netted of the cached prefix (see their notes); each run record keeps the raw ledger values alongside the corrected ones, and the netting is pinned by per-call provider usage lines in the bench log, not inferred from the totals. Qwen3.8-Max reported no cached-prompt split, and the two claude-cli rows report input natively non-cached, so all three are unaffected. Zero npm, zero JS: this page is static HTML.
Qwen3.8-Max: first finisher
Plan-first execution: a 9-task plan
scaffold-to-deliverables, carried to completion with in-session browser verification. Qwen's
signature pattern held — few giant turns instead of many small ones (its armory regen:
$10.14, 4 turns). Full numbers in the leaderboard; the build is playable at the link.
Superseded: the earlier v1 open-world prompt run ($13.78) is excluded from this board,
preserved on disk under runs-excluded.
GLM-5.3: fastest, with one save
Delivered in 30m14s of build + a 7m57s fix
after review caught a first build that did not load: a CDN dependency plus an
rng is not a function TypeError. One follow-up prompt fixed both and vendored
three.js locally (107 files now). Cost of the save: +2.13M tokens, +$0.71 (the raw ledger read +4.18M / +$3.58).
The row includes the fix; the review gap is excluded. Figures corrected 2026-09-19, not re-run:
this row read 27.34M tokens / $23.39 until then, both inflated by the same cached-prefix
double-count that hit GLM-5.3-Flash. The 145 billed API calls in the bench log sum to 13.86M gross
input + 13.36M cache reads + 115K output, reproducing the two ledger rows to the cent; netting the
prefix out leaves 504,711 truly uncached input, so the run is 13.98M tokens / $4.69 at a 96.36%
cache hit rate. Four further calls between the build and the fix turn (295,754 raw tokens) never reached the ledger at all, so that is a floor. Qwen3.8-Max shipped a build that
loaded first try with zero follow-ups.
GLM-5.3-Flash: figures corrected, not re-run
This row read
43.83M tokens / $4.01 until 2026-09-19. Both numbers were an OpenCrabs accounting defect,
not the model and not the provider: on OpenAI-compatible providers
prompt_tokens
already contains the cached prefix, and OpenCrabs stored it in a field documented as non-cached
input, then added the cache read on top again, pricing the prefix twice. The bench log carries a per-call
usage line for every API call, and the 159 billed calls of this run sum to
22.02M gross input +
21.71M cache reads + 107K output, reproducing the ledger row to the cent. The same lines show the
cached count sitting inside the prompt count, which is what makes the second addition a double-count. Netting the prefix out leaves 315,832 truly
uncached input, so the run is
22.13M tokens / $0.75 at a 98.57% cache hit rate. Wall
clock, turns, tool counts and the build itself are untouched. The same defect inflates any row
here whose provider reports a cached-prompt split:
the GLM-5.3 row is corrected the same way,
while Qwen3.8-Max reported no cache split and the Opus 5 and Fable 5.1 rows come through
Anthropic, which reports input natively non-cached, so those three are unaffected. Tracked as
opencrabs#1636.
Opus 5: OpenCrabs run, loaded first try
37m30s wall, 10.43M tokens, $8.00
native-ledger cost. Self-caught 3 bugs in-session (unwinnable OC-3, dead cover state, backwards film
grain) before shipping, 80/80 headless smoke checks green. Zero fix prompts; figures cover the opus leg —
the cancelled pre-switch leg is out of scope (operator ruling: the run counts from the model switch).
Built while a parallel process was rewriting its workdir — flagged the intrusion in its own final message
and still delivered. The 12.6MB dir includes ~8MB test PNGs + test-vendored three.js; the game proper is
~210KB. Same model through Claude Code TUI directly:
1h36m23s, 55.07M
tokens.
Roster
Six models staged for the FPS lane; the prompt card above is
verbatim what every
finisher received (Visual Options menu requirement included — verified against session records and the
Claude Code transcript). Done: Qwen3.8-Max (0902), GLM-5.3, GLM-5.3-Flash, Opus 5, Fable 5.1. Queued: Kimi K3, DeepSeek V4 Flash.
Queued: Kimi K3, DeepSeek V4 Flash. The Opus 5 row is the
OpenCrabs TUI run — same provenance as
every other row on this board, native ledger numbers. Its twin, the same model run directly through
Claude Code TUI's own agent loop, lives on the
harness bench (1h36m23s, 55.07M
tokens, $46.59 est) — OpenCrabs orchestration finished 2.57x faster on 5.3x fewer tokens. Dirs staged under
/srv/bench/fps/;
each finished run lands here via the harvest script — no manual HTML edits.