OC Model Bench — FPS 🎮

playable three.js first-person shooter · prompt v2, verbatim · no timeout · 5/7 models finished · harness lane · Opus 5 → harness lane · GLM-5.3-Flash →
Fastest finish
Fable 5.1
32m59s
Lowest cost
GLM-5.3-Flash
$0.75
Biggest build
Opus 5
12.6 MB
Tokens burned
82.06M
$86.90 total · 5 finishers

Leaderboard

#ModelWallTokensCostTurnsToolsFiles / SizeBuild
🥇 1 Fable 5.1OpenCrabs TUI · one-shot, zero interventions 32m59s8.00M $17.64 7665 ok / — err 31 / 5.6 MB view build
🥈 2 Opus 5OpenCrabs TUI · loaded first try 37m30s10.43M $8.00 9657 ok / — err 52 / 12.6 MB view build
🥉 3 GLM-5.3TUI · shooter v2 (build + 1 fix prompt) 38m11s13.98M $4.69 3161 ok / 15 err 107 / 1.8 MB view build
4 GLM-5.3-FlashTUI · shooter v2 (one-shot) 46m32s22.13M $0.75 222148 ok / 14 err 18 / 111.8 KB view build
5 Qwen3.8-Max (0902)TUI · shooter v2 (post-update trigger) 1h56m27.52M $55.82 11352 ok / 20 err 49 / 299.3 KB view build
· Kimi K3queued —0 — 0— — / — queued
· DeepSeek V4 Flashqueued —0 — 0— — / — queued

Notes

The task (prompt v2)

Playable Three.js first-person shooter: index.html + js/ modules (one concern each), importmap CDN, no build step. Three weapons (hitscan, launcher, scanner) with recoil bloom, tracers, decals, hitmarkers. Enemy LOS/hearing state machines. Six chapters OC-1..OC-6, each with own palette/fog/lighting. 60 fps, pooled particles, zero per-frame allocations, F3 debug overlay. Deliverables: README, design doc, chapter-3 playthrough. Plan, then execute; never stall.
Build OC FPS: a playable Three.js FPS demo. Shell index.html + js/ modules (one concern each, <300 lines), importmap CDN, no build step, pointer lock, WASD/sprint/crouch/ADS. Assets procedural or CDN only. Realism bar: ACES tone mapping, PBR materials, shadow maps, bloom/vignette/grain. Movement has accel curves, head bob, landing dip, weapon sway, ADS FOV shift. Three weapons (hitscan, launcher, scanner) with recoil bloom shown on the crosshair, tracers, decals, hitmarkers, timed reloads. Enemies run LOS/hearing state machines (patrol→alert→engage→cover) with hit reactions. Positional WebAudio, no files. HUD: spread-tracking crosshair, damage vignette, ammo, compass, kill feed. Six escalating chapters, OC-1..OC-6: each gets its own palette/fog/lighting, one objective, one dialogue line, one journal unlock. 60 fps, pooled particles, zero per-frame allocations. F3 overlay: fps, draw calls, enemy states. 60FPS. Visual Options. Deliver: README, one-page design doc, the game, beat-by-beat playthrough script for chapter 3. Plan, then execute. Never stall — decide and ship. Remember that your thinking/reasoning you cant execute tool calls, first execute tool calls ti know your environment, then plan and execute. You must call tools immediatelly

Methodology

All numbers harvested from the session ledger on the bench VPS (tokens, cost, turns, tool calls). Wall time = prompt trigger to last write (pre-run stall era excluded); one-shot auto-approve runs; human review gaps between build and follow-up fixes are excluded from wall time. No timeout limit on this lane — runs may take hours. Builds are browser-verified in-session via headless Chrome, then served from /fps/<model>/ directly off the workdirs. Prompt verified byte-identical across all five finishers. Figures are as harvested except the GLM-5.3-Flash and GLM-5.3 rows, both netted of the cached prefix (see their notes); each run record keeps the raw ledger values alongside the corrected ones, and the netting is pinned by per-call provider usage lines in the bench log, not inferred from the totals. Qwen3.8-Max reported no cached-prompt split, and the two claude-cli rows report input natively non-cached, so all three are unaffected. Zero npm, zero JS: this page is static HTML.

Qwen3.8-Max: first finisher

Plan-first execution: a 9-task plan scaffold-to-deliverables, carried to completion with in-session browser verification. Qwen's signature pattern held — few giant turns instead of many small ones (its armory regen: $10.14, 4 turns). Full numbers in the leaderboard; the build is playable at the link. Superseded: the earlier v1 open-world prompt run ($13.78) is excluded from this board, preserved on disk under runs-excluded.

GLM-5.3: fastest, with one save

Delivered in 30m14s of build + a 7m57s fix after review caught a first build that did not load: a CDN dependency plus an rng is not a function TypeError. One follow-up prompt fixed both and vendored three.js locally (107 files now). Cost of the save: +2.13M tokens, +$0.71 (the raw ledger read +4.18M / +$3.58). The row includes the fix; the review gap is excluded. Figures corrected 2026-09-19, not re-run: this row read 27.34M tokens / $23.39 until then, both inflated by the same cached-prefix double-count that hit GLM-5.3-Flash. The 145 billed API calls in the bench log sum to 13.86M gross input + 13.36M cache reads + 115K output, reproducing the two ledger rows to the cent; netting the prefix out leaves 504,711 truly uncached input, so the run is 13.98M tokens / $4.69 at a 96.36% cache hit rate. Four further calls between the build and the fix turn (295,754 raw tokens) never reached the ledger at all, so that is a floor. Qwen3.8-Max shipped a build that loaded first try with zero follow-ups.

GLM-5.3-Flash: figures corrected, not re-run

This row read 43.83M tokens / $4.01 until 2026-09-19. Both numbers were an OpenCrabs accounting defect, not the model and not the provider: on OpenAI-compatible providers prompt_tokens already contains the cached prefix, and OpenCrabs stored it in a field documented as non-cached input, then added the cache read on top again, pricing the prefix twice. The bench log carries a per-call usage line for every API call, and the 159 billed calls of this run sum to 22.02M gross input + 21.71M cache reads + 107K output, reproducing the ledger row to the cent. The same lines show the cached count sitting inside the prompt count, which is what makes the second addition a double-count. Netting the prefix out leaves 315,832 truly uncached input, so the run is 22.13M tokens / $0.75 at a 98.57% cache hit rate. Wall clock, turns, tool counts and the build itself are untouched. The same defect inflates any row here whose provider reports a cached-prompt split: the GLM-5.3 row is corrected the same way, while Qwen3.8-Max reported no cache split and the Opus 5 and Fable 5.1 rows come through Anthropic, which reports input natively non-cached, so those three are unaffected. Tracked as opencrabs#1636.

Opus 5: OpenCrabs run, loaded first try

37m30s wall, 10.43M tokens, $8.00 native-ledger cost. Self-caught 3 bugs in-session (unwinnable OC-3, dead cover state, backwards film grain) before shipping, 80/80 headless smoke checks green. Zero fix prompts; figures cover the opus leg — the cancelled pre-switch leg is out of scope (operator ruling: the run counts from the model switch). Built while a parallel process was rewriting its workdir — flagged the intrusion in its own final message and still delivered. The 12.6MB dir includes ~8MB test PNGs + test-vendored three.js; the game proper is ~210KB. Same model through Claude Code TUI directly: 1h36m23s, 55.07M tokens.

Roster

Six models staged for the FPS lane; the prompt card above is verbatim what every finisher received (Visual Options menu requirement included — verified against session records and the Claude Code transcript). Done: Qwen3.8-Max (0902), GLM-5.3, GLM-5.3-Flash, Opus 5, Fable 5.1. Queued: Kimi K3, DeepSeek V4 Flash. Queued: Kimi K3, DeepSeek V4 Flash. The Opus 5 row is the OpenCrabs TUI run — same provenance as every other row on this board, native ledger numbers. Its twin, the same model run directly through Claude Code TUI's own agent loop, lives on the harness bench (1h36m23s, 55.07M tokens, $46.59 est) — OpenCrabs orchestration finished 2.57x faster on 5.3x fewer tokens. Dirs staged under /srv/bench/fps/; each finished run lands here via the harvest script — no manual HTML edits.