Skip to content
Hozu
Menu
All trials

Trial 0016 — 0.5 with the trial 0015 fixes, four runs per step, and a second agent

Question: trial 0015 measured the 0.5 build at 1.46× Nuxt, but its change step was noise (one outlier at 177 k), and it found four sources of friction. After fixing them:

what does 0.5 cost with four runs per step, and does a second agent (Codex) get it right?

Setup

Results

Correctness: every run passes everything.

h1 h2 h3 h4 Codex
Build (15) 15 15 15 15 15
Change: new (6) 6 6 6 6 6
Change: regression (15) 15 15 15 15 15

Cost (Claude, weighted tokens):

h1 h2 h3 h4 Mean vs Nuxt Trial 0014 Trial 0015
Build 87.5 k 103.0 k 101.0 k 122.1 k 103.4 k 1.38× 1.80× 1.46×
Change 74.5 k 101.4 k 82.7 k 96.7 k 88.8 k 1.45× 1.81× 2.36×
Build + change 192.2 k 1.41× 1.80× 1.86×
Calls, build / change 12 / 11 20 / 13 13 / 10 19 / 12 16.0 / 11.5 6 / 8.5 18 / 12 16 / 17

Codex:

Anatomy (bench/trial/anatomy.mjs, means, trial 0014 → 0016):

Total Calls Output × 5 Docs, with carry Docs calls Read app Verify calls
Build 134.7 → 103.4 k 17.0 → 12.5 36.5 → 28.2 k 29.9 → 28.3 k 5.5 → 3.3 14.5 → 8.4 k 3.0 → 3.3
Change 110.6 → 88.8 k 12.0 → 11.0 36.3 → 28.4 k 32.5 → 22.6 k 4.0 → 3.0 4.8 → 4.2 k 3.5 → 2.5

Reading the result

Conclusion