Skip to content
Hozu
Menu
All trials

Trial 0014 — The notes app on 0.5 (ADR 0037 D7)

Question: 0.5 generates less code (contracts only for decisions, busy states by rule) and catches more mistakes. Does that lower the agent's cost on the notes task, and does correctness hold?

Targets stated in ADR 0037 before the run:

Setup

Results

Correctness:

Hozu run 1 Hozu run 2
Build (15 checks) 15 15
Change: new behaviour (6) 6 6
Change: regression (15) 15 15

Cost (weighted tokens = fresh + cached × 0.1 + output × 5):

Run 1 Run 2 Mean vs Nuxt Trial 0013 (0.3) Trial 0012 (0.3, no auth scaffold)
Build 106.0 k 163.3 k 134.7 k 1.80× 124.0 k (1.66×) 209.1 k (2.79×)
Build turns 15 21 18.5 20
Change 109.1 k 112.1 k 110.6 k 1.81× — 125.7 k (2.06×)
Cost (USD, build + change) 0.98 1.19 1.08 1.45

Generated code: the tasks feature (--with detail,toggle,filter,remove) is 10.5 KB, down from 14.0 KB. The ≤ 8 KB target is not met.

Where the tokens went

Reading the result

Conclusion