Trial 0014 — The notes app on 0.5 (ADR 0037 D7)
Question: 0.5 generates less code (contracts only for decisions, busy states by rule) and catches more mistakes. Does that lower the agent's cost on the notes task, and does correctness hold?
Targets stated in ADR 0037 before the run:
- build ≤ 1.3× Nuxt;
- a generated tasks feature ≤ 8 KB;
- correctness unchanged.
Setup
- As in trials 0012 and 0013:
- the same spec, change request and prompts;
- the same model (
claude-opus-5-5), hidden acceptance (bench/trial/notes/accept.mjs) andclaude -pin the app directory.
- Hozu: two runs on the packed 0.5.0, from
create-hozu --agent claude. - Nuxt was not re-run: its setup did not change. It is compared with trial 0012's runs: build mean 74.8 k, change mean 61.1 k.
- Codex (the second agent):
- a third app was created with
--agent agentsand run withcodex exec -s workspace-write; - it stopped after reading the skill: "You've hit your usage limit … try again at 3:55 AM";
- the run is void, and no Codex result is reported.
- a third app was created with
Results
Correctness:
| Hozu run 1 | Hozu run 2 | |
|---|---|---|
| Build (15 checks) | 15 | 15 |
| Change: new behaviour (6) | 6 | 6 |
| Change: regression (15) | 15 | 15 |
Cost (weighted tokens = fresh + cached × 0.1 + output × 5):
| Run 1 | Run 2 | Mean | vs Nuxt | Trial 0013 (0.3) | Trial 0012 (0.3, no auth scaffold) | |
|---|---|---|---|---|---|---|
| Build | 106.0 k | 163.3 k | 134.7 k | 1.80× | 124.0 k (1.66×) | 209.1 k (2.79×) |
| Build turns | 15 | 21 | 18.5 | 20 | ||
| Change | 109.1 k | 112.1 k | 110.6 k | 1.81× | — | 125.7 k (2.06×) |
| Cost (USD, build + change) | 0.98 | 1.19 | 1.08 | 1.45 |
Generated code: the tasks feature (--with detail,toggle,filter,remove) is 10.5 KB, down from 14.0 KB. The
≤ 8 KB target is not met.
Where the tokens went
- Build:
- both runs read the spec, then the scaffold's files in full (10.9 k and 17.0 k characters) despite the guide saying not to;
- run 2 also read the skill and
patterns.md; - both added a duplicate check and wrote a contract for it. That is a decision, so the contract is correct.
Both got its
effectswrong once (HZ015) and fixed it in one turn; - both started a server and used
curlat the end.
- Change:
- each run read
SKILL.md(about 10 k), the app's files (10–20 k), andpatterns.md/reference.md(5–16 k) before one edit; - verification used
hozu post --next.
- each run read
Reading the result
Correctness held: 72/72 checks across both runs, like trials 0012 and 0013.
Cost did not measurably fall.
- The build mean is 10 k above trial 0013, but the two runs differ by 57 k. As in trial 0011, one run's detour (printing files, a server check) is larger than the effect being measured.
- The change is 12 % below trial 0012's Hozu mean, which is within the same noise.
0.5 lowered what agents write, not what they read. The cost is reading:
- the guide (10–17 k characters);
- the app's own code (10–20 k);
- sometimes the patterns (5–16 k).
A Nuxt agent reads about 3 k before writing, because it already knows the framework. Shorter generated code lowers the second part a little. Nothing in 0.5 lowers the first.
The rule "do not print the generated files" is not followed. Agents read the code they are about to edit, whatever the guide says (trial 0011 found the same).
Conclusion
- Measured on 0.5: correctness 72/72, build 1.80× Nuxt, change 1.81× Nuxt.
- Target missed: the cost target (≤ 1.3×) is missed. The difference from 0.3 is within the noise of two runs.
- What remains: the cost that remains is the fixed reading cost of a framework the model does not know. Removing written code does not reach it.
- Next directions: the only levers left are to make the guide itself shorter and closer to what models already know, or to accept roughly 1.7–1.8× as the price of the checks that kept every run correct.