跳到主要內容
HOZU0.26.1
選單
All trials

Trial 0014 — The notes app on 0.5 (ADR 0037 D7)

Question: 0.5 generates less code (contracts only for decisions, busy states by rule) and catches more mistakes. Does that lower the agent's cost on the notes task, and does correctness hold?

Targets stated in ADR 0037 before the run:

  • build ≤ 1.3× Nuxt;
  • a generated tasks feature ≤ 8 KB;
  • correctness unchanged.

Setup

  • As in trials 0012 and 0013:
    • the same spec, change request and prompts;
    • the same model (claude-opus-5-5), hidden acceptance (bench/trial/notes/accept.mjs) and claude -p in the app directory.
  • Hozu: two runs on the packed 0.5.0, from create-hozu --agent claude.
  • Nuxt was not re-run: its setup did not change. It is compared with trial 0012's runs: build mean 74.8 k, change mean 61.1 k.
  • Codex (the second agent):
    • a third app was created with --agent agents and run with codex exec -s workspace-write;
    • it stopped after reading the skill: "You've hit your usage limit … try again at 3:55 AM";
    • the run is void, and no Codex result is reported.

Results

Correctness:

Hozu run 1 Hozu run 2
Build (15 checks) 15 15
Change: new behaviour (6) 6 6
Change: regression (15) 15 15

Cost (weighted tokens = fresh + cached × 0.1 + output × 5):

Run 1 Run 2 Mean vs Nuxt Trial 0013 (0.3) Trial 0012 (0.3, no auth scaffold)
Build 106.0 k 163.3 k 134.7 k 1.80× 124.0 k (1.66×) 209.1 k (2.79×)
Build turns 15 21 18.5 20
Change 109.1 k 112.1 k 110.6 k 1.81× — 125.7 k (2.06×)
Cost (USD, build + change) 0.98 1.19 1.08 1.45

Generated code: the tasks feature (--with detail,toggle,filter,remove) is 10.5 KB, down from 14.0 KB. The ≤ 8 KB target is not met.

Where the tokens went

  • Build:
    • both runs read the spec, then the scaffold's files in full (10.9 k and 17.0 k characters) despite the guide saying not to;
    • run 2 also read the skill and patterns.md;
    • both added a duplicate check and wrote a contract for it. That is a decision, so the contract is correct. Both got its effects wrong once (HZ015) and fixed it in one turn;
    • both started a server and used curl at the end.
  • Change:
    • each run read SKILL.md (about 10 k), the app's files (10–20 k), and patterns.md / reference.md (5–16 k) before one edit;
    • verification used hozu post --next.

Reading the result

  • Correctness held: 72/72 checks across both runs, like trials 0012 and 0013.

  • Cost did not measurably fall.

    • The build mean is 10 k above trial 0013, but the two runs differ by 57 k. As in trial 0011, one run's detour (printing files, a server check) is larger than the effect being measured.
    • The change is 12 % below trial 0012's Hozu mean, which is within the same noise.
  • 0.5 lowered what agents write, not what they read. The cost is reading:

    • the guide (10–17 k characters);
    • the app's own code (10–20 k);
    • sometimes the patterns (5–16 k).

    A Nuxt agent reads about 3 k before writing, because it already knows the framework. Shorter generated code lowers the second part a little. Nothing in 0.5 lowers the first.

  • The rule "do not print the generated files" is not followed. Agents read the code they are about to edit, whatever the guide says (trial 0011 found the same).

Conclusion

  • Measured on 0.5: correctness 72/72, build 1.80× Nuxt, change 1.81× Nuxt.
  • Target missed: the cost target (≤ 1.3×) is missed. The difference from 0.3 is within the noise of two runs.
  • What remains: the cost that remains is the fixed reading cost of a framework the model does not know. Removing written code does not reach it.
  • Next directions: the only levers left are to make the guide itself shorter and closer to what models already know, or to accept roughly 1.7–1.8× as the price of the checks that kept every run correct.