Skip to content
Hozu
Menu
All trials

Trial 0017 — a widget-heavy app (Leaflet, Chart.js, GSAP, Three.js)

Question: trials 0012–0016 used the notes task: forms, sessions, lists. Does Hozu hold up on an app whose interesting parts are third-party DOM libraries (a map, a chart, animations, a WebGL globe), where every library goes through ui.widget / ui.use / implement?

Setup

Acceptance fixes (before scoring)

All six apps failed the same checks, which pointed at the checks rather than at the apps. The three checks were rewritten to test what the spec says; the reference app still passes everything.

Check Was Now
W7 fade-in the opacity of the region and its ancestors the opacity from the region's <h2> upwards, so a fade on an inner wrapper counts
W9 no reload the URL is unchanged a window marker survives, so keeping ?q= in the URL with replaceState is allowed
C5 tour start the first sample is the first station the first selected station is the first station; the spec does not say the first step is immediate

Results

Correctness: every run passes everything after the fixes above.

h1 h2 Codex n1 n2
Build (15) 15 15 15 15 15
Change: new (8) 8 8 8 8 8
Change: regression (15) 15 15 15 15 15

Cost (Claude, weighted tokens):

h1 h2 Mean n1 n2 Mean Hozu / Nuxt
Build 373.1 k 317.8 k 345.4 k 125.5 k 135.1 k 130.3 k 2.65×
Change 155.8 k 236.7 k 196.2 k 49.0 k 63.7 k 56.3 k 3.48×
Calls, build / change 34 / 16 26 / 22 11 / 7 13 / 8

Codex (Hozu):

Where the cost went

bench/trial/anatomy.mjs shows three parts that the notes task did not have.

1. Code volume.

2. Verifying in a browser.

3. Three defects found by the runs:

Docs reading stayed in the range of trial 0016 (22–118 k carried per run), but it now comes on top of the three parts above.

Conclusion