284 post-session analyses actually served to athletes in the ZON app, January to September 2026, scored as stored.
217 of 284 analyses (76.4%) pass every quality check, and the weighted score across all checks is 96.6%. Zero forbidden coaching cliches served in 284 analyses, zero invalid ratings, and the 1 to 10 session rating actually discriminates instead of clustering.
Where it is weaker
- next_step cites neither a number nor an exercise (10.9% of analyses). A deterministic server-side guard shipped in September 2026; the behavior passes 97.5% when isolated, so the fix is enforcement, not rewording.
- next_step drifts into lifestyle advice (explicitly banned) (5.3%). Same server-side guard: an intensity push after an injury warning is now blocked before it reaches the athlete.
- An analysis contradicted its own injury warning (2 cases in 183 read end to end (1%)). The worst defect we found anywhere. A deterministic coherence guard now blocks a next step that raises intensity when the analysis carries a risk alert.
- Quality is stable across months (95.3% to 98.1%), so these are current numbers, not history.
- The corpus skews toward a small set of heavy users; scores are per-output, not per-athlete.
122 authentic input to output pairs persisted by the production endpoint, January 19 to September 10, 2026.
48 of 48 briefs on the current production code (July 2026 onward) pass every quality check, including the one that matters most: a clear train, modify or rest call backed by the athlete’s actual numbers. Across all eras the figure is 85.3%, and we publish why: it mixes two code versions, and segmenting by version is the honest read.
Latency: p50 4.1 to 4.8 s, p95 6.9 to 7.9 s (controlled replays of real production inputs through the live engine).
Where it is weaker
- Raw engine output can overshoot the message length cap (1 brief in a small replay run (486 characters against a 480 cap)). The deterministic safety layer absorbs it: the stored production briefs never exceed the cap. The safety net is real, not decorative.
- A replay cross-check on the raw engine output (before the safety layer) scored 95.7% and 100% on two small runs: a verification rather than a measurement.
- Inputs are archived by the endpoint itself, which is why this is the one surface where production quality can be re-measured without generating anything.
129 structured weekly reports actually generated for athletes, February to September 2026.
48 of 48 reports since August 2026 pass every quality check, from complete metrics to an actionable directive that names a number or a movement. Earlier months score 70 to 78% for a reason we verified in git history: today’s quality bar did not exist yet, so scoring old rows against it measures the age of the row, not the engine.
Where it is weaker
- Latency and token telemetry are not instrumented on this endpoint (no data). Instrumentation is scheduled ahead of any further quality work; we list the gap instead of estimating around it.
- Scored at zero inference cost: the measurement reads what production already stored.
All 14 real athlete requests archived by the production copilot (the entire population, not a sample), each replayed twice under controlled conditions: 28 scored runs.
99.2% for the planning logic now in production, against 97.2% for its predecessor. The revision eliminated three of the four observed failure categories: informational questions that triggered actions, action sprawl (one request produced 8 actions; it now produces 1), and language mismatch.
Latency: p50 4.2 s, p95 7.8 s (the shipped revision also halved p95 latency, from 13.8 s, while cutting output tokens 32%).
Where it is weaker
- An explicit request produced no action at all (1 of 28 runs (about 4%)). Unresolved variance, published as such rather than smoothed over.
- Historical production (May to July 2026, an earlier engine configuration) attached unsolicited nutrition actions (11 of 17 real proposals (65%)). Zero of 28 on the current configuration. An upstream change silently fixed the dominant defect, and only this measurement could see it. That is the argument for a standing benchmark.
- The archived requests are authentic; their surrounding context is reconstructed from real sessions, because production does not archive it. Composition of that context changes behavior, which we document rather than hide.
20 coach documents (12 core, 8 adversarial) written to look exactly like what coaches actually send, each scored twice: 40 scored runs.
98.5% for the shipped import logic, against 89.6% for its predecessor. The gap concentrates where real coach documents live: structured intensity (tempo, RPE, percent 1RM) went from 13% to 100% preserved, explicit periodization from 50% to 100%. Hostile inputs (prompt injection, an invoice instead of a program) fail gracefully at 100% on both.
Latency: p50 about 12 s, p95 up to 88 s on a 48-exercise document (measured across the eval runs; the endpoint budget is 130 s).
Where it is weaker
- The improved contract under-generated on the largest document (one run of two collapsed 48 exercises to 17). We refused to ship a known regression: a completeness rule was added and the case re-tested 3 of 3 at 100% before the contract went to production.
- p95 latency on very large documents (88 s against a 130 s budget). Published as a watch item. Very large imports are the endpoint’s hardest case and we say so.
- This is the one measured endpoint scored on an authored corpus: production does not archive source documents, so the document mix is our hypothesis while the scoring contract is production’s own.
- The image and PDF path is not yet covered; text path only.