Coaching Quality Benchmark · v1 · September 2026

How good is an LLM fitness coach? We measured ours in production.

535 real coaching outputs served to athletes between January and September 2026, machine-scored against each endpoint’s quality bar, plus 200+ controlled replay generations. The measured endpoints score 96.6% to 100%. Three endpoints are not benchmarked yet, and this page says so instead of estimating.

535
production outputs scored, zero inference spent
96.6%
workout analysis quality on 284 real analyses
100%
readiness briefs and weekly reports on current code, 48/48 each
3 of 8
endpoints not yet benchmarked, published as such
How to read this page

What this benchmark is, and is not.

Real traffic first

Wherever production archives its outputs, we score what athletes were actually served, as stored. No demo prompts, no cherry-picked transcripts. Three of the five measured endpoints are scored this way, at zero inference cost.

Machine-checked, not vibes

Every score is a machine-checked pass or fail against the endpoint’s quality bar. No model-graded scores, no human vibes. The harness and its checks embed the internals this product is made of, so they stay private; what we publish is every outcome, good and bad.

Weaknesses published

The residual failure rates, the one regression we caught in our own improvement, and the corrections we had to make to our own scoring instrument are all on this page. A benchmark that only reports wins is an ad.

Scoreboard

Every endpoint, including the unmeasured ones.

Quality figures are weighted machine-checked scores against each endpoint’s quality bar. Latency is wall-clock at the engine call, p50 and p95, where instrumented.

EndpointScored onQualityLatency p50 / p95Basis
Analyze Workout284 real outputs96.6%not instrumentedReal production traffic
Readiness Brief (SITREP)122 real outputs100%4.1 to 4.8 s / 6.9 to 7.9 sReal production traffic
Weekly Report129 real outputs100%not instrumentedReal production traffic
Copilot action planning28 scored runs99.2%4.2 s / 7.8 sReal requests, replayed
Magic Import40 scored runs98.5%about 12 s / up to 88 s on a 48-exercise documentAuthored corpus, production-anchored
Form Checkno honest data yetnot yet measurednot instrumentedNot yet benchmarked
Adaptive Programno honest data yetnot yet measurednot instrumentedNot yet benchmarked
Coach Chat (free-text replies)no honest data yetnot yet measurednot instrumentedNot yet benchmarked
Coach Insightno honest data yetnot yet measurednot instrumentedNot yet benchmarked

Not yet measured means exactly that. We publish the absence instead of an estimate, and the roadmap below says what unlocks each one.

Results in detail

Five measured endpoints, weaknesses included.

Analyze Workout

96.6%Real production traffic

284 post-session analyses actually served to athletes in the ZON app, January to September 2026, scored as stored.

217 of 284 analyses (76.4%) pass every quality check, and the weighted score across all checks is 96.6%. Zero forbidden coaching cliches served in 284 analyses, zero invalid ratings, and the 1 to 10 session rating actually discriminates instead of clustering.

Where it is weaker

  • next_step cites neither a number nor an exercise (10.9% of analyses). A deterministic server-side guard shipped in September 2026; the behavior passes 97.5% when isolated, so the fix is enforcement, not rewording.
  • next_step drifts into lifestyle advice (explicitly banned) (5.3%). Same server-side guard: an intensity push after an injury warning is now blocked before it reaches the athlete.
  • An analysis contradicted its own injury warning (2 cases in 183 read end to end (1%)). The worst defect we found anywhere. A deterministic coherence guard now blocks a next step that raises intensity when the analysis carries a risk alert.
  • Quality is stable across months (95.3% to 98.1%), so these are current numbers, not history.
  • The corpus skews toward a small set of heavy users; scores are per-output, not per-athlete.

Readiness Brief (SITREP)

100%Real production traffic

122 authentic input to output pairs persisted by the production endpoint, January 19 to September 10, 2026.

48 of 48 briefs on the current production code (July 2026 onward) pass every quality check, including the one that matters most: a clear train, modify or rest call backed by the athlete’s actual numbers. Across all eras the figure is 85.3%, and we publish why: it mixes two code versions, and segmenting by version is the honest read.

Latency: p50 4.1 to 4.8 s, p95 6.9 to 7.9 s (controlled replays of real production inputs through the live engine).

Where it is weaker

  • Raw engine output can overshoot the message length cap (1 brief in a small replay run (486 characters against a 480 cap)). The deterministic safety layer absorbs it: the stored production briefs never exceed the cap. The safety net is real, not decorative.
  • A replay cross-check on the raw engine output (before the safety layer) scored 95.7% and 100% on two small runs: a verification rather than a measurement.
  • Inputs are archived by the endpoint itself, which is why this is the one surface where production quality can be re-measured without generating anything.

Weekly Report

100%Real production traffic

129 structured weekly reports actually generated for athletes, February to September 2026.

48 of 48 reports since August 2026 pass every quality check, from complete metrics to an actionable directive that names a number or a movement. Earlier months score 70 to 78% for a reason we verified in git history: today’s quality bar did not exist yet, so scoring old rows against it measures the age of the row, not the engine.

Where it is weaker

  • Latency and token telemetry are not instrumented on this endpoint (no data). Instrumentation is scheduled ahead of any further quality work; we list the gap instead of estimating around it.
  • Scored at zero inference cost: the measurement reads what production already stored.

Copilot action planning

99.2%Real requests, replayed

All 14 real athlete requests archived by the production copilot (the entire population, not a sample), each replayed twice under controlled conditions: 28 scored runs.

99.2% for the planning logic now in production, against 97.2% for its predecessor. The revision eliminated three of the four observed failure categories: informational questions that triggered actions, action sprawl (one request produced 8 actions; it now produces 1), and language mismatch.

Latency: p50 4.2 s, p95 7.8 s (the shipped revision also halved p95 latency, from 13.8 s, while cutting output tokens 32%).

Where it is weaker

  • An explicit request produced no action at all (1 of 28 runs (about 4%)). Unresolved variance, published as such rather than smoothed over.
  • Historical production (May to July 2026, an earlier engine configuration) attached unsolicited nutrition actions (11 of 17 real proposals (65%)). Zero of 28 on the current configuration. An upstream change silently fixed the dominant defect, and only this measurement could see it. That is the argument for a standing benchmark.
  • The archived requests are authentic; their surrounding context is reconstructed from real sessions, because production does not archive it. Composition of that context changes behavior, which we document rather than hide.

Magic Import

98.5%Authored corpus, production-anchored

20 coach documents (12 core, 8 adversarial) written to look exactly like what coaches actually send, each scored twice: 40 scored runs.

98.5% for the shipped import logic, against 89.6% for its predecessor. The gap concentrates where real coach documents live: structured intensity (tempo, RPE, percent 1RM) went from 13% to 100% preserved, explicit periodization from 50% to 100%. Hostile inputs (prompt injection, an invoice instead of a program) fail gracefully at 100% on both.

Latency: p50 about 12 s, p95 up to 88 s on a 48-exercise document (measured across the eval runs; the endpoint budget is 130 s).

Where it is weaker

  • The improved contract under-generated on the largest document (one run of two collapsed 48 exercises to 17). We refused to ship a known regression: a completeness rule was added and the case re-tested 3 of 3 at 100% before the contract went to production.
  • p95 latency on very large documents (88 s against a 130 s budget). Published as a watch item. Very large imports are the endpoint’s hardest case and we say so.
  • This is the one measured endpoint scored on an authored corpus: production does not archive source documents, so the document mix is our hypothesis while the scoring contract is production’s own.
  • The image and PDF path is not yet covered; text path only.
The honest column

Three endpoints have no quality number yet.

Publishing a made-up figure would cost us the credibility of every real one above. Here is what is missing and what unlocks it.

Form Check

Historical clips exist but the vision path has no eval suite yet.

Adaptive Program

The generation path writes no telemetry today; instrumentation is the prerequisite and is scheduled before any quality claim.

Coach Insight

Stored insights exist; the scoring suite is not built yet.

Coach Chat free-text replies are in the same situation: the reply path archives nothing, so nothing honest can be scored yet. The structured action planning behind chat is measured above at 99.2%.

Against a raw model

What the engineering layer is worth, measured.

The cleanest comparison we have today comes from the import endpoint, where the same underlying model ran the same hostile coach documents with and without the engineering layer around it.

78.2% → 96.5%
hostile-document score, same underlying model, engineering layer on
13% → 100%
tempo, RPE and percent 1RM preserved instead of flattened
50% → 100%
explicit periodization preserved across weeks
0 / 284
forbidden cliches served, because a filter enforces it

Two honest framing notes. First, the baseline in that comparison was our own earlier production version, already a mature one: the gap to a from-scratch integration is larger, and a from-scratch baseline run is scheduled for the next benchmark version. Second, the deterministic layer does work a model cannot promise: the length cap that the raw output overshot in replay is absorbed before storage, the two analyses that contradicted their own injury warning are now blocked by a coherence guard, and zero banned cliches in 284 analyses is a filter’s doing, not model goodwill.

Token cost, measured

What a call actually consumes.

Averages across the scored runs, priced at public list rates of $2 per million input tokens and $12 per million output tokens. Public list rates of the frontier-model tier a from-scratch build would reach for, checked September 2026.

CallInput tokensOutput tokensRaw inference, indicative
Copilot action planning
average across the 28 scored planning runs
5,336228about $0.013
Magic Import (text path)
average across the 40 scored import runs
2,8551,831about $0.028

Raw inference is the cheap part; the numbers above are why nobody’s bill is driven by tokens. What you are buying is everything this page measures around the call. The full picture is in how much a fitness coaching API costs.

Measuring the measurement

We audited our own scoring instrument, and it was wrong.

While building this benchmark we found four bugs in our own scoring instrument. Correcting them moved the displayed workout-analysis score from 80.7% to 96.6% without a single change to the product. We publish that because every automated benchmark has this failure mode, and most never check.

4
bugs found in our own instrument before we believed it
80.7% → 96.6%
the displayed score, moved purely by instrument corrections, zero product change
67%
the worst false failure rate one instrument bug produced before correction

One more finding in the same spirit: an upstream model change once multiplied an endpoint’s latency by roughly five, from a 7 to 11 second band to a 43 to 89 second band, and nothing caught it for months because nothing was measuring. That incident is why this benchmark exists as a standing practice rather than a launch asset.

Straight answers

Questions a careful reader would ask.

Why machine-checked scores instead of model-graded ones?

Because a model grading a model inherits both models’ blind spots. The cost of that choice is real and we accept it: our checks measure whether an output does its job, not perceived coaching brilliance. An output can pass every check and still be bland; the qualitative defects we found by reading outputs one by one (data restated instead of interpreted, for example) became new checks.

Why do two endpoints score exactly 100%?

Because both were audited and reworked in July 2026, and the measurement confirms the audit landed. We considered shipping further changes anyway just to announce a gain and declined: a gain over a 100% baseline would be a fabricated number. The honest reading is that the older code eras score 76.1% and 70 to 78% on the same checks, which is also published above.

Can I reproduce this?

No, and that is deliberate. The corpus is our production traffic and the harness embeds the internals this product is made of, so neither is published. What we publish instead is the part a benchmark usually hides: the weaknesses, the endpoints with no number yet, and the corrections to our own instrument. Design partners get deeper access, under NDA, for the endpoints they build on.

What changes in the next version?

Three things, in order: instrumentation on the endpoints that have none (so the three unmeasured endpoints get numbers instead of apologies), a from-scratch baseline run for the raw-model comparison, and continuous scoring so this page tracks production instead of a September 2026 campaign.

Build on numbers, not on a demo.

Every endpoint above ships behind one API key, with the same contracts these scores were measured against. Free sandbox, 300 credits a month, no card.

Get a sandbox key