Comparison · Build vs buy

Should you build your coaching layer on a frontier model yourself?

The honest answer

If LLM infrastructure is your core competency, build. You will get exactly the behavior you design, you keep every prompt, and nothing on this page argues otherwise.

For everyone else, the trap is that the model call is the easy 10%. The other 90% is output contracts, deterministic validators, a scoring harness, and standing measurement across model migrations. We did not estimate that 90%: we measured it on production traffic, and the numbers below come from that work.

Decision table

Build it yourself when, buy the layer when.

Build on the model directly ifUse a coaching layer if
You have a team whose job is prompts, evals and model operations.Your team ships app features, and coaching is one of them.
Coaching logic IS your product’s moat and you must own every line of it.Coaching is table stakes your users expect, not your differentiation.
Your scale makes per-call economics dominate engineering cost.You are at the scale where four months of engineering costs more than years of API fees.
Your data may not transit any third party, contractually.A stateless per-call proxy that stores nothing satisfies your data posture.
The other 90%

What you actually rebuild, with the measured evidence.

Each line below exists because production traffic proved it necessary. The full detail, including where our own numbers are weak, is on the benchmark page.

Output contracts per endpoint

The import contract alone moved adversarial-document fidelity from 78.2% to 96.5% on the same model; structured intensity (tempo, RPE, percent 1RM) went from 13% preserved to 100%.

Deterministic validators behind the model

A coherence guard now blocks a next step that raises intensity after an injury warning (found twice in 183 real analyses). A length cap absorbs raw-model overshoot. Zero banned cliches in 284 production analyses is a filter, not luck.

A scoring harness you can trust

Four bugs in our own instrument produced false failure rates up to 67% before correction. Building the benchmark was as hard as building the prompts.

Standing measurement across model migrations

One migration silently fixed a 65% unsolicited-action defect; another silently multiplied latency by 5. Neither is visible without a suite that replays real traffic.

Domain post-processing

PR detection and session classification are computed, not asked of the model. The model does interpretation; arithmetic stays deterministic.

Cost reality

Raw inference is not where the money goes.

Measured per-call token averages from our scored runs, at public list rates of $2 in / $12 out per million tokens:

$0.013
Copilot action planning: 5,336 tokens in, 228 out
$0.028
Magic Import (text path): 2,855 tokens in, 1,831 out
4 months
a realistic first pass at contracts, validators and an eval harness for a handful of endpoints, before the first migration breaks something silently
×5
the latency regression one upstream model change caused in our stack while nothing was measuring. The suite exists because of this.

So a build decision priced on tokens is priced wrong in both directions: tokens are cheaper than most teams expect, and the engineering loop around them is more expensive. The full per-unit math, including our own prices next to every alternative, is in how much a fitness coaching API costs.

Fair questions

Pushback we agree with.

Frontier models keep improving. Does the layer become worthless?

Model upgrades genuinely absorbed one of our defects: a migration took unsolicited actions from 65% to zero, and we published that. What upgrades do not absorb: contracts (a better model still needs to know your schema), safety coherence rules, and the measurement that tells you what a migration changed. Better models make the layer thinner, and they make knowing-what-changed more valuable, not less.

Is a wrapper around someone else’s model defensible at all?

A naked wrapper is not, and the market is right to be cynical about them. What we sell is not model access: it is eight endpoint contracts extracted from a shipped training product, the validators behind them, and a production benchmark with our weaknesses printed on it. Judge whether that clears the bar on the numbers.

We already prototyped coaching with a system prompt. It looks fine.

Ours looked fine too. Then reading 183 real outputs found two analyses that told an athlete they were at elevated injury risk and, in the next sentence, to push intensity. One percent of outputs, invisible in any demo, and exactly the kind of thing that ends up in a screenshot. Fine-in-a-demo and safe-at-volume are different claims; only measurement separates them.

Skip the 90%, keep the proof.

Eight production-tested endpoints behind one key, with the benchmark to hold us to it.

Get a sandbox key