What An AI Cycling Coach Actually Does (And What It Only Claims To Do)
There are three completely different technologies being sold to you under the same two letters, and only one of them will change your FTP this winter.
I’ve spent the last eighteen months feeding the same ride files into everything that calls itself an AI cycling coach app: TrainerRoad’s Adaptive Training, Xert, Athletica, Join, TrainAsONE, Humango, the various intervals.icu custom fitness models, plus a fair amount of ChatGPT and Claude with my full CSV export pasted in. The conclusion is not the one the marketing pages want. Rule-based plan adjustment is dressed up as machine learning and mostly isn’t useful. Genuine LLM reasoning is genuinely impressive and mostly can’t touch your training. The boring middle layer, statistical modelling of your own power data, is the only one that reliably makes you faster, and it’s the one nobody markets because “we fit a curve to your ride files” doesn’t raise a funding round.
Layer one: the if-then engine wearing a lab coat
Most of what gets called AI in this space is a decision tree somebody wrote by hand.
Here’s the pattern. You complete a workout. The platform asks how hard it felt, on a 1-to-10 scale or a four-button “easy / moderate / hard / all out”. You say hard. Next week’s equivalent session gets scaled by some fixed percentage. You say easy, it goes up. Miss two sessions, a ramp test gets inserted. That’s it. That’s the whole mechanism in a surprising number of products, and you can verify it yourself: do the same workout twice, give opposite feedback, and watch the next prescription move by exactly the same magnitude in both directions. Symmetrical, deterministic, reproducible. A learned model almost never behaves that cleanly.
TrainerRoad’s Adaptive Training is the honest high end of this layer, and to be fair to them they’ve never claimed otherwise in their technical write-ups: it’s Progression Levels, a per-zone difficulty score from 1.0 to 10.0, moved by survey responses and completion data. Fine work. Well-engineered. But when my Progression Level for Threshold sat at 5.3 and I told it a session was “hard”, the adjustment I got was a session at 5.1. Nothing in that loop knows my threshold actually drifted up 9 watts in the preceding three weeks, because nothing in that loop is looking at my power-duration curve. It’s looking at a button I pressed while sweaty and annoyed.
The tell for layer one is that the system asks you how you feel more than it asks your data. A useful heuristic: count the inputs. If subjective feedback is the primary signal and your files are only checked for did-you-finish, you’re paying a subscription for a spreadsheet with a friendly voice. That’s not worthless. Structure beats no structure, and a plan that nudges itself is better than a static twelve-week PDF you abandon in week five. Just don’t pay a premium for it or expect it to notice anything you didn’t tell it.
Layer two: statistical modelling, which is the part that works
This is where the actual gains live, and it’s the layer that has the weakest marketing and the strongest evidence.
The underlying idea is old. Monod and Scherrer published the critical power model in 1965. Morton and Billat extended it. Skiba built W’bal on top of it in 2012. What’s changed in the last five years is that fitting these models to messy real-world ride files, rather than lab tests, became computationally trivial, so your platform can refit your entire physiological profile every night against every ride you’ve ever uploaded.
Take a concrete case. My intervals.icu eFTP over one build, derived from the best efforts embedded in ordinary training rides rather than any test:
Date eFTP CP W' Source effort
2026-01-14 282 W 289 W 19.4 kJ 6min @ 341W, club chaingang
2026-02-02 288 W 294 W 18.1 kJ 12min @ 305W, Box Hill repeats
2026-02-19 296 W 301 W 17.2 kJ 20min @ 299W, TT sim on Zwift
2026-03-08 301 W 307 W 16.6 kJ 8min @ 328W, gravel race
2026-03-24 299 W 306 W 15.9 kJ 5min @ 345W, club chaingang
Read the W’ column, not the eFTP column. Threshold went up 19 watts across ten weeks, which is nice, but my anaerobic work capacity fell from 19.4 to 15.9 kJ, an 18% drop. That is the model telling me something I could not feel and did not report in any survey: my training had become so threshold-heavy that I’d traded away the exact capacity a UK crit or a punchy gravel race is decided by. No rule-based engine surfaced that, because I never pressed a button saying “I feel less snappy above 400 watts.” The curve knew. I added two sessions a week of 30/15s at 130% CP, and W’ came back to 18.3 kJ in a month with eFTP holding at 300.
That’s a training outcome changed by a model. Not by a chatbot, not by a survey.
Xert is the most aggressive commercial implementation of this. Its three-parameter model (Threshold Power, High Intensity Energy, Peak Power) updates from every ride via what it calls signature extraction, and its Maximal Power Available metric estimates in real time how much you have left. When Xert told me my MPA was down and my form was “tired” against my own feeling of freshness, it was right roughly four times in five, and the one time in five it was wrong was always after a ride with dodgy power data (dropouts on a cold Garmin dual-sided pedal, reading 40 watts low for a 90-second stretch, which the model ingests as a real effort and misfits). Athletica does something comparable with a per-zone impulse-response approach. intervals.icu gives you the raw machinery free and expects you to think.
The practical numbers matter here. A well-fitted CP model on a rider with reasonable data density predicts a genuine 20-minute maximal effort to within about 2 to 4%. Your estimated FTP from a ramp test carries error bars wider than that. So: your model is already more accurate than your testing, which means the argument for testing every six weeks has quietly collapsed, and if your platform still demands ramp tests it’s telling you it doesn’t have a model worth the name. If you want the side-by-side on which platforms actually fit models versus which ones fit vibes, that comparison is in AI Coaching Platforms, Tested.
Layer three: real LLM reasoning, and why it can’t coach you yet
Now the fun part, because this is where the technology is most genuinely remarkable and least useful for your training.
I gave Claude Opus and GPT-5 the same input: 400 rides of exported summary data, my full power curve, six weeks of HRV, and a target event (a hilly 25-mile TT, 8 weeks out, 1,100 feet of climbing, needing a 55-minute ride). The reasoning was good. Better than good. It correctly identified that my 5-minute power was disproportionately strong relative to my 40-minute power, inferred a fast-twitch bias, flagged that a hilly TT would punish my pacing discipline specifically because of that profile, and proposed over-under work at 95–105% CP with the overs placed on the climbs. That is coaching-quality analysis. A £100-an-hour human coach would give you the same read.
Then it wrote me a plan with Tuesday 4x8 at 105% FTP, Thursday 2x20 at 95%, Saturday endurance, and here is where it falls apart: it had no idea what happened on Tuesday. Nothing in that loop closes. The model has no persistent state, no automatic ingest of Wednesday morning’s HRV, no knowledge that I abandoned the third interval because of a headwind on the A road, and no ability to refit anything. Every conversation restarts from whatever I remember to paste in. Coaching is not an analysis problem, it’s a state-tracking problem across months, and LLMs are currently spectacular at the former and structurally bad at the latter.
Two more failure modes worth knowing. First, arithmetic on long series. Ask an LLM to compute your weekly TSS from raw ride data and it will confidently hand you a figure that’s 8% off, because it’s pattern-matching over numbers rather than summing them. It’ll get 4x8min at 105% of 300W right; it won’t reliably get a 42-day exponentially weighted average right. Second, sycophancy. Tell it you want to do three threshold sessions a week on 6 hours of training time and it will find a way to agree with you. Xert will just show you a falling signature and let you draw your own conclusion.
Where LLMs do earn their keep, and it’s a real use case: interpretation. Paste in a race file and ask why you cracked at 40km. Ask what the difference between your two best 20-minute efforts tells you about pacing. Ask for the physiological argument against a plan someone sold you. The prompt that’s worked best for me, by a distance:
Here is my power curve, my CP/W’ estimates for the last six months, and my last four race files with lap splits. Do not write me a training plan. Identify the single limiter most likely to decide my result at [event], argue against your own answer, then tell me which specific number in this data would change your mind.
Forcing the counter-argument kills most of the flattery. Asking what would change its mind tells you what to go measure.
What this means for where your money goes
Rank the three layers by watts-per-pound-sterling and the order is unambiguous. Pay for the model. A £10-a-month platform that refits your CP, W’ and per-duration curve nightly from real rides gives you information you cannot get any other way and cannot feel. Use a free or cheap LLM for interpretation on the occasions you have a specific question about a specific file. Treat rule-based plan adjustment as a scheduling convenience, priced accordingly, and stop reading its survey prompts as intelligence.
One caveat on the middle layer, because it isn’t magic: statistical models are only as good as their input density. If you ride three times a week and never go properly deep, your CP fit has nothing to anchor on and your W’ estimate will wander by 4 or 5 kJ on noise alone. The model needs maximal efforts across at least three durations (something near 3, 8 and 20 minutes) inside a rolling six weeks. Hit that and the numbers get trustworthy. Miss it and you’re looking at a curve fitted to a shrug.
So audit your own subscription this week. Open whatever it prescribed for Thursday, then go find the specific number in your ride history that produced it. If you can trace the prescription to a fitted parameter, keep paying. If you can only trace it to a button you pressed last Tuesday, you now know exactly which layer you bought.