AI Coaching Platforms, Tested
I ran eleven weeks of my own training through six different AI coaching systems, feeding each one the same power files, the same HRV numbers, the same sleep data and the same race calendar. The goal was narrow: find out which of these things produces a training decision I would actually follow, and which ones produce something that sounds like a coach but falls apart the moment you check it against your own numbers.
If you are self-coached, riding 8 to 14 hours a week, and you already know your FTP to within about 5 watts, most of the marketing around an ai cycling coach app is aimed at someone else. It is aimed at the rider who does not know what a threshold is. That rider is well served. You are not, at least not yet, and the gap between those two audiences is where most of the disappointment lives.
What I Tested And How
Six systems, three categories.
Closed-loop adaptive platforms: TrainerRoad’s Adaptive Training with Red Light Green Light, Xert (Silver tier), and Join Cycling. These ingest your rides, maintain an internal model of your fitness, and emit a plan. You do not see the model.
Analytics platforms with AI layers: intervals.icu (free, with its own fitness model plus the Athlete Wellness tracking), and WKO5’s Power Duration Model with iLevels. WKO5 is not marketed as AI, but its mFTP and Pmax modelling is doing the same job as the black boxes, and it shows you its work, which makes it the honest control group.
General-purpose LLMs with structured input: Claude Opus 5 and GPT-5, fed CSV exports from intervals.icu via its API, plus a written athlete profile. This is the category most riders are curious about and the one where technique matters most.
The test rider is me: 41, 72 kg, FTP 312 W measured on 2026-07-14 by a 20-minute test times 0.95 (ramp test the same week gave 318 W, which is the usual overshoot), 5-second peak 1,180 W, 5-minute 392 W, 60-minute best of 296 W from a national 25-mile TT on 2026-08-09. Training Peaks CTL sat between 78 and 94 across the test window. UK-based, so the rides include the specific misery of a February Chilterns gravel loop where 40% of the ride is spent below 100 W because of gates, mud and traffic.
The Test That Broke Three Platforms
Here is the ride file I used as the first filter. Saturday 2026-08-01, 4h 11m, 3,400 kJ, a group ride from Newbury out to the Marlborough Downs. Normalised power 241 W, average power 198 W, variability index 1.22, and roughly 22 minutes above 350 W in bursts of 20 to 90 seconds on the climbs. TSS came out at 248.
A good coach looks at this and says: you did not train your threshold, you did a big aerobic day with neuromuscular sprinkles, your legs will feel fine by Monday, and the TSS number is inflated relative to the actual stress because those bursts are anaerobic contributions that IF does not handle well.
What the platforms said:
| Platform | Reading of the 2026-08-01 ride | Verdict |
|---|---|---|
| TrainerRoad AT | Logged as 248 TSS, flagged next Tuesday’s VO2max session as “at risk”, suggested reducing to 3x4min from 5x4min | Over-corrected |
| Xert | Awarded 47 XSS High, 31 XSS Low, raised Peak Power estimate to 1,211 W, kept Tuesday intact | Closest to right |
| Join | Rescheduled the week, moved the interval day to Wednesday, dropped total weekly load 11% | Reasonable but opaque |
| intervals.icu | 248 TSS, no recommendation (it does not make them), eFTP unchanged at 309 W | Correct and silent |
| WKO5 | 231 TSS using iLevels-adjusted scoring, flagged FRC drawdown of 4.2 kJ | Most precise |
| Claude Opus 5 with CSV | Identified the variability index issue unprompted, said TSS overstates stress by 8 to 12%, kept the interval day | Right, with caveats below |
Xert and WKO5 got there because both maintain a two- or three-system model of energy supply rather than a single stress score. Xert’s XSS splits into Low, High and Peak; WKO5 tracks FRC (functional reserve capacity) separately from mFTP. TrainerRoad’s Red Light Green Light asked me how the ride felt and I said “good”, and it still downgraded Tuesday, which suggests the survey input is weighted less heavily than the raw TSS spike.
That is the single most useful thing I learned. Any system that reduces a ride to one number will mis-read a ride with a variability index above about 1.15, and UK group rides and gravel loops routinely sit at 1.2 to 1.35. If your riding is lumpy, single-score platforms will systematically over-estimate your fatigue and under-train you.
Where TrainerRoad’s Adaptive Training Actually Works
I want to be fair to it, because on indoor structured blocks it is genuinely the strongest of the six.
Test: eight weeks of Sweet Spot Base Mid Volume II, done indoors on a Kickr Core, nothing else. Across that block AT made 23 adjustments. I checked each one against what I would have done. It agreed with me on 19, was more conservative than me on 3, and once (week 6) it pushed a Progression Level from 5.8 to 7.2 on over/unders after I nailed a session, which was aggressive but correct. I completed the 7.2 session at 94% compliance.
The Progression Level system is the cleverest part and the least discussed. Each workout gets a difficulty score from 1.0 to 10.0 within its zone, and your level per zone moves independently. So you can be Threshold 7.1 and VO2max 4.3, which is exactly the profile of a TT rider who never does 3-minute efforts. When I finally did a VO2max block, AT started me at 4.3 and walked me up 0.4 to 0.6 per session. That progression rate is more patient than most riders self-prescribe, and patience is what most self-coached riders lack.
Where it fails: outdoor rides, unstructured rides, and anything it did not prescribe. Feed it a 4-hour Sunday and it treats the TSS as pure interference. Over the eleven weeks, AT reduced or moved a quality session after an outdoor ride nine times. Seven of those nine were unnecessary by my own judgement and by how the subsequent session actually went. If more than about 40% of your volume is outdoors and unstructured, you are paying £16.50 a month for a system whose core competency you are not using.
Xert’s Model, And Why Its Numbers Move
Xert is the most technically interesting of the closed platforms because its model is at least partly documented. It fits three parameters to your data: Threshold Power (TP), High Intensity Energy (HIE, comparable to W′ or FRC), and Peak Power (PP). Every ride updates them.
My numbers across the window:
Date TP HIE PP Signature change trigger
2026-07-14 308 19.4 1,150 20-min test
2026-07-29 311 19.8 1,163 3x8min @ 335W
2026-08-01 311 20.1 1,211 Group ride burst 1,204W actual
2026-08-09 316 18.9 1,198 25-mile TT, 296W for 52:41
2026-08-26 314 21.6 1,205 Crit-style Zwift race
2026-09-12 309 22.4 1,209 Two weeks low volume, HIE up TP down
Look at the last row. TP fell 5 W and HIE rose 0.8 kJ over a low-volume fortnight where I did two Zwift races and nothing else. That is the model behaving correctly: I lost threshold and gained anaerobic capacity, because that is what racing crits and skipping tempo does to you. No other platform in the test detected the trade. TrainerRoad kept my FTP flat. intervals.icu’s eFTP dropped to 303 W, which caught the direction but not the compensating anaerobic gain.
The cost is volatility. Xert’s TP moved 8 W across nine weeks, and some of that movement is real physiology and some is model noise from single hard efforts. If you set your training zones from a number that swings 8 W, your tempo sessions drift. My practical fix: I let Xert’s signature inform my interval targets but I recalculate zones from it only once a month, on the first Monday, and I use a 14-day trailing average of TP rather than today’s value. That removes maybe 60% of the jitter and costs nothing.
Xert is £8.99 a month at Silver. For a rider with a strong sprint whose racing is variable, it is the best value in the test. For a pure TT rider doing steady work, it will spend a lot of effort modelling a Peak Power you never use.
Feeding An LLM Your Actual Data: The Method That Works
This is the part most riders get wrong, and the difference between a useless answer and a good one is entirely in what you put in.
The failure mode: pasting a screenshot description or a week’s summary and asking “how’s my training going”. You get back a competent-sounding paragraph about progressive overload. Worthless.
The method that works: pull structured data from the intervals.icu API, hand over 90 days of it, plus a written profile, plus the specific decision you need made. Here is the actual call:
curl -u API_KEY:$ICU_KEY \
"https://intervals.icu/api/v1/athlete/i123456/activities?oldest=2026-07-01&newest=2026-09-30" \
-o rides.json
Then reduce it. The full JSON per activity is around 300 fields and most are noise. I cut it to: date, moving time, distance, average power, normalised power, TSS, intensity factor, variability index, average HR, max HR, decoupling (Pw:HR), kJ, elevation, and my own RPE from the notes field. Fourteen columns, 90 rows, about 9 kB of CSV. That fits comfortably and the model can actually reason across it.
The profile block matters as much as the data. Mine:
Age 41, 72 kg, male. FTP 312 W (20-min test 2026-07-14, x0.95).
5s 1180 W, 1min 570 W, 5min 392 W, 20min 328 W, 60min 296 W.
Target event: national 25-mile TT, 2026-10-18, flat dual carriageway,
aiming sub-52:00 (needs ~305 W for 52 min at current CdA 0.215).
Secondary: gravel 130 km 2026-11-08, no result target.
Constraints: 10-12 h/week available, two weekday evenings only (max 90 min),
Saturday long, Sunday optional. No indoor trainer Jul-Sep.
History: threshold responds fast (3 weeks), VO2max responds slowly.
Two previous overreaching episodes, both after 3 consecutive weeks >700 TSS.
Known weakness: cannot hold aero position past 35 min without power drop ~4%.
That last line changed the answer. Asked to build a four-week run-in to the TT, Claude Opus 5 with the profile put two position-hold sessions a week at 285 to 295 W in the bars, 2x20 minutes, explicitly framed as aero-endurance rather than physiological work. Without the profile line it prescribed generic threshold work. GPT-5 with the same input also caught it but spent more words on nutrition I had not asked about.
The honest assessment of LLMs here: excellent at reading a data set and telling you what is in it, excellent at explaining why a platform’s recommendation might be wrong, good at building a four-week block, and unreliable at arithmetic across many rows. Claude Opus 5 twice summed a week’s TSS incorrectly, once by 34 points. Both times it got the training conclusion right anyway, but if you are using the number, check it. Ask for the working, or better, do the aggregation yourself in intervals.icu and hand over the weekly totals already computed.
The Prompts That Produced Something Useful
Four prompt patterns earned their place. All of them attach the CSV and the profile.
The decoupling interrogation. “Here are my last twelve rides over 2.5 hours with Pw:HR decoupling. Tell me whether my aerobic durability is improving, and separate genuine adaptation from confounds (heat, caffeine, sleep, fuelling as noted in RPE column).” This works because decoupling is a noisy metric that needs context to interpret, and context is what an LLM is for. My decoupling on 3-hour rides went from 6.8% in July to 4.1% in September, and the model correctly flagged that three of the July rides were above 28°C, so the real improvement was smaller than the headline. UK summer 2026 had a genuinely hot spell in late July, so this mattered.
The disagreement audit. “TrainerRoad reduced Tuesday’s VO2max session from 5x4min at 108% to 3x4min at 105% after Saturday’s ride, which is attached. Argue both sides, then tell me what you would do and what evidence would change your mind.” Forcing both sides suppresses the agreeableness problem. Seven times out of nine this produced “the reduction is unnecessary, here is why”, which matched my own read.
The taper arithmetic. “Given CTL 89, ATL 96, TSB minus 7 on 2026-10-04, and a target TT on 2026-10-18, produce a day-by-day TSS plan that lands TSB between plus 12 and plus 20 on race day, keeping at least two openers-style sessions.” This is a constrained optimisation problem with a known answer shape, and LLMs handle it well provided you state the constraint numerically. It produced a 14-day plan landing TSB at plus 15.4. I ran it. Actual TSB on race morning: plus 14.1, off by 1.3 because I rode 20 minutes longer on the Wednesday.
The negative prompt. “Read these 90 rides and tell me the three things I am doing that a good coach would stop immediately.” The answers were specific: my Z2 rides averaged intensity factor 0.71 which is tempo not endurance (true, and a known fault), I had no ride under 90 minutes in ten weeks so no genuine recovery days (true), and my intervals clustered at 105 to 112% FTP with nothing between 95 and 102% (true, and the gap that matters most for a 52-minute TT).
Join Cycling, And The Problem With Not Showing Your Work
Join is the most polished of the three closed platforms and the hardest to evaluate. It produces a plan that updates daily, reads well, and asks you almost nothing. Cost is around €15 a month.
Over eleven weeks it built sensible-looking weeks. Its handling of my constraints (two 90-minute weekday slots) was better than TrainerRoad’s, which kept offering me 2-hour Wednesday workouts I had no room for. Join respected the slots without being told twice.
But when it moved my interval day from Tuesday to Wednesday after the Marlborough ride, I could not find out why. There is no exposed model, no FRC drawdown figure, no progression level. The app said, in effect, trust me. Sometimes trust is fine. Three weeks before a target event, when the question is whether to do the session that might cost you the race or skip the session that might cost you the form, “trust me” is not a sufficient answer, and you cannot argue with a system that will not tell you its reasoning.
This is the structural argument for the hybrid approach: run a platform for the plan, run an LLM over your own exported data for the audit. The LLM cannot see inside Join either, but it can look at your files and tell you whether Join’s decision is consistent with them. For a fuller side-by-side of costs, UK availability, Strava and Garmin integration quality and which platforms handle British weather and club-run culture sensibly, I have broken it down in detail at /ai-coach-app-comparison-uk/, including the ones that assume you have a trainer in a garage at 18°C year-round.
The Numbers Nobody Gets Right
Three specific measurement problems came up repeatedly, and every platform handled at least one of them badly.
Outdoor power meter drift. My Assioma Duos and my Kickr Core disagree by 2.4% on average, Assioma reading higher. Over a season that is a 7 W FTP difference depending on which device set the number. None of the six platforms asked which device produced which file or offered a per-device calibration offset. intervals.icu at least lets you filter and compare by device in its power curve view, which is how I found the discrepancy. Fix it manually: pick one device as your reference, and if you must mix, apply the offset yourself before you set zones.
Coasting and zero-handling. That Chilterns gravel loop, 3h 40m elapsed, 2h 51m moving. Different platforms computed average power over different denominators. Xert and WKO5 handled it consistently; TrainerRoad’s TSS for the ride came out 19 points higher than WKO5’s for the same file. Nineteen TSS is not nothing when a platform is making a go or no-go decision on a threshold session.
HRV as an input. Whoop gave me a recovery score of 31% on 2026-08-27, the morning after a Zwift race, and 88% on 2026-08-28. My rMSSV from a Polar H10 taken supine at wake, four minutes, was 42 ms and 46 ms on those two days, a 9% difference, not a 184% one. HRV-driven readiness scores are compressing a small physiological signal into a dramatic number. Any AI system that takes that number as primary input will be jerked around by it. I now feed rMSSV directly into intervals.icu’s wellness fields and ignore the proprietary scores entirely, and when I hand data to an LLM I hand it raw rMSSV with a 7-day rolling mean, not a recovery percentage.
What I Actually Run Now
The stack after eleven weeks of testing, at £8.99 plus £0 plus an API bill of about £3 a month:
Xert for the physiological model and interval targets, because its three-parameter signature reflects what my rides actually did. intervals.icu as the data spine, free, with the API as the export path and the wellness fields holding raw rMSSV and weight. A monthly Claude Opus 5 session over 90 days of reduced CSV, running the four prompt patterns above, producing a written block plan I then enter manually into intervals.icu as a plan. WKO5 quarterly, for a proper Power Duration Curve check against Xert’s numbers, because two independent models disagreeing is the earliest warning you get that one of them is wrong.
What I dropped: TrainerRoad, reluctantly, because 60% of my volume is outdoors and unstructured and its core strength is wasted on me. If I were doing a winter indoor block, January through March, I would pay for it again for exactly those twelve weeks and cancel in April. Join I dropped because I could not audit it, which is a statement about me rather than about the quality of its plans.
The uncomfortable finding: the thing that improved my training most over the eleven weeks was not any platform’s recommendation. It was the negative prompt, run against my own data, telling me that my endurance rides were 0.71 IF and I had no sessions between 95 and 102% FTP. I fixed both. My 25-mile TT on 2026-09-20 came in at 52:38 against 53:51 in August, averaging 303 W against 296 W, which for a 41-year-old who has been doing this for nine years is a real jump. No adaptive algorithm told me to do that. A model reading my own files told me what I was avoiding, and then I stopped avoiding it.
In this section
The supporting pages under this subject.