AI Cycling Coach Apps Compared For UK Riders
Every AI coaching app looks convincing until you feed it a British winter. Three weeks of 4°C drizzle, a Tuesday chaingang that wasn’t in the plan, a Sunday club run that turned into a 92-minute tempo block because someone attacked at the café stop, and two turbo sessions abandoned at 40 minutes because the kids wouldn’t sleep. That’s the input. The question is which tools produce something sane out the other end, and which ones quietly decide you’ve “regressed” and hand you sweet spot for the fourth week running.
This page is the long version of one comparison: what happens when the same rider’s actual data goes through six platforms at once. For the wider landscape of what these systems are doing under the bonnet, and how the big physiological models differ, the parent overview at /ai-coaching-platforms/ covers the theory. Here we’re only interested in output quality against real UK ride files.
The test case
One rider, one dataset, six subscriptions running in parallel from January to June.
Profile: 41 years old, 74 kg, FTP 291 W confirmed on a Kickr Core in early February (20-minute test, 306 W average, 0.95 multiplier), CTL 68 at the start of the block. Works 07:30 to 18:00 weekdays. Rides Tuesday chaingang (outdoor, April onward), Wednesday and Thursday turbo, Saturday club run, Sunday long. Two A-events: the Gralloch gravel race in May, and a club 25-mile TT in June against a standing PB of 56:42.
All six platforms received the same history via Strava or a direct .fit sync, plus the same stated goals and the same declared weekly availability (roughly 8 to 10 hours, with a hard 75-minute ceiling Monday to Thursday). Nothing was hand-tuned afterwards. The point was to see what each one does when a normal person just uses it.
What each tool actually is
Worth being precise here, because “AI” covers three very different things in this market.
| Tool | What the “AI” does | Adapts to outdoor rides? | Approx. UK cost |
|---|---|---|---|
| TrainerRoad | Machine-learning compliance model (Adaptive Training) reshapes future workouts after each session | Partially, via Ride Analysis | ~£18–23/mo annual, billed USD |
| Xert | Continuous three-parameter model (Threshold, HIE, Peak Power) recalculated from every file | Yes, natively, this is its core strength | ~£8–12/mo |
| Athletica | Polarised/pyramidal plan generation with weekly re-sequencing | Yes, though it prefers structure | ~£18–22/mo |
| AI Endurance | Physiological modelling plus race-date-driven periodisation | Yes | ~£12–16/mo |
| JOIN Cycling | Plan generator built around availability changes and rescheduling | Yes | ~£12/mo |
| intervals.icu | No AI plan generation; excellent open model, API and custom charts | N/A | Free, donation ~£3/mo |
Prices move with the exchange rate. TrainerRoad and AI Endurance bill in dollars, so the number leaving your account in March is not the number that left in November, and there’s no UK VAT line on some of them because of where the entity sits. Small point, but over a year the TrainerRoad annual plan swung by more than £20 on FX alone.
TrainerRoad: brilliant indoors, deaf outdoors
Adaptive Training is still the best thing in the category at one specific job: taking a failed interval session and producing a slightly easier but structurally identical one next week. The test rider failed a 4x8 at 105% on a Thursday in February. The following Thursday’s session came back as 4x8 at 102%, not 4x6 or 3x8. That’s the right adjustment, and almost nothing else does it that precisely.
Where it fell over was the outdoor season. From mid-April the Tuesday chaingang was producing 250 to 280 TSS in 95 minutes with 14 efforts above 400 W. TrainerRoad ingested it, marked it as an unstructured ride, and continued to schedule a VO2 session for Wednesday. Over six weeks, the plan accumulated an extra 180 to 220 TSS per week beyond what the plan believed it was prescribing. CTL climbed from 68 to 94 in nine weeks, a ramp rate averaging 2.9 per week, which is well above the 3 to 5 per week upper bound most coaches would call sustainable for a working masters rider. Two chest infections later, the block ended early.
If you train indoors from October to March and your outdoor riding is genuinely easy, this is still the strongest product on the list. If your bunch rides are hard, and in the UK they usually are, you’ll be doing the load management yourself.
Xert: the only one that read the chaingang correctly
Xert’s model treats every file as a signal about three parameters rather than as a workout to be graded. After the same Tuesday chaingang, Xert reported this kind of change:
Tue 14 Apr | 1:34:12 | XSS 212 | Difficulty 4.2 (Extreme)
Threshold Power 291 → 294 W
High Intensity Energy 19.8 → 21.4 kJ
Peak Power 1042 → 1038 W
Focus: Rouleur (Breakaway Specialist)
Wed 15 Apr | Recommended: Endurance, 60-75 min, XSS 55-70
That last line is the whole argument for Xert. It saw a hard ride, revised the fitness signature upward, and then told the rider to go easy the next day. No other platform in the test did that automatically from an unstructured outdoor file.
The cost is interface friction. Xert’s terminology (XSS, MPA, Focus, Specificity) takes a fortnight to internalise, and its workout player is noticeably less pleasant than TrainerRoad’s. It also over-reacts to a single big day: one 5-minute personal best on Sunday pushed Threshold Power up 4 W, which then made Monday’s recommended endurance zones slightly too hard. Manageable, but you have to watch it.
Athletica and AI Endurance: sensible periodisation, less sensitive to chaos
Both of these generate genuinely well-structured macrocycles. Given the 17 May Gralloch date, Athletica built a nine-week base with a 3:1 loading pattern, then moved to 30/30s and 4-minute work six weeks out, then tapered over 11 days with a 42% volume reduction and intensity preserved. That is textbook, and the textbook is correct.
What Athletica handled less well was the gravel-specific demand. The Gralloch is roughly 110 km with about 1,400 m of climbing on fireroad, and the power profile is punchy and repeated, not steady. The prescribed sessions skewed toward 8 to 12 minute threshold blocks when what the event needed was repeated 45-second to 2-minute surges off a tempo baseline. Asking the plan builder to weight “punchy” higher improved it, but you have to know to ask.
AI Endurance produced a more polarised distribution, around 82% of weekly time below LT1 and the rest genuinely hard, which suited the 8 to 10 hour budget better than the sweet-spot-heavy alternatives. Its FTP estimate tracked a hair conservative throughout: it held 284 W in April when a 20-minute field test on a local drag gave 302 W. Conservative is a defensible error, but if you race off percentages you’ll be training slightly soft.
JOIN: built for a diary that keeps changing
JOIN is the only one of the six that treats “I can’t ride Thursday any more” as a normal event rather than a failure. Move a session in the app and the whole week reshuffles, protecting the key workout and downgrading the filler. Over the 12 weeks, 23 sessions were moved or shortened. JOIN absorbed all of them without ever producing a week with two hard days back to back.
Its ceiling is lower. The intensity prescriptions are less precise than Xert’s, the analysis after each ride is thin, and it has no real concept of a time trial as a distinct discipline. For a rider whose main problem is life rather than physiology, it’s the most usable thing here.
Pointing an LLM at your own files
The most interesting result in the test didn’t come from a coaching app at all. intervals.icu exposes a clean API, and exporting 90 days of activity summaries as CSV takes about two minutes:
GET https://intervals.icu/api/v1/athlete/{id}/activities.csv
?oldest=2026-01-05&newest=2026-04-05
Feed that to a capable model with a prompt that specifies constraints rather than asking for a plan:
Here are 90 days of ride summaries: date, duration, TSS, IF, normalised power, average HR, decoupling, and time in each of six power zones. FTP is 291 W, weight 74 kg. Available: 75 min max Mon–Thu, 2.5 h Sat, 4 h Sun. Target event 17 May, 110 km gravel, 1,400 m, expect repeated 1–2 min efforts at 110–130% FTP off a 78% baseline. Identify the three clearest weaknesses in this data, cite the specific rides that show them, then write next week’s sessions. Do not propose anything that exceeds my stated time limits.
The response correctly identified that Sunday rides showed HR-to-power decoupling above 8% after 2h45 (six of nine rides), that time above 110% FTP totalled 41 minutes across the entire 90 days, and that Wednesday and Thursday sessions were near-identical in stimulus. All three were true, and none of the six paid apps flagged the second one.
What it got wrong was continuity. Asked again the following week, without the transcript, it produced a plan that ignored the previous week’s progression entirely. LLMs are excellent diagnosticians of a fixed dataset and poor at holding a 16-week arc in their head. Used as an analyst that reports to you, rather than a coach that directs you, the value is real and the cost is roughly zero.
The 25-mile TT, and what it exposed
Six weeks out from the club 25, the useful question is narrow: can the platform build toward 45 to 55 minutes at 88 to 93% of FTP in a position that costs you 15 to 20 W? Only Xert and AI Endurance produced sessions of the right shape, and neither knew anything about the position. TrainerRoad’s TT-specific plans exist but assume you’ll do the aero work yourself.
The rider went 55:58 in June, a 44-second improvement, averaging 268 W on a sporting course. The sessions that plausibly mattered were 2x25 and 3x20 at 92 to 95% held on the TT bike outdoors, which came from Xert’s recommendations plus an hour’s manual editing. Not one of the six would have produced that block unprompted from a cold start.
If you’re choosing one thing to start with this winter, take a month of your own files, run them through Xert’s free trial and through an LLM with the prompt above, and compare what each says your weakness is. If they agree, that’s your next training block. If they disagree, you’ve learned something more useful than either answer.