AI Cycling
005 FTP Estimation And Fitness Models 1,543 words · 7 min

How Accurate Is Automatic FTP Detection? Testing Five Platforms Against A Lab Number

Every platform you ride with has an opinion about your threshold, and none of them agree. I got tired of guessing which one to trust, so I bought a lactate test and treated the whole thing as a measurement problem. One rider, one week, five auto-detected FTP values, and a physiologically defined reference number to score them against.

The result: a 40 W spread on a rider whose actual maximal lactate steady state sits at 288 W. That is 13.9% of the reference value, and it is the difference between a threshold session that works and one you abandon in rep two. If you care about ai ftp detection accuracy because you’re prescribing your own intervals off these numbers, the spread is the finding. Not the average, not the “they’re all roughly in the same ballpark” reassurance. The spread.

The reference number

Test subject: 34, male, 71 kg, UK 2nd cat, 11 hours a week, mostly road with a spring TT block. Nothing exotic about the physiology, which is the point.

MLSS was determined the boring way, over four visits inside nine days. Each visit was a single 30-minute constant-load ride on a Cyclus 2 ergometer with capillary lactate drawn at 10 and 30 minutes. MLSS is defined as the highest power at which the lactate rise between those two samples stays at or below 1.0 mmol/L.

Visit   Power    La @10min   La @30min   Δ        Verdict
1       275 W    2.8         3.1         +0.3     sub-MLSS
2       292 W    3.4         4.9         +1.5     supra-MLSS
3       285 W    3.2         3.9         +0.7     sub-MLSS
4       288 W    3.1         4.0         +0.9     MLSS

So: 288 W, 4.06 W/kg. Before anyone blames hardware, the power meters were cross-checked by dual recording across three rides. The Favero Assioma Duo pedals read 1.4% higher than the Wahoo Kickr Core (about 4 W at threshold), and the Cyclus 2 sat between them. Nothing in the platform spread below is a calibration artifact. All five platforms were fed the same files from the same pedals.

What the five platforms said

The rider’s last 90 days of history were already in all five systems. No fresh tests, no ramp protocols, nothing done to game any of them. This is what each one was displaying on the Monday after the lab visit.

PlatformHow it gets thereAuto FTPvs 288 W MLSS
Strava Premium (estimated FTP)best 20-minute power × 0.95, no decay312 W+24 W (+8.3%)
Garmin Edge 1040 (auto FTP, Firstbeat)HR/power response within a single ride305 W+17 W (+5.9%)
Xert (Threshold Power)fitness signature fitted to near-maximal efforts296 W+8 W (+2.8%)
intervals.icu (eFTP)power-duration curve fit to 3 to 20 min bests283 W-5 W (-1.7%)
TrainerRoad (AI FTP Detection)ML model over ride history272 W-16 W (-5.6%)

Two of the five were within 2% of a lab-defined MLSS, which is genuinely good. The other three were not, and they missed in opposite directions, which is worse than all of them missing the same way. A consistent bias you can correct with a multiplier. A bidirectional 40 W spread just tells you the platforms are answering different questions.

Strava’s 312 W traces back to a single 20-minute effort of 328 W set during a Tuesday chaingang on the A4130. Sitting fourth wheel for eleven of those twenty minutes. Strava’s estimate does not know or care that the effort was draft-assisted and not maximal in the way the 0.95 multiplier assumes, and it never decays, so 312 W sat there as gospel for six weeks.

Garmin’s auto FTP landed at 305 W after a hot Sunday road race. Firstbeat’s method infers threshold from the heart rate response to the power you produced, and heat plus a race-day sympathetic load distorts that relationship in the direction of “you were producing this power at a heart rate that implies a higher threshold.” It is a one-ride judgement, and it showed.

The number changes what happens to you in rep two

Here is the part that turns a benchmarking exercise into a training problem. The rider ran the same session, 3 × 15 minutes with 5 minutes recovery, in three separate weeks at three different anchor values, each prescribed as “100% of FTP.” Lactate sampled at 3 and 14 minutes of each rep.

Anchor 272 W (TrainerRoad)   = 94% of MLSS
  Rep 1: 2.4 -> 2.7 mmol/L   RPE 6/10   HR drift +2 bpm
  Rep 3: 2.6 -> 3.0 mmol/L   RPE 7/10   completed, could have done a fourth

Anchor 288 W (lab MLSS)      = 100% of MLSS
  Rep 1: 3.1 -> 3.9 mmol/L   RPE 7/10   HR drift +4 bpm
  Rep 3: 3.5 -> 4.4 mmol/L   RPE 8/10   completed, nothing left

Anchor 312 W (Strava)        = 108% of MLSS
  Rep 1: 3.4 -> 7.1 mmol/L   RPE 9/10   HR drift +9 bpm
  Rep 2: abandoned at 8:40, cadence fell 92 -> 81 rpm

Three sessions, one label, three entirely different physiological events. The 272 W version is a good tempo workout that will not do much for lactate clearance at threshold. The 312 W version is a VO2max session with a bad recovery structure, and it cost four days of usable training afterwards. Only the middle one is the session the plan intended.

Zone boundaries move with the anchor too, and the overlap is ugly:

ZoneAt 272 WAt 288 WAt 312 W
Tempo (76-90%)207-245 W219-259 W237-281 W
Threshold (91-105%)248-286 W262-302 W284-328 W

At 265 W the rider is doing tempo on one anchor and threshold work on another. If you’re building sweet spot blocks in Zwift or intervals.icu off a number from Strava, a substantial share of your “threshold” minutes are landing above MLSS, where the whole mechanism you’re targeting stops operating.

Stability matters as much as bias

Accuracy is not only about where a number sits on one Monday. I re-read all five every Monday for six weeks through a build block where the rider’s real threshold genuinely moved, confirmed by a second lactate visit at week six that put MLSS at 294 W.

TrainerRoad crept 272 to 279, understating the change but moving in the right direction and never bouncing. intervals.icu went 283 to 290, tracking the real 6 W gain closely, on the strength of two genuine 8-minute climbing efforts. Garmin rattled between 298 and 311 with no pattern anyone could defend: 13 W of week-to-week noise on a rider whose threshold moved 6 W in six weeks. Xert dropped 296 to 291 during a deliberately easy recovery week and recovered once hard efforts returned, which is decay working as designed but still a signal you cannot prescribe from. Strava did not move at all until a new 20-minute best appeared, then jumped.

The pattern under all of it: platforms that fit a power-duration curve to efforts you actually produced behave like measurement. Platforms that infer from heart rate within a single ride, or that apply a fixed multiplier to your best 20 minutes regardless of context, behave like a guess with a confidence interval nobody shows you. The underlying maths of critical power and the three-parameter models, and why the fitted-curve approach has a structural advantage here, is worth understanding properly if you’re going to depend on one: the FTP estimation and fitness models pillar covers the derivations.

What to actually do with this

Pick one platform as your prescribing anchor and never mix. Based on both bias and stability, intervals.icu eFTP is the defensible default for a self-coached rider, provided you feed it a genuine near-maximal effort of 8 to 20 minutes every two or three weeks. Without that input it goes stale and quietly drifts low.

Treat Strava’s estimated FTP and Garmin’s auto FTP as ceilings rather than prescriptions. They answer “what have you shown you can do in favourable conditions,” which is a useful thing to know before a TT and a terrible thing to base 3 × 15 on.

Then validate your chosen number with a confirmation session, because none of the five is a measurement and you can make one yourself in 50 minutes. Ride 2 × 20 minutes at your anchor with 5 minutes easy between. Pass criteria: heart rate drift inside rep two under 5%, cadence held within 4 rpm of your opening minutes, and a final-minute RPE of 8 rather than 10. Fail it and drop 3%, then retest the following week. A cheap Lactate Plus meter and two fingertip strips at minute 3 and minute 18 turns that into an actual lactate steady state check for under £2 a session.

Next in this series: the same protocol on a 58 kg gravel racer with a very different anaerobic profile, where the Xert signature and the power-curve fits start disagreeing in a way that suggests the 40 W spread here was the easy case.