AI Cycling
§3 Section 3 of 6 2,939 words · 13 min

FTP Estimation And Fitness Models

Every power-based training platform you use is making a guess about you. Not a measurement, a guess: a number derived from a model fitted to whatever you happened to do in the last 42 days, weighted by recency, filtered by some heuristic about what counts as a “good” effort. The number gets a name (eFTP, mFTP, Threshold Power, rFTPw) and a two-decimal confidence you didn’t ask for, and then every zone, every TSS figure and every workout target downstream inherits its error.

This page is about how accurate those guesses actually are, how to test them against your own files, and where the newer AI-branded layers sit in relation to the older curve-fitting maths they’re usually wrapping.

What “ai ftp detection accuracy” Actually Means

There are three different questions hiding inside that phrase, and conflating them is why forum threads about it go nowhere.

Question one: does the model recover a known number? You do a 20-minute test, or a 40k TT, or an actual 60-minute maximal effort. Does the platform’s estimate land near it? This is the question most people think they’re asking. It’s the least interesting one, because the ground truth is itself noisy. Your 20-minute test result on a Tuesday after a rest day and your 20-minute test result on a Saturday after a hard Thursday differ by 3-5% with no change in physiology at all.

Question two: is the model self-consistent? If the platform says 285 W today and 291 W tomorrow off a ride that contained nothing above 250 W, the model has a stability problem regardless of whether either number is “right”. Run-to-run variance under stable fitness is measurable, and it’s the thing that most affects whether you can trust a rising trend.

Question three: does the number work as a prescription? This is the one that matters for training. If your zones built from the estimate produce a 2x20 at “95% FTP” that you can hold comfortably, and a 5x5 at “108%” that buries you by rep three, the estimate is wrong in a way that will actually damage your training, even if it’s within 2% of your lab-measured MLSS.

I’d argue accuracy in sense three is the only one worth optimising, and it’s the one that model documentation almost never addresses.

The Maths Underneath, Briefly

Nearly every estimate you’ll see comes from one of four families.

Percentage-of-a-test. 20-min × 0.95, 8-min × 0.90, ramp-test MAP × 0.75. The 0.95 multiplier for 20 minutes has a standard deviation of roughly ±0.04 across riders, which on a 320 W twenty-minute power is a ±13 W band. The ramp test’s 0.75 is worse: Zwift’s own ramp test applies 0.75 to your best 1-minute power at failure, and for riders with a big anaerobic contribution that overshoots badly. A track sprinter-turned-gravel-rider I know tested 348 W FTP on Zwift’s ramp and could not hold 300 W for 20 minutes.

Critical power / W′ (two-parameter hyperbolic). Fit P = CP + W′/t across two or more maximal efforts, typically 3 min and 12 min. Clean, interpretable, and biased high when your short effort is disproportionately strong. CP from a 3/12 pairing typically lands 2-6% above 60-minute mean maximal power.

Power-duration curve fitting with a fatigue term. Morton’s three-parameter model, Puchowicz’s Bayesian variants, Golden Cheetah’s extended CP model. These add a maximum instantaneous power asymptote and start to behave sensibly at the 5-second end, at the cost of needing more data.

Recency-weighted curve extension. This is what intervals.icu’s eFTP does: it looks for efforts that punch above your current modelled curve, then re-fits the curve upward. It’s specifically designed so that a hard 8-minute climb on a group ride can raise your estimate without you ever doing a test. TrainingPeaks’ mFTP (from the WKO family) is a relative of this, fitting a full power-duration curve and reading FTP off it at the point where the curve’s slope crosses a threshold.

Then there’s the AI layer. Xert’s “signature” tracking (Threshold Power, High Intensity Energy, Peak Power) runs a recursive fitness-signature update on every ride, treating your parameters as a state to be tracked rather than a value to be fitted periodically. Whether you call that AI or a Kalman-flavoured filter is mostly marketing, but the behaviour is genuinely different: it can drop your threshold after a bad week, which most curve-fitters cannot.

A Nine-Month Test You Can Replicate

I ran intervals.icu’s eFTP against periodic ramp tests for nine months and wrote the whole dataset up separately, because the detail deserves its own page: intervals.icu eFTP versus a ramp test, nine months of data. The headline is that eFTP tracked the ramp-derived number with a mean offset of about +6 W and a spread wide enough that any single-day comparison was close to meaningless, but the direction of change agreed in 11 of 13 test intervals.

The method is what I’d like you to steal. It’s four steps and needs nothing but your existing files.

1. Pick a reference protocol. Same protocol every time. Same
   trainer or same road. Same time of day. 48h with nothing
   above Z2 beforehand.

2. Freeze the model estimate on the morning of each test, before
   you upload the test file. Write it down manually. Do not
   trust yourself to reconstruct it later; most platforms
   retroactively rewrite history when the curve changes.

3. Test every 5-6 weeks. Fewer than 6 data points and you
   cannot distinguish a trend from noise.

4. Record a third column: how the prescribed zones felt.
   "4x8 at 100% of estimate: completed, last rep ragged" is
   data. Log it.

Step two is where almost everyone gets it wrong. If you pull “what did eFTP say in March” out of intervals.icu today, you get the value as reconstructed by the current model from the current data, not the value the model was showing you in March. They are different numbers. Your zones in March were built on the March value.

Worked Example: One Rider, Five Models, One Week Of Files

Here’s a concrete case from a rider I’ll call S, 42, UK-based, TT focus, 8-10 hours a week. Same 28 days of files pushed into five platforms. Reference: a 40k TT ridden 9 days after the snapshot in 52:41 at 281 W average, 284 W normalised, and a 45-minute best-effort on the turbo at 288 W three weeks earlier.

PlatformEstimateMethodDelta vs 40k avg
Zwift ramp test312 W1-min MAP × 0.75+11.0%
intervals.icu eFTP291 Wrecency-weighted curve extension+3.6%
TrainingPeaks mFTP287 WWKO power-duration curve+2.1%
Xert Threshold Power279 Wrecursive signature−0.7%
Golden Cheetah CP (3/12)296 Wtwo-parameter CP+5.3%
Manual 20-min × 0.95294 W309 W × 0.95+4.6%

Spread: 33 W, or 11.8% of the lowest estimate. If S had built zones off the Zwift number, “sweetspot” at 90% would have been 281 W, which is their actual 52-minute race power. Their sweetspot sessions would have been threshold-or-above sessions. Six weeks of that and you have a rider who is chronically flat and cannot work out why their “easy tempo” is hurting.

Notice what the pattern is: the two models that ran continuously on the whole file history (Xert, mFTP) sat closest. The two that keyed off a single short maximal effort (Zwift ramp, CP 3/12) sat furthest high. This shows up over and over in TT riders and diesel-type road riders whose anaerobic capacity is modest relative to their aerobic ceiling. Runs the other way for crit racers: if you have a big W′, single-test models will often underread your threshold because the fit gets dragged by your monstrous 3-minute power in a way that flattens the long end.

Where The AI Tools Fit, And What They’re Actually Doing

Three categories, and the distinction matters for how much you should trust their FTP numbers.

Category one: statistical models with AI branding. Xert, Wahoo SYSTM’s 4DP, Form’s threshold tracking. These are parametric or state-space models. They’re often good. Calling them AI is a stretch. Their failure modes are predictable and documented: Xert’s signature can drift during a long block of exclusively low-intensity riding because it has no fresh high-intensity data to constrain High Intensity Energy, and it will start quietly shifting Threshold Power to compensate.

Category two: ML models trained on large ride corpora. This is where things like Athletica’s models and some of the newer TrainerRoad AI FTP Detection live. TrainerRoad’s AI FTP Detection is the most interesting because it’s the most honest about what it is: a machine-learning model trained on their own database of rides and test results, which outputs an FTP and asks you to accept or reject it. TrainerRoad claim it’s within a few percent for most users. In my own testing across four riders it came in tighter than any curve-fit, with one glaring caveat: it works well if most of your riding is inside TrainerRoad. Feed it a summer of unstructured gravel and it degrades noticeably, because the features it was trained on (structured workout compliance, specific interval shapes) aren’t present.

Category three: LLMs reasoning over your data. ChatGPT, Claude, Gemini with your ride files or your intervals.icu data pasted in. This is not an FTP model and should never be treated as one. An LLM cannot fit a power-duration curve reliably by inspection; ask one to estimate FTP from a list of mean maximal powers and you’ll get something plausible-looking derived from half-remembered heuristics. Where they are genuinely useful is in the interpretation layer, which I’ll come back to.

Fitness Models: CTL/ATL Are Not Fitness

Your FTP estimate feeds your zones, and your zones feed your TSS, and your TSS feeds CTL. Which means CTL error compounds FTP error, and the compounding is not linear.

TSS is proportional to the square of intensity factor. If your FTP is overestimated by 5%, every workout’s IF is understated by roughly 4.8%, and TSS is understated by around 9.3%. Over a 12-week block at 600 TSS/week you’d be accumulating a CTL roughly 9% lower than a correctly-calibrated rider doing identical work. That’s not a rounding error, it’s the difference between a CTL of 85 and a CTL of 93 on your chart, and if you’re using CTL ramp rate to govern load you’ll be systematically undercooking or overcooking depending on the direction of the FTP error.

Worse: the error isn’t constant. eFTP tends to run high after a block containing a lot of short hard efforts and drift low during base. So your CTL is being computed against a moving yardstick, which means CTL comparisons across seasons are only as valid as the FTP stability across those seasons. Anyone who has looked at their own five-year CTL chart and concluded something about their 2023 base has probably been reading FTP drift.

If you want CTL you can actually compare across years, fix your FTP value manually for the duration of a block and only change it at block boundaries. It makes the number less “live” and much more useful. intervals.icu supports this directly: set FTP manually and turn off “use eFTP for zones”, but leave eFTP visible so you can watch it and decide. Best of both.

The newer alternatives to impulse-response are worth a look but don’t solve this. The Banister-derived models (PerfPot, Fitness-Fatigue with individually fitted time constants) still take training load as input, and training load still comes from your FTP. Garmin’s Firstbeat-derived Training Status has the same dependency through a different route. There’s no fitness model that escapes the calibration problem by being cleverer about the maths downstream.

A Prompt Pattern That Works On Your Own Files

LLMs can’t estimate your FTP. They’re quite good at spotting inconsistency between a model’s claim and your actual efforts, which is a different and more useful job. The pattern that has held up for me: give the model the evidence, ask it to argue against the estimate, never ask it to produce the estimate.

Here is my mean maximal power curve from the last 42 days
(intervals.icu > Fitness > Power, export the table):

5s 1102 W
1m 498 W
3m 392 W
5m 351 W
8m 322 W
12m 306 W
20m 294 W
30m 288 W
40m 281 W
60m 268 W

My platform estimates FTP at 291 W.

Do not produce your own FTP estimate. Instead:
1. Which of these durations are genuine maximal efforts and
   which are likely submaximal? Justify from the shape of
   the curve.
2. If 291 W were correct, what would you expect my 30m and
   40m numbers to be? How far off are the actuals?
3. What single test would most efficiently discriminate
   between an FTP of 275 W and 291 W?

That third question is the valuable one, and it’s the kind of thing an LLM handles well because it’s reasoning about experimental design rather than doing numerics. In the case above, the answer is a 35-minute effort: the 60-minute number is almost certainly submaximal (a 268 W hour off a 294 W twenty means either the hour was a ride not a test, or the twenty was), and a 35-minute maximal sits in the region where a 275 W and a 291 W threshold predict clearly different outcomes.

What you should not do: paste a .fit file and ask for an analysis. Token limits mean the model will see a sampled or summarised version, the summarisation will lose exactly the interval structure that matters, and you’ll get confident commentary on data the model never actually had.

Reading The Failure Signatures

Once you’ve collected six or more paired data points, the shape of the disagreement tells you which model is broken.

Estimate consistently 8%+ high, in a rider with a strong sprint. Almost certainly an anaerobic contamination problem. The model is reading your W′ as aerobic capacity. Switch to a method anchored on longer efforts: 45-minute best effort, or CP fitted from 12/40 rather than 3/12.

Estimate jumps 10 W after every group ride, then decays. Recency weighting doing exactly what it says on the tin. Your race-day 8-minute power genuinely is above your trained threshold, because adrenaline and drafting and a wheel to follow. The estimate is not wrong so much as measuring a different thing. Fix: exclude races from eFTP detection (intervals.icu lets you mark an activity as not eligible) or just hold FTP manually.

Estimate flat for eight weeks during a block you know went well. Look at what you actually did. Most likely you did no efforts in the 5-20 minute window, so the curve has nothing new to fit. Hidden fitness. One 12-minute maximal effort will reveal a 12 W jump instantly, and the jump was real the whole time.

Estimate drifting down 1-2 W a week with no obvious cause. Check whether your power meter has drifted, particularly if it’s a single-sided crank-based unit in a season where your pedalling balance has shifted. I’ve seen a Stages left-only unit read 4% low over six months after a bearing replacement changed the strain-gauge zero, and every model faithfully tracked the decline.

Estimate high and zones that feel wrong in a specific way: tempo too hard, VO2 too easy. This is a zone-model problem, not an FTP problem. Your threshold might be fine while your %-based zone boundaries don’t match your physiology. Riders with high aerobic durability often find 88-93% genuinely sustainable for hours, which breaks the standard Coggan zone assumption that Z3 ends at 90%.

The Practical Setup

For what it’s worth, the configuration I’d defend for a self-coached rider who wants numbers they can trust:

Set FTP manually, from a real test, once every 5-6 weeks. Use a 40-45 minute maximal effort or a 20-minute test with the multiplier you’ve personally calibrated rather than 0.95. Keep a running log of what your own multiplier is, because it’s stable within a rider and varies a lot between riders. Mine is 0.935; a clubmate who is all diesel runs 0.96.

Let intervals.icu compute eFTP in the background and watch it as an indicator, not an input. When eFTP rises 8 W or more above your manual value and stays there across two weeks, that’s a signal to test. When it falls that far, it’s a signal to look at your last two weeks of sleep and stress before you conclude anything about fitness.

Keep your CTL computed against the manual value so the curve means something across the season. Accept that this makes CTL laggy at block boundaries; that’s a much smaller cost than having a metric whose baseline moves.

Cross-check with a second model annually. If Xert and intervals.icu and your own test all land within 6 W of each other, stop thinking about it and go training. If they spread 25 W, the spread itself is information: look at which model is anchored on which duration, and you’ll usually find your power-duration curve has an unusual shape that one of them is mishandling.

One thing that has changed my own practice more than any model comparison: recording how the prescribed sessions felt, in one line, every time. Six weeks of “4x8 at prescribed threshold: rep 4 failed at 5 min” is a more reliable indictment of an FTP estimate than any amount of curve-fitting argument, and it costs you fifteen seconds a session.

In this section

The supporting pages under this subject.