FTP, Critical Power And eFTP: Three Numbers, Three Different Meanings
Open intervals.icu, WKO5 and your own 20-minute test spreadsheet on the same afternoon and you will get three numbers that disagree. Most riders treat the gap as noise: three instruments pointing at one true threshold, each slightly miscalibrated, so average them or trust the one you like. That framing is wrong, and it is expensive. FTP, critical power and eFTP are outputs of three different models, built on three different assumptions, answering three different questions. The disagreement is the models talking, not the sensors.
This matters because almost nothing in your training plan uses these numbers directly. Zones do. Every zone boundary below threshold is a percentage of whichever number you typed into the box, so choosing the wrong model doesn’t produce one slightly wrong figure, it produces a whole ladder of wrong figures, each wrong by the same proportion, in the same direction, for as long as you leave it there.
FTP vs critical power is not a measurement dispute
FTP, in the Coggan and Allen formulation, is a single-parameter model. One number describes the boundary of quasi-steady-state, and everything else is a ratio of it. The classic estimate, 95% of a 20-minute maximal effort, is not a physiological derivation; it’s a regression shortcut that happened to fit the riders in the original dataset reasonably well. The model’s core assumption is that riders are shaped alike: that the relationship between your 20-minute power and your hour power is roughly the same as everyone else’s.
Critical power comes from a two-parameter hyperbolic model, traced back to Monod and Scherrer in 1965 and refined heavily since. It says that work above a threshold comes out of a finite tank: P = CP + W'/t. CP is the asymptote, the power at which the tank stops draining. W’ is the tank, measured in kilojoules. Two parameters instead of one, which means the model can tell two riders apart. It also means it makes falsifiable predictions about how long you can hold a given power, which FTP cannot.
Then there’s eFTP, which is a different kind of thing entirely. intervals.icu computes it by taking your best recent maximal effort of roughly three minutes or longer, assuming a standard curve shape, and reading off what FTP that effort implies. Zwift’s zFTP works on similar logic against your longer efforts. Neither is a physiological model. Both are shape assumptions applied to a single data point, plus a bookkeeping rule about when to revise. Useful, genuinely, because they update from racing without you ever doing a test. But an eFTP that moved 8 W after Tuesday’s crit is telling you something about the curve’s assumed curvature, not about your mitochondria.
| FTP (20-min × 0.95) | Critical power | eFTP / zFTP | |
|---|---|---|---|
| Parameters | 1 | 2 (CP, W’) | 1, fitted to a fixed shape |
| Inputs | One 20-min effort | Two or more maximal efforts, ~2-15 min | Any single recent maximal effort |
| Assumes | Your curve shape matches the population | Work above CP comes from a fixed finite reserve | Your curve’s curvature matches the population |
| Breaks when | W’ is unusually large or small | You include efforts longer than ~20 min | Your best recent effort is short and you are punchy |
| Actually predicts | Nothing, it’s a reference value | Time to exhaustion at any power above CP | Nothing, it’s a reference value |
A worked example where the three numbers collide
Take Ellie, 63 kg, a category 2 crit and gravel rider with a big anaerobic reserve. Here is her intervals.icu power curve from a solid six-week block:
dur power
3:00 406 W
5:00 343 W
8:00 308 W
12:00 289 W
20:00 268 W
Her 20-minute test gives FTP = 268 × 0.95 = 255 W.
Fit the two-parameter CP model to her 3-minute and 12-minute efforts. Work done at 3:00 is 406 × 180 = 73,080 J; at 12:00 it’s 289 × 720 = 208,080 J. So CP = (208,080 − 73,080) / (720 − 180) = 250 W, with W’ = (406 − 250) × 180 = 28.1 kJ. That is a large reserve, which is exactly what you’d expect from a rider who wins by attacking.
Now fit the same model to her 5-minute and 20-minute efforts instead, which plenty of people do because those are the efforts they actually have. CP drops to 243 W and W’ inflates to 30.0 kJ. Same rider, same week, two defensible protocols, 7 W apart. The 20-minute effort is outside the model’s valid range, and because real riders fall below the hyperbolic line past about 15 minutes, including it drags CP down and pushes W’ up to compensate. If you want the model to behave, feed it efforts between about two and twelve minutes and leave the long ones out.
Meanwhile intervals.icu is showing her eFTP as 272 W, because that 406 W for three minutes sits far above the standard curve for a 255 W rider. The platform isn’t malfunctioning. It’s correctly reporting that a 406 W three-minute effort would imply a 272 W FTP in a rider with average curvature. Ellie does not have average curvature.
What the error does to every zone underneath
Set her zones off eFTP 272 instead of FTP 255 and the whole ladder shifts up by 6.7%:
| Zone | On FTP 255 | On eFTP 272 |
|---|---|---|
| Endurance (56-75%) | 143-191 W | 152-204 W |
| Tempo (76-90%) | 194-230 W | 207-245 W |
| Sweet spot (88-94%) | 224-240 W | 239-256 W |
| Threshold (95-105%) | 242-268 W | 258-286 W |
Her critical power is 250 W. Read the sweet spot row again. On the eFTP figure, “sweet spot” runs from 239 to 256 W, which straddles and then exceeds the power at which her W’ starts draining. A session she has been told is comfortably sub-threshold aerobic work is, at the top of the range, above the heavy-severe boundary.
The CP model will tell you precisely how that goes wrong, which is the whole reason to have two parameters. Prescribe 3 × 12 min at 272 W. That’s 22 W above CP, so W’ drains at 22 J/s, giving 28,100 / 22 = 1,277 s, about 21 minutes of total accumulated time above CP before the tank is empty. Recovery valleys reconstitute some of it, but slowly and incompletely. The prediction is that interval three falls apart somewhere around the nine-minute mark, and that she will feel it as a motivation problem rather than a maths problem. Run the same arithmetic at the FTP-derived 255 W: 5 W above CP, 5,620 s of budget, 94 minutes. Trivially completable across 36 minutes of work. One number produces a session she finishes; the other produces a session she fails, and neither number is measuring anything different about her fitness.
Flip the rider and the error flips too. Tom, 78 kg, time trials, W’ of 14 kJ and CP of 300 W. His three-minute power is only 378 W, lower than Ellie’s despite being 50 W stronger at threshold. His 20-minute best is 308, so FTP lands at 293 W, seven watts below his CP. For a rider like Tom, the 0.95 multiplier is systematically pessimistic, and the endurance and tempo zones built on 293 are all slightly too easy. Over a 14-hour week that’s real aerobic work left on the table.
Where the AI tools currently let you down
Paste a power curve into an LLM and ask for your FTP and you will usually get one of three failure modes. It averages the numbers. It states that FTP and CP are “essentially the same thing, around 260 W”. Or it does the hyperbolic algebra correctly and then silently uses your 40-minute effort as one of the two points, which invalidates the fit without any warning. TrainerRoad’s AI FTP Detection avoids the problem differently, by not estimating a physiological quantity at all: it predicts what number will let you complete their workouts, which is a legitimate answer to a question you may not have asked.
A prompt that holds up against a real file has to pin down the model, the valid input range and the output you want to check:
Here are my maximal mean powers from the last 42 days:
2:00 458W, 3:00 406W, 5:00 343W, 8:00 308W, 12:00 289W, 20:00 268W
1. Fit the 2-parameter CP model using ONLY efforts between 2 and 12 minutes.
Report CP and W', and show the work-vs-time linear fit residual for each point.
2. Separately report 20min x 0.95 as the Coggan FTP estimate.
3. Do not reconcile them. Tell me which of CP or FTP is higher and by how much,
and state what that gap implies about the size of my W' relative to average.
4. Using CP and W', predict time to exhaustion at 260W, 272W and 285W.
5. Flag any effort that looks pacing-limited rather than physiologically maximal.
Point 5 is the one that catches the most real problems. A 12-minute effort set on a climb where you eased off at the top is not a maximal effort, and it will pull CP down by several watts while looking perfectly respectable in the table. Point 3 matters because the sign and size of the gap is itself information. When CP sits well above your 20-minute-derived FTP, you have a small W’ and a high aerobic ceiling. When eFTP sits well above both, you have a big W’ and eFTP is reading your anaerobic capacity as aerobic fitness. More on how these fits behave across a season, and how to tell a genuine threshold shift from a model artefact, in our guide to FTP estimation and fitness models.
None of the three numbers knows anything about durability, incidentally. CP measured fresh is not CP measured after 2,500 kJ, where most riders lose 5-10% and some lose considerably more. That’s a third axis the two-parameter model doesn’t have, and it’s where the interesting work is happening.
So pick deliberately. Use CP and W’ for prescribing anything above threshold, because it’s the only one of the three that makes a prediction you can falsify by Thursday. Use a 20-minute-derived FTP for setting sub-threshold zones if your W’ is near average, and adjust the multiplier if the CP fit tells you it isn’t. Treat eFTP as a change detector rather than a value: when it jumps, go and look at which effort caused the jump before you let it rewrite your zones.
Your next move is a test, not a spreadsheet. Two maximal efforts, three minutes and twelve, on separate days, on the same road, fresh. Then compare what the model says you should be able to hold for 30 minutes against what you actually hold, and let the residual tell you which of your three numbers has been lying.