AI Cycling
§3.1 FTP Estimation And Fitness Models 1,799 words · 8 min

intervals.icu eFTP Versus A Ramp Test: Nine Months Of Data

I started logging this in January because I was annoyed. intervals.icu kept telling me my eFTP was 312W while my last ramp test said 288W, and I had no idea which number to build my winter blocks around. So I did the boring thing: every time eFTP moved, I wrote down the date, the ride that triggered it, the value, and what my power duration curve looked like around it. Then I ramp tested roughly every six weeks and logged those too. Nine months, 41 eFTP revisions, seven ramp tests, one 20-minute test and two 40km TTs as reality checks.

What follows is what the data actually said about intervals icu eftp accuracy, where the estimate drifts, and the specific ride types that make it lie to you. If you want the wider context on how FTP estimation models differ from each other, the FTP Estimation And Fitness Models pillar covers the landscape. This page is narrower: one athlete, two methods, nine months, and the gap between them.

What intervals.icu eFTP Is Actually Doing

eFTP is not a test result. It’s a curve fit. intervals.icu takes your power duration curve, fits a hyperbolic model to the points between roughly 3 and 30 minutes, and derives an FTP from where that curve sits. David Tinker’s implementation uses a modified Morton/Monod three-parameter approach rather than the naive two-parameter critical power model, which matters because the two-parameter version overestimates badly when your only hard efforts are short.

The practical consequence: eFTP responds to whatever your hardest recent effort in that window happened to be. It doesn’t know whether you were fresh, whether you were racing, or whether the 5-minute PB came off a 20-minute descent. It sees watts and duration.

You can see the mechanics in the activity page. When a ride triggers a revision, intervals.icu tells you which effort did it:

Activity: Thursday Chaingang
eFTP: 312W (was 305W)
Triggered by: 8:42 @ 341W
Model: 3-param CP, W' = 21.4 kJ, CP = 298W

That 8:42 is the whole story. One effort, one revision. Nothing about the other 92 minutes of the ride mattered.

The Nine-Month Numbers

Here’s the comparison at each ramp test date. Ramp tests were Zwift’s standard protocol (20W/min after warm-up, best 1-minute × 0.75), done on a Wahoo Kickr V5, same bike, same fan setup, morning, fasted, 48 hours after a rest day.

DateRamp FTPeFTP that weekGapeFTP trigger effort
8 Jan288W312W+24W4:11 @ 372W (Zwift race)
19 Feb295W318W+23W5:02 @ 358W (club chaingang)
2 Apr306W321W+15W12:38 @ 327W (hill repeat)
14 May314W327W+13W20:04 @ 318W (TT warm-up block)
25 Jun318W324W+6W24:17 @ 314W (sportive climb)
6 Aug322W319W−3W18:41 @ 311W (solo threshold)
17 Sep319W331W+12W3:48 @ 391W (crit-style group ride)

The pattern is not subtle. The gap tracks the duration of the triggering effort, not my fitness. January and February, when my longest hard efforts were four and five minutes, eFTP ran 23-24W high. Through late spring and summer, when I was doing actual sustained work, the two converged and at one point eFTP read three watts below the ramp. Then September happened, I did one fast group ride full of 3-4 minute surges, and eFTP jumped 12W over a ramp test that had just shown me losing a bit of form.

The 40km TT Reality Check

Two time trials, both on the same dual carriageway course (a V718-style out-and-back, flat, roughly 200m of elevation across the whole thing), both on the same TT bike with the same Assioma Duo pedals I use for everything else.

11 May, 53:48, normalised power 306W, average 299W. My ramp FTP three days later was 314W. eFTP that week was 327W. A 54-minute effort at 306W NP puts real FTP somewhere very near 300-305W, because a well-executed 40 on a flat course is close to a 55-minute maximal effort and sits a couple of percent under true threshold for most riders at that duration.

23 August, 52:31, normalised power 318W, average 314W. Ramp test on 6 Aug said 322W. eFTP said 319W. Both were right. This is the block where I’d been doing 2×20s and over-unders religiously for eleven weeks, and my power duration curve had proper data between 10 and 40 minutes.

So the TTs backed the ramp test in May and backed both in August. What they never backed was a winter eFTP built entirely on Zwift race anaerobic efforts.

Why Ramp Tests Fail In A Different Direction

I’m not arguing the ramp test is the truth. It has its own well-known distortion: it rewards anaerobic capacity. The 0.75 multiplier on best minute is a population average, and if you’re a sprinter-ish rider with a big W’, a ramp test will flatter you. If you’re a diesel with a small W’ and a high aerobic ceiling, it will underrate you.

I’m the diesel type. My W’ sits around 21kJ, which is low for my weight (74kg). That’s why my ramp tests read conservative relative to the TTs. A rider with W’ of 28-30kJ would likely see the opposite: ramp test above both eFTP and real 40km pace.

Work out which you are before you trust either number. Take your best 1-minute power and your best 20-minute power from the last six months. If the ratio is above about 1.35, you’ve got meaningful anaerobic contribution and your ramp test is probably reading high. Mine ran 412W/318W = 1.30. Below 1.30 and the ramp is likely underselling you.

The Ride Types That Corrupt eFTP

Nine months of logging surfaced four repeat offenders. Each one produced an eFTP revision I could prove was wrong within two weeks.

Zwift races under 30 minutes. Six of my 41 revisions came from these. The starts are anaerobic, the efforts are 3-6 minutes, and the draft means your power spikes don’t correspond to sustained load. Every single one pushed eFTP up and every single one was contradicted by the next honest threshold session.

Short steep climbs done as repeats. A 4-minute hill done five times gives you a lovely 4-minute PB and nothing in the 15-40 minute window. eFTP extrapolates. It extrapolates generously.

Descending-heavy group rides. This one’s sneaky. If the hard efforts are separated by long freewheeling descents, you can produce a 10-minute peak that you could never hold as a continuous effort, because the model doesn’t know you got 90 seconds at 80W in the middle of it. intervals.icu works from the best rolling window, so a stochastic effort with gaps looks the same as a steady one.

Anything with a power meter dropout. I had a Kickr firmware issue in March that produced a 22-second spike to 780W. eFTP moved 9W. I had to manually flag the activity and rebuild the curve. Check your PD curve after any ride where the numbers look strange; intervals.icu lets you exclude an activity from your power curve from the activity’s options menu, and one bad file will sit in your 90-day curve for three months if you let it.

How To Read eFTP Properly

After nine months I’ve settled on three rules that make the number useful rather than misleading.

First rule: look at the trigger duration, always. intervals.icu shows it. If the effort that set your eFTP is under about 8 minutes, treat the number as an upper bound with a 15-25W error bar. If the trigger is 20 minutes or longer, treat it as roughly as good as a 20-minute test, because functionally that’s what it is.

Second: check your power duration curve for holes. On the Fitness page, the PD curve overlays your 42-day and 90-day bests. If the 12-30 minute section is visibly smooth and flat compared to the rest, you have no data there and eFTP is interpolating across a gap. My January curve had a 40W drop between 5 and 12 minutes with nothing in between. That’s an estimate built on nothing.

Third: use eFTP for direction, ramp or TT for level. eFTP’s genuine strength is that it updates continuously and catches gains between tests. Over my nine months, eFTP’s month-over-month change correlated well with what the ramp tests eventually confirmed, even when the absolute values were 20W apart. When eFTP started climbing in late March, that was real and the April ramp confirmed it. The trend was honest; the level wasn’t.

A Practical Calibration You Can Run This Week

If you want to know your own personal eFTP offset rather than borrowing mine, do this. It takes three weeks.

Week one, ride normally but include one 2×20 session at what you currently believe is threshold. Week two, do a ramp test on a Tuesday and a 30-minute all-out effort on the Saturday (outdoors on a consistent drag if you can, indoors if not). Week three, note your eFTP.

Now compare. Your 30-minute best × 0.95 gives you a solid FTP anchor. The ramp gives you a second. eFTP gives a third. Write all three down along with the eFTP trigger duration:

30-min best: 312W  →  anchor 296W
Ramp test:   288W
eFTP:        312W  (trigger: 4:11)
Offset: eFTP reading +16W vs anchor, +24W vs ramp

Keep that offset. Apply it whenever eFTP updates off a short effort. Recheck it every couple of months, because the offset moves as your training changes: mine went from +16W in winter to roughly zero by August, purely because my rides started containing the kind of efforts the model needs.

Where This Leaves The Two Methods

Neither number deserves to be your training zones on its own. The ramp test is repeatable and cheap and biased by your anaerobic profile. eFTP is free, continuous and biased by whatever you happened to ride hard last. The honest answer from nine months is that they agree when your training contains sustained threshold work, and diverge by 20W or more when it doesn’t, which is exactly the period when you’re least equipped to notice.

The thing I’d change about my own winter: I built sweet spot blocks off a 312W eFTP that should have been 290W. Eleven percent too hard on every interval, for six weeks, and I couldn’t work out why session three of the week kept falling apart. The estimate wasn’t broken. I just hadn’t given it anything to estimate from.