Aerobic Decoupling: The One Metric Worth Asking An AI To Compute Every Endurance Ride
Most of the data you generate on a four-hour ride is noise dressed as insight. TSS tells you the ride happened. Normalised power tells you roughly how hard, if your FTP is current, which it usually isn’t. Your Strava fitness curve is a 42-day exponential average of a number derived from another number you last tested in February. None of it answers the question that actually matters for an endurance athlete: is your aerobic engine getting more durable, or are you just accumulating hours?
Aerobic decoupling answers that. It costs nothing, needs no test protocol, and it’s computed from a file you already have. The catch is that almost nobody computes it correctly, and if you hand the job to a chatbot without constraints you’ll get a confidently wrong number to three decimal places.
What Pw:HR actually measures
Joe Friel’s formulation is simple arithmetic. Split a steady aerobic effort into two equal halves. For each half, divide normalised power by average heart rate to get a watts-per-beat efficiency figure. Then express the change from the first half to the second as a percentage.
Pw:HR (first half) = NP_1 / HR_1
Pw:HR (second half) = NP_2 / HR_2
Decoupling % = (Pw:HR_2 - Pw:HR_1) / Pw:HR_1 x 100
A negative result means your heart rate climbed to hold the same watts. Friel’s working threshold is 5%: under that, you’re aerobically coupled for that duration and intensity. Above it, you’ve found the edge of your current durability.
Why is this higher signal than anything else on your dashboard? Because it’s a ratio of a ratio, it cancels out most of what makes other metrics lie. You don’t need a correct FTP, because you’re comparing you to you within a single ride. You don’t need a lab. You don’t need to trust your HR strap’s absolute calibration, only its consistency over three hours. And the failure it detects is precisely the one that decides road races, 200km gravel days and late-season 100-mile TTs: not what you can produce fresh, but what you can still produce at hour four.
The metric is also unusually sensitive to things you can actually fix. Underfuelling shows up as decoupling. Sleeping badly shows up as decoupling. Riding an endurance block too hard shows up as decoupling long before it shows up as a blown FTP test.
Why eyeballing it does not work, and why default LLM answers are worse
Here’s a real-ish scenario. Three hours in the Surrey lanes, average power 205 W, average heart rate 141 bpm. You glance at the graph, see heart rate finishing seven beats higher than it started, and conclude the ride was fine.
Compute it properly, with the warm-up clipped and coasting removed:
| Normalised power | Average HR | Pw:HR | |
|---|---|---|---|
| First half (0:20 to 1:40) | 212 W | 138 bpm | 1.536 |
| Second half (1:40 to 3:00) | 209 W | 145 bpm | 1.441 |
Decoupling: -6.2%. Not fine. That ride sits outside Friel’s band, and it was ridden at 68% of FTP, which is not an intensity where you should be falling apart.
Now the LLM problem. Paste a summary of that ride into a chat window and ask for decoupling, and a language model will produce a number. It will look reasonable. It will often be wrong, because the model did one or more of the following: used average power instead of normalised power, split the ride by elapsed time rather than moving time so a 12-minute café stop landed in the middle of the second half, included a 20-minute warm-up in which heart rate was still rising toward steady state, averaged heart rate including zeros from strap dropouts, or simply performed the arithmetic in its head and got it slightly off.
Worse, models are cheerful. Ask “was this a good ride?” and you’ll get validation. Ask for a specific number computed under specific rules, with code, and you get something you can act on.
The prompt that produces a defensible number
The fix is to stop treating the model as an analyst and start treating it as a calculator that writes its own script. Use a tool with real code execution: Claude with the analysis tool (Opus 5 for the interpretation, Sonnet 5 is fine for the arithmetic), or ChatGPT’s data analysis mode. Upload the actual .fit file or a per-second CSV, never a summary.
Attached is a per-second activity file. Compute aerobic decoupling (Friel Pw:HR).
Write and execute Python. Do not estimate any value by inspection.
Cleaning rules, apply in this order:
1. Drop the first 20 minutes (warm-up) and any auto-pause gaps.
2. Drop all samples where speed < 3 km/h or power == 0 for
more than 15 consecutive seconds (stops, long descents).
3. Treat heart_rate == 0 or < 60 or > 200 as missing. If more than
3% of samples in either half are missing, STOP and tell me the
file is unusable rather than reporting a number.
4. Smooth power with a 30s rolling average before computing NP.
Then:
- Split the remaining samples into two halves of equal MOVING time.
- Per half: normalised power (30s rolling avg, 4th power mean,
4th root) and mean heart rate.
- Report Pw:HR for each half to 3 dp and decoupling to 1 dp.
Also print: sample count per half, % samples dropped, mean and max
temperature per half, average gradient per half, and the elapsed
time of the midpoint split.
Finally: state whether the average intensity was below 80% of
FTP 268 W. If it was not, say the result is not a valid
durability read.
That last instruction matters more than it looks. Decoupling is only meaningful sub-threshold. A 25-mile TT will decouple by 8% or more and it tells you nothing except that you were riding a 25-mile TT.
The diagnostic printouts matter too. Here’s the second worked example, a gravel ride in the Chilterns where the raw file said the athlete was in perfect shape:
| NP | Mean HR (raw) | Mean HR (cleaned) | |
|---|---|---|---|
| First half | 212 W | 151 bpm | 141 bpm |
| Second half | 208 W | 148 bpm | 148 bpm |
Raw decoupling: +0.1%. Textbook durability. Except the chest strap was dry for the first eight minutes and reported 210 to 220 bpm before settling, which dragged the first-half mean up by ten beats. Clean those samples out and the real figure is -6.5%. The artifact didn’t just add noise, it inverted the conclusion. This is exactly the class of error a model that “reads the chart” will never catch and a model that prints its sample counts will surface immediately.
Reading the trend, not the ride
One decoupling figure is a data point. Six of them across a block is a training decision. Run the same prompt on every endurance ride above 90 minutes, log the output, and you get something like this:
Date Dur NP IF Temp Decoupling
2026-08-04 3:02 204 0.71 19C -6.2%
2026-08-11 3:10 208 0.72 21C -5.4%
2026-08-18 3:05 211 0.73 18C -4.1%
2026-08-25 4:15 201 0.69 17C -5.9%
2026-09-01 3:08 215 0.74 16C -3.1%
2026-09-08 4:30 206 0.70 18C -3.6%
Read that as a story: durability held while duration went up, and the 4h30 on 8 September decoupling less than the 3h02 five weeks earlier is the whole point of the block. Compare it to a fitness chart, where the same progress would show as CTL going from 71 to 79, a number that would also have gone up if you’d done the same hours badly.
The one thing you must not do is compare across conditions. A 28°C July ride will decouple more than a 9°C March ride at identical power, because cardiac drift from thermoregulation is real and has nothing to do with your aerobic base. Same for caffeine, altitude, and a tucked TT position that restricts stroke volume. Keep the comparisons like-for-like, which in practice means one designated route ridden roughly monthly at the same target intensity.
Getting the file into the model
intervals.icu is the path of least resistance for UK riders already using it. The activity page has a Decoupling field, and it computes per-interval too, which is genuinely useful if you want to see whether the last of five 20-minute sweet-spot efforts fell apart. Download the original file from the activity page, or pull streams through the API if you want to automate the whole block at once. Zwift writes .fit files to Documents\Zwift\Activities. Strava’s “Export Original” gives you the file your head unit recorded, which is what you want, not the Strava-processed version.
For parsing, fitdecode or fitparse in Python handle Garmin and Wahoo files without drama, and any code-execution model will install them without being asked. If you want the broader toolkit for turning ride files into answers, including the parsing boilerplate and the other questions worth asking of a season’s data, we cover that in analysing ride data with LLMs.
What to do with a bad number
Suppose the answer comes back at -7.8% on a three-hour ride at 0.70 IF. Resist the urge to reduce it to “I’m unfit.” Work the list.
Fuelling first, because it’s the most common cause and the cheapest to fix. Under 60g of carbohydrate per hour on a three-hour endurance ride will produce decoupling that looks like detraining. Check the second half against the first for a power fade you didn’t notice: if NP dropped from 212 to 195 while heart rate held, that’s a different problem (you ran out of matches) than heart rate climbing at stable power.
Then look at accumulated load. Decoupling that worsens across three consecutive weeks of a block, with everything else held constant, is the clearest early signal of non-functional overreaching available to a self-coached rider, and it shows up seven to ten days before the subjective “legs feel heavy” phase.
Duration is the third lever. If you decouple by 6% at three hours but 3% at two, your durability ceiling is somewhere around 2h30 and your endurance rides should be sitting just under it, not an hour past it. Build from there, adding fifteen minutes when the number comes back under 5%.
The riders who get value out of AI coaching tools are not the ones asking for a training plan. They’re the ones who’ve picked a small number of questions that have correct answers, written a prompt strict enough to force those answers, and run it on every file. Aerobic decoupling analysis is the best first candidate because the arithmetic is trivial, the cleaning rules are where all the difficulty hides, and cleaning is exactly the part a model with a Python interpreter does better than you do at 9pm on a Sunday with a spreadsheet open.