Fitting A Fitness-Fatigue Model To Yourself (And Why It’s Harder Than It Looks)
I spent a winter fitting the Banister impulse response model to my own ride files. Eleven months of data, a power meter on every ride, and eight maximal tests scattered through it. The model is old (Banister, Calvert and colleagues, 1975, originally fitted to swimmers and hammer throwers) and it sits underneath nearly every “fitness” line you have ever looked at on Strava, intervals.icu, TrainingPeaks or Zwift. If you want an impulse response model for cycling that tells you your FTP next Tuesday, I have bad news. If you want one that draws a useful shape, read on.
The model in five lines
Performance on day n is a baseline plus a fitness term minus a fatigue term:
p(n) = p0 + k1 * sum_{i<n} w(i) * exp(-(n-i)/tau1)
- k2 * sum_{i<n} w(i) * exp(-(n-i)/tau2)
w(i) is training load on day i (I used TSS). Five free parameters: p0, k1, k2, tau1, tau2. Textbook values are tau1 around 42 days and tau2 around 7, which is where CTL and ATL come from. Note the difference though: CTL and ATL in intervals.icu or TrainingPeaks fix the time constants at 42 and 7 and never fit anything to your tests. Banister fits all five to your own performance data. That is the whole promise, and the whole problem.
What I fitted it against
Eight performance tests: 20-minute maximal efforts, ridden on the same Wahoo-Kickr-and-Zwift setup, and converted to a 0.95 × 20-minute figure. Ranges from 268 W in a bad February to 301 W in late June. Daily TSS came from the intervals.icu export. Fitting was ordinary least squares via scipy.optimize.least_squares, with bounds of 10 to 90 days on tau1 and 1 to 20 on tau2.
Eight data points, five parameters. Already you should be wincing. The original Banister studies had the same issue, and the literature since (Hellard and others, 2006, put it politely) has flagged that you want dozens of tests, not a handful.
Run one: the fit looks great
Fit A (all 8 tests)
p0 = 247.1 W
k1 = 0.118 tau1 = 38.2 d
k2 = 0.341 tau2 = 6.4 d
R^2 = 0.91 RMSE = 3.6 W
Lovely. Plausible time constants, close to the textbook ones. I nearly stopped there, which is how most blog posts about this model end.
Run two: drop one test
Leave-one-out is the cheapest honesty check there is. Here are the results of refitting seven times with a different test removed each time:
Dropped test tau1 tau2 k1 k2 p0
(none) 38.2 6.4 0.118 0.341 247.1
Feb 11 52.7 9.1 0.071 0.198 251.8
Mar 25 31.0 5.2 0.164 0.502 243.0
May 06 44.9 7.8 0.093 0.262 249.6
Jun 24 27.5 4.1 0.221 0.688 239.4
Aug 12 61.3 11.0 0.052 0.134 254.0
Sep 30 35.8 6.9 0.129 0.360 246.2
tau1 wanders from 27.5 to 61.3 days. k2 varies more than fivefold. Removing one test out of eight, 12% of the data, changes the “personal recovery time constant” by a factor of two. That is not a parameter. That is a number the optimiser landed on because the error surface is a long flat valley, and the valley floor tilts depending on which points you give it.
The reason is structural. k1 and tau1 trade off against each other, as do k2 and tau2, because both exponentials are smoothed versions of the same load series. Training load is also autocorrelated: hard blocks follow hard blocks, easy weeks cluster. The fitness and fatigue signals are highly collinear, so the data cannot tell “bigger k1 with shorter tau1” from “smaller k1 with longer tau1”.
Run three: does it predict anything?
Fit on the first five tests (to end of May), predict the last three:
Test Actual Predicted Error
Jun 24 301 W 294 W -7
Aug 12 289 W 281 W -8
Sep 30 283 W 297 W +14
The first two are respectable, since the model correctly knows I was getting fitter in the summer. The third missed by 14 W, about 5%. My real FTP moved by maybe 18 W across the whole season. An error of 14 W is most of the signal. For comparison, simply guessing “same as the last test” would have missed by 12, 6 and 6 W. A model with five fitted parameters lost to a naive persistence forecast on one of three holdouts and was barely better on the others.
So what is the usable output?
The curve. Take any of those parameter sets from the leave-one-out table, including the wild ones, and plot the fitness-minus-fatigue line across the year. The shapes agree. Same build in March, same dip through the July holiday and the heavy August block, same rebound in September. Correlation between the best and worst parameter sets’ curves was 0.96.
Different parameters, same picture. The shape is robust because it is dominated by your load history, which is data you actually have, rather than by the fitted constants, which you do not. Rankings of “which fortnight was my freshest” or “when did I peak” came out identical across all seven fits, give or take a few days.
So the practical rules I now use:
- Read timing and direction, not watts. “Form was positive for about ten days before the June test” is trustworthy. “Your FTP is 296 W today” is not.
- Don’t trust personalised time constants. If a tool such as an LLM-driven coach or a Python notebook gives you “your tau1 is 34.7 days”, ask for the interval. If it cannot give one, ignore the decimal.
- Fixed 42/7 is fine. The default CTL/ATL in intervals.icu is no less valid than my fitted version, and it does not shift when you add a test.
- Test more often if you want to fit. Monthly tests for a year gives twelve points. Still thin. Fortnightly rides with a ramp or 20-minute effort start to get you somewhere, at the price of a lot of suffering.
A note on AI tools
If you paste your ride files into a chatbot and ask it to fit Banister, it will cheerfully return five parameters and an R². Mine did, in about forty seconds, with a confident paragraph about my “unusually long fitness decay of 51 days”. That was one of my leave-one-out fits. Ask it to rerun with one test dropped and watch the story change. A tool worth paying for will show you that spread unprompted. Most do not.
For the wider picture on how these models and FTP estimates fit together, the pillar page on FTP estimation and fitness models covers the rest.
Try it yourself
Export your activities from intervals.icu, pull your best 20-minute power for each month, and fit with bounded least squares. Then do the one thing most people skip: drop each test in turn and tabulate the parameters. If tau1 moves by more than 20% you have your answer about whether to trust the numbers. Plot the curve regardless. It will still be telling you something true.