AI Cycling
§5.1 AI Route Planning And Race Pacing 1,938 words · 9 min

AI-Generated Time Trial Pacing Plans, Tested On Real Courses

Ask ChatGPT for a pacing plan for a 25-mile time trial and it will give you one. It will look plausible. It will have zones, watts, a warm-up, and some encouraging language about “settling into rhythm.” Then you ride the course and finish 90 seconds down on your own seat-of-the-pants effort, having blown up on the drag at mile 18 that the model never knew existed.

This page is about the gap between those two things. I’ve run the same three courses through five different AI tools across a 2025 season, compared the outputs against my actual power files, and worked out where an ai pacing plan for time trial efforts genuinely beats what you’d do yourself, and where it’s confidently wrong in ways that cost you real time.

What a pacing plan has to get right

Before testing anything, it’s worth being precise about the job. A time trial pacing plan is a mapping from course position to target power. That’s it. Everything else is packaging.

The physics constraint is that aerodynamic drag scales with velocity cubed, so power spent at high speed buys less time than power spent at low speed. The optimal strategy is variable pacing: push above threshold on climbs and into headwinds, drop below on descents and tailwinds. The theoretical gain is real but modest. Published modelling (Swain’s work in the early 90s, and the later Atkinson and Brunskill studies) puts it around 1-3% of total time for typical UK courses, which on a 55-minute 25 is 30 to 100 seconds. Worth having. Not worth blowing up for.

The constraint that actually matters more is the anaerobic work capacity budget. Every watt above your critical power draws down a finite tank of roughly 15-25 kJ for most trained riders. Spend it in the first 5km and you’ve got nothing for the final drag. Any AI plan that hands you surges without accounting for the recovery cost between them is arithmetic dressed up as coaching.

So the test for any tool is: does it know the course, does it know your physiology, and does it respect the work capacity budget? Three questions. Most tools fail at least one.

The three test courses

I picked these because they break AI pacing plans in three different ways.

Course A: a classic dual carriageway 10 (Q10/19-style, flat, out-and-back). Almost no gradient. Whatever advantage exists here comes from wind, not terrain. Total elevation gain 42m over 16.1km.

Course B: a sporting 25 in the Cotswolds. 310m of climbing, three distinct climbs of 2-4 minutes each, and a technical descent with two junctions where you have to sit up. This is where work-capacity budgeting gets tested.

Course C: a 90km gravel event in the Peak District. Rolling, with surface changes that make power-to-speed modelling unreliable. Included specifically to see which tools admit they don’t know rather than inventing numbers.

Rider profile used throughout, so you can scale to your own: FTP 285W, critical power 272W, W’ 19.4 kJ, mass 74kg, bike + kit 9kg, CdA measured at 0.229 m² in TT position via the aerolab function in Golden Cheetah. That measured CdA matters enormously and almost no AI tool asks for it.

Running the same brief through five tools

Identical prompt each time, including the rider profile above, the course GPX summarised as a gradient table, and forecast wind. Here’s what came back.

ChatGPT (GPT-5, no plugins). On Course A it produced a sensible flat-course plan: 268W steady with permission to go to 275W in the final 2km. Defensible. On Course B it fell apart. It suggested 320W on all three climbs and 230W on descents, with no note on how much W’ that spends. I modelled it afterwards in Golden Cheetah’s W’bal: the plan takes W’bal to 2.1 kJ remaining by the top of the second climb, meaning you arrive at the third climb with essentially nothing. That’s not a plan, it’s a detonation schedule.

Claude (Opus 4.5 and later Opus 5). Better on the budgeting question, mostly because it asked about W’ when I gave it partial data. Given the full profile, its Course B plan came out at 302W on climb one, 298W on climb two, 310W on the final climb, and a specific instruction to accept the speed loss on the descent rather than pedalling hard into the junction. W’bal at the summit of climb three: 6.8 kJ. Survivable. When I ran Course C, it flagged that it couldn’t estimate rolling resistance across the surface changes and asked for split times from a previous attempt instead of guessing. That refusal to invent is the single most valuable behaviour in this whole test.

intervals.icu (the Power Curve and Pace Plan features). Not a chat model, but it’s the tool most self-coached UK riders already have open. Its strength is that the numbers come from your actual files. The weakness is that pace planning is basic: it will hold you to a target normalised power but won’t optimise the distribution across gradient. Excellent for the physiology half of the problem, silent on the course half.

BestBikeSplit. The specialist, and the one that keeps winning. It’s not a generative model, it’s a physics engine, and that’s the point. Given the same inputs it produced this for Course B:

Segment      Dist    Grad    Target   Est.Speed   Cumulative
1  flat      3.2km   +0.4%   271W     41.2 km/h   4:39
2  climb 1   1.8km   +4.1%   301W     20.8 km/h   9:50
3  descent   2.4km   -3.8%   198W     54.1 km/h   12:29
4  rolling   6.1km   +0.9%   276W     38.6 km/h   21:57
5  climb 2   2.1km   +3.6%   297W     22.4 km/h   27:35
6  flat      9.8km   +0.2%   274W     40.9 km/h   41:52
7  climb 3   1.4km   +5.2%   309W     18.1 km/h   46:36
8  run-in    13.4km  -0.3%   279W     42.7 km/h   55:26

Predicted 55:26. Actual on the day, riding to that plan with a head unit showing the target: 55:51. Twenty-five seconds off across 55 minutes is 0.75% error, which is better than most riders’ day-to-day variance.

TrainerRoad’s Red Light Green Light and Zwift’s workout mode. Neither does course-specific TT pacing. Worth stating plainly so you don’t waste an afternoon trying.

Where the chat models actually earn their place

The physics engines beat the LLMs on number generation. That’s settled. But there are three jobs the chat models do better, and they’re not trivial.

Translating a plan into something you can execute is the first. BestBikeSplit gives you a segment table; it doesn’t tell you what to do when your legs feel terrible at 12km. Feeding the BBS output into Claude or ChatGPT with “give me three decision rules for when to deviate from this” produced genuinely useful output: if heart rate is more than 6bpm above the same point last time, hold target rather than chasing; if you’ve overshot on climb one, subtract 8W from the climb two target, not from the flat sections.

Debugging a bad ride is the second. Upload the .fit file summary and ask why the plan failed. On my second attempt at Course B I ran 4W over target on the flat sections, which felt like nothing, and the model correctly identified that the 4W cost roughly 900 joules of W’ across 19km of flat riding, which is why climb three hurt. I would not have found that by staring at the graph.

Third, and most underrated: the models are good at interrogating your assumptions before you commit. Ask “what would make this plan wrong?” and you get a list worth reading. Wind forecast resolution, whether your CdA holds when you’re fatigued (it doesn’t; most riders lose 3-6% of position quality in the final third), whether the course has a start ramp, whether your power meter has drifted since the last calibration.

A prompt that produces usable output

Vague prompts produce vague plans. This structure has been the most reliable across both Claude and GPT-5:

Rider: FTP 285W, CP 272W, W' 19.4kJ, 74kg, CdA 0.229, bike+kit 9kg.
Event: [distance], [date], start time [time].
Course profile by segment:
  km 0.0-3.2, +0.4% avg
  km 3.2-5.0, +4.1% avg
  [...]
Wind forecast: 14km/h from 240deg.
Previous best on this course: 57:12 at 268W NP.

Produce a segment-by-segment power target table. For each segment
give the target watts and the running W'bal in kJ. State explicitly
which segment is the highest-risk for going too deep. Do not give me
zone names. If you cannot estimate something, say so rather than
approximating.

The W’bal instruction is doing most of the work there. It forces the model to carry state across segments rather than treating each one independently, which is exactly the failure mode in that first ChatGPT Course B output. The “do not give me zone names” line kills the padding. The “say so rather than approximating” line is what got Claude to refuse on the gravel course.

One warning: models will still occasionally produce W’bal numbers that are internally inconsistent. Check them. The arithmetic is simple enough to verify with the Skiba differential equation, or just paste the table back into intervals.icu as a planned workout and look at what it says.

The gravel problem, and why nothing solved it

Course C defeated everything. BestBikeSplit’s physics engine assumes a rolling resistance coefficient you supply once, and on a course that moves between tarmac, hardpack, loose limestone and a 400m hike-a-bike, one Crr is a fiction. Predicted time came out at 3:11:04 against an actual 3:34:18, a 12% error that makes the segment targets meaningless.

The chat models were more honest but not more useful. Claude asked for previous split data and, when given it, produced something closer to a fuelling and effort-distribution plan than a power plan: hold 240-250W on the smooth sections, accept whatever power the rough sections demand, don’t chase a number on the climbs at 65km because the surface will decide the outcome. That’s correct advice and it’s also nearly content-free as a plan.

If you race gravel, the honest position is that power-based pacing plans work on the tarmac portions and nowhere else. Build the plan for the road sections, use RPE and heart rate ceilings elsewhere, and treat any tool that gives you confident watt targets for a rocky descent as a tool that doesn’t know what it’s talking about. The broader question of how AI handles route selection and mixed-surface events is covered in the AI Route Planning And Race Pacing pillar, which is the better starting point if you’re deciding what to adopt across your whole season rather than for one event.

What to actually do next Tuesday

Measure your CdA before you trust any plan. Everything downstream is wrong if this number is guessed. An aerolab estimate from a quiet out-and-back with a known elevation profile costs you one evening and improves every prediction you’ll make for the next two years. Most riders assume something around 0.25 and are actually at 0.27 or worse.

Get a real W’ figure too. intervals.icu will fit CP and W’ from your existing files if you’ve got a 3-5 minute effort and a 12-20 minute effort in the last 90 days. If you haven’t, go and do them; the model quality depends entirely on having both ends of the curve.

Then run the physics engine for the numbers and the language model for the execution rules, and check the second against the first. The division of labour is the whole trick: one tool that knows Newtonian mechanics, one tool that knows how to talk to a tired person at 40km/h, and a rider who knows that neither of them has ridden the course.