AI Route Planning And Race Pacing
An ai cycling route planner is only worth your time if it does something your own local knowledge and a heatmap can’t. For most riders that threshold is crossed in exactly three places: finding rideable surface you’ve never seen, resolving a route against wind that hasn’t happened yet, and converting the resulting profile into a power plan you can actually hold. Everything else is a nicer-looking version of RideWithGPS.
This page is about the parts that survive contact with a FIT file. It assumes you can read a power curve, know roughly what your 20-minute number is, and have been burned at least once by a tool that promised 1,400 m of climbing and delivered 1,120 m.
What an AI cycling route planner is actually doing under the bonnet
There are two completely different things being sold as AI route planning, and conflating them is why people end up disappointed.
The first is classical graph routing with heuristics tuned on usage data. Komoot’s sport profiles, cycle.travel’s surface-aware routing, Strava’s global heatmap and route builder, RideWithGPS’s popularity weighting: these all run Dijkstra or A* variants over an OpenStreetMap-derived graph with per-way cost functions. Komoot weights surface, smoothness and highway tags differently for “Road Cycling” versus “Gravel Riding”. cycle.travel penalises A-roads hard and has genuinely good UK surface data because Richard Fairhurst has spent years on it. None of this is a language model. It’s good software, and for a 120 km Peak District loop it will beat an LLM every time.
The second is an LLM sitting on top of those engines, plus weather and terrain APIs, doing the bit that graph search can’t: interpreting a fuzzy brief. “Give me 90 km from Hathersage that avoids the A57, has at least two cafés open on a Tuesday, keeps me out of a north-easterly for the last 30 km, and tops out under 1,300 m of climbing.” That’s a constraint-satisfaction problem with soft preferences, and it’s exactly what a tool-using model is for.
The failure mode is tools that skip the engine. Ask a model to emit GPX directly and it will oblige, confidently, with coordinates it has invented.
Why an LLM cannot draw a route, and what to ask for instead
Here is the shape of what you get back when you ask for raw GPX. This is a real pattern, not a strawman: the model produces well-formed XML with plausible-looking decimal degrees that drift off the road network.
<trkpt lat="53.3272" lon="-1.6501"><ele>167</ele></trkpt> <!-- on the B6001, fine -->
<trkpt lat="53.3341" lon="-1.6388"><ele>212</ele></trkpt> <!-- 340 m east of any way -->
<trkpt lat="53.3410" lon="-1.6275"><ele>268</ele></trkpt> <!-- crosses the Derwent, no bridge -->
Three points, roughly 1.5 km apart, is about the density a model will produce for a 60 km loop: 40 or so trackpoints. Load that into a Garmin and it will draw straight lines between them, which means your Edge will happily route you across a field boundary and a river. The elevation values are invented too, which matters more than it sounds, because that’s the input to any pacing plan built downstream.
The fix is to force tool use. A well-built AI cycling route planner never emits geometry itself; it emits a request. Test any tool by asking it to show you the call it made. If it can’t, it’s guessing.
curl -X POST 'https://api.openrouteservice.org/v2/directions/cycling-road/geojson' \
-H "Authorization: $ORS_KEY" -H 'Content-Type: application/json' \
-d '{
"coordinates": [[-1.6501,53.3272],[-1.7812,53.3695],[-1.6501,53.3272]],
"elevation": true,
"preference": "recommended",
"options": {"avoid_features": ["ferries","steps"]},
"instructions": false
}'
OpenRouteService gives you cycling-road, cycling-gravel, cycling-mountain and cycling-electric. BRouter is the better choice for UK gravel because you can edit the profile text directly: bump the cost multiplier on surface=asphalt and watch it thread bridleways instead. GraphHopper and Valhalla both take custom cost models if you want to host your own. The point is that the coordinates come from a graph that knows where bridges are.
A good prompt makes this explicit:
Plan a 95 to 105 km gravel loop from Kielder Castle. Do not write GPX. Instead: (1) propose 4 to 6 waypoints as lon,lat pairs with a one-line reason for each; (2) call BRouter with the
Trekking-Dryprofile between them; (3) report total distance, ascent from the returned elevation array, and the surface breakdown by percentage; (4) if total distance is outside 95 to 105 km, adjust waypoints and re-call, up to three attempts. Show every call and its raw distance/ascent numbers.
That structure (brief, engine, critic) is the whole trick. The third stage is where most tools fall down: they accept the first result. Insisting on a self-check loop against a numeric target turns a 112 km “100 km loop” into a 99 km one.
Elevation data will make your plan wrong by 15% before you turn a pedal
Ascent figures disagree because the underlying digital terrain models disagree, and the smoothing algorithms disagree more.
| Source | Resolution | Typical behaviour on UK terrain |
|---|---|---|
| SRTM v3 | 30 m (1 arc-sec) | Under-reads short steep ramps, smooths hedgerow-scale relief |
| OS Terrain 50 | 50 m grid | Free, coarse, fine for 100 km totals, useless for a 400 m climb |
| EA LIDAR DTM (England) | 1 m / 2 m | Excellent, open under OGL, but includes drainage ditches that look like descents |
| Barometric altimeter (Edge 840, Wahoo Roam) | continuous | Drifts with pressure change, usually the most trustworthy for a single ride |
A concrete case. Holme Moss from the Woodhead side is about 4.7 km averaging just over 7%. Plan it in one tool and you get 1,180 m for a loop containing it; ride the same loop with a barometric head unit and you record 1,043 m. That’s a 13% gap, and it is not a bug in either system. Strava applies its own elevation basemap to non-barometric uploads, which is why two mates on the same ride post different climbing.
Why you care: if you hand an AI pacing model an ascent figure that’s 13% high, every climb-segment duration it returns is long, your projected finish time is slow, and your fuelling plan over-provisions. Ask the tool which DTM it used. If the answer is “the route platform’s number”, treat the profile as directionally useful and the totals as ±12%.
From GPX to watts: the physics a model will fudge
The steady-state cycling power equation is not controversial:
P_wheel = v · [ Crr·m·g·cosθ + m·g·sinθ + ½·ρ·CdA·(v + w)² ]
P_rider = P_wheel / η η ≈ 0.975 for a clean chain
Where v is ground speed (m/s), w is the head-on wind component (m/s, positive into you), ρ is air density (1.225 kg/m³ at sea level, 15 °C), and m is rider plus bike plus kit.
Worked example, because this is the number you should use to audit any tool. Take a 75 kg rider, 8 kg of bike and kit, so m = 83 kg. Road position CdA = 0.32 m². Crr = 0.005 on decent tarmac with 28 mm tyres at sensible pressure. No wind.
On the flat at 250 W: P_wheel = 243.75 W. Rolling resistance force is 0.005 × 83 × 9.81 = 4.07 N. The aero constant ½ρCdA = 0.196. Solve 243.75 = 4.07v + 0.196v³ and you get v = 10.11 m/s, which is 36.4 km/h.
On a 6% gradient at the same 250 W: gravity contributes 83 × 9.81 × 0.0599 = 48.8 N, rolling drops slightly to 4.07 N, so 243.75 = 52.85v + 0.196v³ gives v = 4.32 m/s, which is 15.5 km/h. A 1.2 km ramp at 6% therefore takes 278 seconds, 4:38.
Now go and ask whichever tool you’re evaluating for those two numbers, giving it exactly those inputs. Models that get 36.4 km/h and 15.5 km/h within a few tenths are doing the algebra (usually by writing and running code). Models that come back with 33 km/h flat and “about 5 minutes” for the climb are pattern-matching against forum posts. The difference compounds across a 40-segment course into minutes.
Two errors to watch for specifically. First, using tanθ where sinθ belongs: at 6% the difference is 0.2%, at 20% (Hardknott, Rosedale Chimney) it’s 2%, which is real. Second, holding ρ at 1.225 for a plan that spends an hour above 500 m, where 1.16 is closer, worth roughly 1.5% of speed at 45 km/h.
Wind is the single biggest lever in UK racing, and it’s computable
Nothing else in a British road or TT plan moves the numbers like wind. The good news: you can resolve it per-segment for free.
Open-Meteo needs no API key and returns hourly 10 m wind speed and direction:
https://api.open-meteo.com/v1/forecast
?latitude=52.7&longitude=-1.3
&hourly=wind_speed_10m,wind_direction_10m,wind_gusts_10m,temperature_2m,surface_pressure
&windspeed_unit=ms
The prompt that makes this useful is the one that asks for vector decomposition along your actual track:
From the attached GPX, compute the bearing of each 500 m segment. Using Open-Meteo hourly wind for the segment’s forecast timestamp, resolve wind into head-on and cross components:
w_head = wind_speed · cos(wind_from_bearing − segment_bearing). Reduce 10 m wind to rider height by a factor of 0.65 for open roads and 0.35 for hedged or wooded sections. Output a table of segment, bearing, w_head, and a target power, and flag every segment where the cross component exceeds 4 m/s.
MyWindsock already does a version of this for UK courses and Strava segments, and it is the reference point to check an LLM against. If your model’s wind-adjusted 10-mile prediction differs from MyWindsock’s by more than about 40 seconds, the model is wrong, not MyWindsock.
What 52 seconds of variable pacing looks like
Take a flat out-and-back 40 km, 20 km out on a bearing of 235°, 20 km back at 055°. Wind from 240° at 22 km/h (6.1 m/s), giving essentially a pure 6.08 m/s head component out and the same as a tailwind back. Same rider as above: 83 kg, CdA 0.32, Crr 0.005.
Ride it at a flat 290 W:
- Out: solve 283 = 4.07v + 0.196v(v + 6.08)² → v = 7.26 m/s, 26.1 km/h, 45:55
- Back: solve 283 = 4.07v + 0.196v(v − 6.08)² → v = 14.86 m/s, 53.5 km/h, 22:26
- Total 68:21, total work 1,189 kJ
Now split it 310 W into the wind, 265 W with it:
- Out at 310 W: v = 7.50 m/s, 27.0 km/h, 44:27
- Back at 265 W: v = 14.47 m/s, 52.1 km/h, 23:02
- Total 67:29, total work 1,193 kJ
Fifty-two seconds for 0.3% more mechanical work. Normalised power on the variable plan is 297 W against an average of 295 W, so VI = 1.007, which is inside anything a judge or a coach would call even pacing. That asymmetry (spend watts where air resistance is highest, save them where you’re already fast) is the single most reliable finding in the whole field, and it’s the core of what a competent tool should be doing for you. The full method, including how the numbers hold up when you take them to a real course with real corners and a real 180° turn, is on AI-generated time trial pacing plans.
One caveat worth more than most: a ±7% split around threshold is about the limit for most riders on a 25-mile TT. Ask an AI for a plan and it may hand you ±15%, which looks clever in a spreadsheet and blows you up at 14 km.
Normalised power, VI, and why 0.85 IF is meaningless on gravel
Intensity factor targets get quoted as though surface doesn’t exist. It does.
| Event type | Realistic VI | Sustainable IF for 3 to 4 h |
|---|---|---|
| Flat 25-mile TT | 1.01 to 1.04 | 0.92 to 0.95 (for ~55 min) |
| Rolling UK road race | 1.06 to 1.12 | 0.82 to 0.88 |
| Sportive, hilly (Struggle Moors, Fred Whitton) | 1.08 to 1.14 | 0.72 to 0.78 |
| Dirty Reiver 200 km gravel | 1.16 to 1.22 | 0.62 to 0.70 |
A Dirty Reiver file with VI 1.19 is not evidence of bad pacing. Forest gravel with 300 m of loose fire road followed by a 12% kicker forces variability regardless of discipline. What that means practically: any AI tool that gives you a gravel plan expressed as “hold 0.80 IF” has not looked at a gravel file. The useful instruction for long off-road events is a cap, not a target. Something like “no more than 4 minutes above 105% of FTP in any 30-minute window, and no single effort above 130% for longer than 40 seconds before km 150.”
Feed the model a target VI and ask it to reverse-engineer what average power that implies. With VI 1.19 and a target NP of 240 W, average power is 202 W, and your kJ total over 7 hours is around 5,090 kJ. That number is what drives fuelling, not NP.
Fuelling maths, done in the right order
Mechanical efficiency on a bike is roughly 22 to 24%, and one kcal is 4.184 kJ, which is the happy coincidence behind “kJ of work ≈ kcal burned”.
Worked: 4 hours at 235 W average. Work = 235 × 14,400 = 3,384 kJ. Metabolic cost at 23% efficiency = 3,384 / 0.23 = 14,713 kJ = 3,517 kcal. At IF 0.72, call it 60% of that from carbohydrate: 2,110 kcal, so 528 g of carbohydrate oxidised. Liver and muscle glycogen give you maybe 500 g total, and you’re not starting full and won’t finish empty.
Intake at a well-drilled 90 g/h over 4 hours is 360 g, which is 168 g short of oxidation. That’s fine for four hours. At seven hours and 5,090 kJ, the same arithmetic says you need to be at 90 g/h from hour one, not hour two, and the deficit is what determines whether km 160 is survivable.
Where AI gets this wrong: quoting 60 g/h as a universal ceiling (that’s the glucose-only limit; glucose plus fructose at 1:0.8 supports 90 to 120 g/h in trained guts), and computing carbohydrate need from kcal burned without splitting fat and carbohydrate by intensity. Ask for the substrate split explicitly.
Getting your own ride files in without wasting the context window
A 4-hour FIT recorded at 1 Hz is 14,400 records. Exported as CSV with eight columns it’s around 850 KB, which is roughly 250,000 tokens. You can technically fit that in a long-context model. You should not, because attention over 14,400 near-identical rows produces mush, and you’ll pay for it.
Pre-aggregate instead. Thirty-second means give you 480 rows and about 8,000 tokens, and every metric you care about (NP, VI, time-in-zone, climb splits) is preserved to within a percent.
import fitdecode, pandas as pd, numpy as np
rows = []
with fitdecode.FitReader("ride.fit") as f:
for frame in f:
if frame.frame_type == fitdecode.FIT_FRAME_DATA and frame.name == "record":
rows.append({fld.name: fld.value for fld in frame.fields})
df = pd.DataFrame(rows).set_index("timestamp")
agg = df.resample("30s").agg({
"power": "mean", "heart_rate": "mean", "cadence": "mean",
"speed": "mean", "altitude": "last", "distance": "last",
})
agg["grade"] = agg.altitude.diff() / agg.distance.diff() * 100
# 30s rolling mean of power, then fourth-root-mean = Coggan NP
r = df.power.rolling("30s").mean().dropna()
np_watts = (np.mean(r ** 4)) ** 0.25
print(f"NP {np_watts:.0f} W | avg {df.power.mean():.0f} W | VI {np_watts/df.power.mean():.3f}")
agg.round(1).to_csv("ride_30s.csv")
For intervals.icu, skip the export step. Basic auth with the literal username API_KEY and your key as the password:
curl -u "API_KEY:$ICU_KEY" \
'https://intervals.icu/api/v1/athlete/i12345/activities?oldest=2026-07-01&newest=2026-09-30'
curl -u "API_KEY:$ICU_KEY" \
'https://intervals.icu/api/v1/activity/i9876543/streams/time,watts,altitude,latlng'
Then the prompt that gets real answers rather than encouragement:
Attached is 90 days of activity summaries and the 30-second streams for four rides tagged “race”. Compute, showing your arithmetic: (a) my mean maximal power for 5 s, 1, 5, 20 and 60 min across the window, with the date each came from; (b) for each race file, the percentage of total work done above 105% of a 292 W FTP, and where in the ride it clustered; (c) whether my 20-minute power is higher on rides with a preceding day of under 90 TSS. If a calculation needs data I haven’t given you, say so and stop rather than estimating.
That last sentence is doing real work. Without it you get fabricated power-curve values that look right.
A test harness for judging any AI coaching tool
Five prompts, scored pass/fail. Run them against whatever you’re paying for.
- Physics audit. 83 kg, CdA 0.32, Crr 0.005, ρ 1.225, 250 W, 6% gradient. Pass = 15.4 to 15.7 km/h.
- Geometry honesty. Ask for a 60 km loop. Pass = it names a routing engine and shows waypoints; fail = inline GPX.
- Elevation provenance. “Which DTM produced this ascent figure, and what’s the uncertainty?” Pass = a named source and a range.
- Wind decomposition. Give it a 235° bearing and wind from 240° at 6.1 m/s. Pass = head component 6.05 to 6.1 m/s, plus a stated 10 m-to-rider-height reduction factor.
- Refusal test. Ask for your FTP from a ride file containing no power data. Pass = it says it can’t.
Most tools fail 2 and 5. The ones that pass all five are, in practice, the ones wired to execute code and call APIs rather than answer from memory.
Recon prompts that earn their tokens
The genuinely high-value use isn’t route drawing at all. It’s interrogating a course you can’t drive to.
Give the model a GPX plus Street View-derived notes and ask for a hazard and decision list: every junction where you’d lose more than 3 seconds, every corner whose radius forces you below 30 km/h, every gradient change longer than 200 m that needs a gear you might not have. On a 25-mile out-and-back, the roundabout turn is usually worth 8 to 14 seconds depending on whether you can take it wide, and knowing that in advance changes where you spend your matches.
Second high-value prompt: contingency branching. “Here are three route options from the same start. Wind forecast is 240° at 22 km/h now, forecast to back to 210° and rise to 30 km/h by 14:00. Rank them for a 13:00 start and tell me the last point on each where I can switch.” Graph routers can’t answer that. Models with tool access can, and the answer is checkable against the ride you then do.
Third: pacing plan diff. After the event, hand the model the plan and the actual file and ask for a segment-by-segment variance table with a column for “cause”, constrained to causes visible in the data (gradient error, wind error, cornering, rider deviation). Keep those diffs in a running log. Four of them and the pattern in your own errors becomes obvious, usually that you go too hard in the first 8 minutes and that your CdA estimate is optimistic by 0.015.
What to do next, before you plan another route: pull your last three race files, run the 30-second aggregation above, and check the number in each one that a model would have had to predict. If the tool you’re using can’t be wrong in a way you’d notice, it isn’t telling you anything.
In this section
The supporting pages under this subject.