A Working ChatGPT Workflow For FIT File Analysis
Most people who try ChatGPT FIT file analysis give up in the first ten minutes. They drag a .fit file straight into the chat box, get told the file can’t be read, and conclude the whole idea is hype. That’s a tooling problem, not a capability problem. A .fit file is a binary format built by Garmin: profile-defined message types, scaled integers, developer fields, the lot. ChatGPT’s code interpreter can parse it, but only if you tell it how, and only if you’ve thought about what you actually want out of the other end.
This page is the workflow I use. It assumes you have a ChatGPT Plus or Pro subscription (the free tier’s file handling and code execution limits will fight you), a power meter, and a specific question about a specific ride. If you want the wider view of what LLMs can and can’t do with training data, the pillar at /ride-data-analysis/ covers the landscape, including intervals.icu’s own AI features and the local-model route. What follows is narrower: getting from a raw binary file to a defensible number.
Step Zero: Decide Whether You Need The FIT File At All
Before you upload anything, ask whether the FIT file is the right input. A lot of the analysis people attempt with a raw file is better done with a summary.
If your question is “how has my 20-minute power trended over eight weeks”, you want a CSV export from intervals.icu, not sixteen FIT files. If it’s “did I go out too hard in that 10-mile TT”, you need the second-by-second file, because the answer lives in the first 90 seconds of data.
Rough rule: FIT file analysis is worth the friction when the answer depends on within-ride structure. Pacing, cadence drift, torque effectiveness, cornering behaviour on a crit lap, the exact shape of a power fade. For anything that’s fundamentally a trend across rides, export a table and skip the binary parsing entirely.
Getting The File Out Of Garmin, Wahoo Or Zwift
Garmin Connect: open the activity on the web (not the app), click the gear icon top right, choose Export Original. You get a .fit, sometimes zipped. The app’s “export” gives you a TCX or GPX in some regions, which throws away developer fields and often rounds power to integers you didn’t record.
Wahoo users get the file from the ELEMNT app’s activity detail, or more reliably by plugging the head unit in and copying from the /Activities directory. Zwift writes to Documents/Zwift/Activities on Windows and ~/Documents/Zwift/Activities on macOS, with filenames like 2026-09-14-18-42-33.fit.
Strava is the awkward one. Individual activity downloads give you the original file only if you uploaded a FIT in the first place. If Strava recorded it from the phone app, you’ll get a GPX with no power data at all. Use the source device wherever possible.
One thing worth checking before upload: file size. A four-hour ride at 1 Hz is around 400 KB to 900 KB as a FIT, which uploads fine. A 1 Hz recording with Cycling Dynamics and Di2 shifting data from a Garmin Edge 1050 can reach 2.5 MB, and that’s still fine. Multi-day bikepacking files over 20 MB will time out during parsing.
The Upload Prompt That Actually Works
Here’s the thing nobody tells you: ChatGPT’s sandbox does not have fitparse or fitdecode installed, and it has no internet access to pip install them. So a bare “analyse this FIT file” fails at the import line.
The fix is to tell it to parse the binary directly. This prompt has worked consistently for me since mid-2025:
Attached is a Garmin .fit file from a road bike ride.
The sandbox has no fitparse/fitdecode and no internet, so write a
minimal FIT parser from scratch using only the standard library and
pandas/numpy. You need:
- Read the 12 or 14 byte file header, confirm ".FIT" at bytes 8-12
- Walk the record stream, handling definition messages (record
header bit 6 set) and data messages
- Handle compressed timestamp headers (bit 7 set)
- Decode global message 20 (record) fields: 253 timestamp,
0 position_lat, 1 position_long, 2 altitude, 3 heart_rate,
4 cadence, 5 distance, 6 speed, 7 power, 13 temperature,
73 enhanced_speed, 78 enhanced_altitude
- Apply the scale/offset from the FIT profile: altitude is
(raw/5)-500 m, speed is raw/1000 m/s, enhanced_altitude
is (raw/5)-500, distance is raw/100 m
- Timestamps are seconds since 1989-12-31 00:00:00 UTC
Output a DataFrame, print df.shape, df.head(10), and
df[['power','heart_rate','cadence']].describe(). Then stop and
wait for my next instruction.
That last sentence matters more than it looks. Without it, ChatGPT will parse the file and then immediately produce four pages of generic observations about your ride, burning context and giving you nothing. Making the first turn purely a parse-and-verify step means you can check the data is right before you build anything on top of it.
What you should see back looks roughly like this:
(11847, 12)
timestamp power heart_rate cadence distance altitude
0 2026-09-14 09:02:17 0.0 88.0 0.0 0.00 142.4
1 2026-09-14 09:02:18 61.0 89.0 52.0 2.41 142.4
2 2026-09-14 09:02:19 94.0 91.0 58.0 6.87 142.6
power heart_rate cadence
count 11847.0000 11847.0000 11847.000
mean 198.4419 141.2216 78.331
std 102.8873 22.4470 26.884
max 1104.0000 183.0000 142.000
Two sanity checks before you go further. First, does count roughly match ride duration in seconds? 11,847 records for a 3h17m ride means near-continuous 1 Hz recording with minimal smart-recording gaps. If you see 4,200 records for a four-hour ride, your head unit was on smart recording and every time-weighted calculation downstream will be wrong until you resample. Second, does max power look plausible? A 1104 W spike is a sprint. A 2400 W spike on a steady endurance ride is a dropout artefact and you should ask for it to be clipped.
Worked Example: Was That TT Pacing Actually Bad?
This is where the workflow earns its keep. Take a 25-mile TT, 56 minutes, average 291 W, with an FTP of 305 W. Strava says you rode 0.95 IF and gives you a Suffer Score. Neither tells you whether you paced it well.
Second turn of the conversation:
Using the parsed df, split the ride into 8 equal-time blocks.
For each block give me: mean power, normalised power, mean HR,
mean cadence, and mean speed. Format as a markdown table with
one decimal place. Then compute the coefficient of variation
of the 8 block mean powers.
What comes back on a badly paced TT:
| Block | Mean W | NP W | HR | Cad | Speed km/h |
|-------|--------|------|-----|-----|------------|
| 1 | 327.4 | 331.2| 158 | 91 | 41.8 |
| 2 | 304.1 | 306.8| 172 | 89 | 42.1 |
| 3 | 295.6 | 297.0| 175 | 88 | 41.3 |
| 4 | 289.2 | 291.4| 176 | 87 | 40.9 |
| 5 | 283.8 | 285.1| 176 | 86 | 40.2 |
| 6 | 277.3 | 279.9| 175 | 85 | 39.8 |
| 7 | 281.0 | 284.6| 177 | 86 | 40.1 |
| 8 | 298.7 | 305.3| 180 | 88 | 40.6 |
CV of block means: 5.3%
Now you have something to coach off. Block 1 at 327 W is 12% above the ride average and 107% of FTP, which for a 56-minute effort is a straightforward over-pace. The monotonic decline through block 6 is the cost. And block 8 at 298 W proves you had more in the tank than blocks 5 to 7 suggest, which is the classic signature of going too deep early and then rationing.
The useful follow-up prompt: Recompute assuming I had ridden the whole thing at the block-average power of 294.6 W. Using the standard cubic relationship between power and speed at these speeds, estimate the time difference. ChatGPT will produce an estimate in the 20 to 40 second range on a course like that. Treat it as an order of magnitude, not a stopwatch, because it’s ignoring wind, gradient and cornering.
The Numbers ChatGPT Gets Wrong, And How To Pin Them Down
This is the part where trust matters. Ask for “normalised power” without specifying and you will get one of three different calculations depending on how the model feels that day.
NP is a 30-second rolling average, raised to the fourth power, averaged, then fourth-rooted. The failure modes are specific: using a centred rolling window instead of trailing, not handling the first 30 seconds correctly, and including zero-power coasting in a way that inflates or deflates the result. Spell it out:
Compute NP as: 30-second trailing rolling mean of power
(min_periods=30, so the first 29 samples are dropped), each
value to the 4th power, mean of those, then the 4th root.
Include zeros. Print the value to one decimal place and also
print how many samples were dropped.
Same discipline for anything else that has a definition:
| Metric | What to specify | Common error |
|---|---|---|
| TSS | (duration_s × NP × IF) / (FTP × 3600) × 100 | Using average power instead of NP |
| Variability Index | NP ÷ average power, zeros included | Silently excluding coasting |
| Efficiency Factor | NP ÷ average HR over the steady portion only | Including the warm-up |
| Best 20 min power | Trailing rolling mean, not any 20-min block boundary | Binning into fixed windows |
| Decoupling | (Pw:HR first half) vs (Pw:HR second half), as % change | Not excluding the warm-up |
| Elevation gain | Sum of positive deltas on a 3-sample smoothed altitude series | Summing raw deltas, which inflates by 30 to 60% |
That last one catches almost everyone. Raw barometric altitude at 1 Hz has noise of ±0.3 m, and summing every positive delta across 12,000 samples turns a 780 m ride into 1,240 m. Always ask for smoothing, and always ask what smoothing was applied.
Verification habit worth building: once per new chat, ask ChatGPT to compute average power and total distance, then compare against what Garmin Connect shows for the same activity. If average power matches to within 1 W and distance to within 0.1 km, the parse is sound and you can trust the harder numbers. If they don’t match, the parser mishandled something (usually scaling on speed or a missed compressed-timestamp record) and everything downstream is suspect.
Which Model, And When It Matters
For the parsing turn, use the strongest reasoning model you have access to. Writing a bit-level FIT parser from memory is genuinely hard, and the cheaper models produce parsers that run without error but silently drop developer fields or misread the definition message architecture. GPT-5.1 Thinking and Claude Opus 5 both handle it. Faster, cheaper models will get there sometimes, which is worse than failing outright because you won’t notice.
Once the DataFrame exists and you’ve verified it against Garmin, model choice stops mattering much. Block splits, rolling averages and table formatting are ordinary pandas work. Switch to a faster model for the analysis turns if you’re doing a lot of them.
Where the reasoning model earns its cost again: interpretation across multiple constraints. “Given this 56-minute TT file, my 4-week CTL of 78, and the fact that I raced on Saturday, is Tuesday’s 5×5 at 110% FTP appropriate?” That’s a judgement question with real tradeoffs, and a thinking model will surface the tension between the decoupling number and the training plan rather than just agreeing with you.
Keeping One Chat Per Question
The temptation is to build a giant chat with twenty rides in it. Don’t. Once a conversation exceeds roughly 30 turns of code execution, the sandbox has usually been reset at least once, the DataFrame is gone, and ChatGPT will silently re-parse or, worse, hallucinate numbers from earlier in the thread rather than recomputing.
Practical structure: one chat per question, parse once, verify once, then three to six analysis turns. If you want the same analysis across many rides, get ChatGPT to write you a standalone Python script in that chat, then run it locally against a folder of FIT files with fitdecode properly installed. That’s the point at which you graduate from the chat interface to actual tooling, and the chat has served its purpose as a way of writing the script.
Where this workflow stops being enough is when you need the answer in under thirty seconds while standing over the bike after a session. For that, intervals.icu’s own charts and a saved custom field will beat any amount of prompting.