Parsing FIT Files Locally With Python: Streams, Records And The Gotchas
Every number you look at after a ride is someone’s opinion. Strava’s Weighted Average Power, intervals.icu’s Normalized Power, Garmin Connect’s Training Effect, Zwift’s “avg power” on a ride where you stopped to answer the door: each of those is a derived metric computed by code you can’t read, on a file you already own, using assumptions nobody published. Most of the time the differences are small enough to ignore. The moment you start asking an AI coaching tool to explain a training block, they stop being small, because the model is reasoning about whichever platform’s numbers you happened to paste in.
The fix is unglamorous. Read the FIT file yourself. A 3-hour gravel ride at 1 Hz is roughly 10,000 record messages and about 380 KB on disk, and you can have every field of it in a pandas DataFrame in under a second. Once you’re doing that, the platforms become renderers, not sources of truth, and you can check any claim an LLM makes about your training against the actual bytes your head unit wrote.
The 40-line version
I use fitdecode rather than python-fitparse or Garmin’s official garmin-fit-sdk. All three work. fitdecode wins because it exposes the file as a stream of frames instead of hiding the structure behind a flat message list, which matters enormously for developer fields and for files your head unit truncated mid-crash.
import gzip
import fitdecode
MOVING_STOP = {"stop", "stop_all", "stop_disable", "stop_disable_all"}
def frames(path):
opener = gzip.open if str(path).endswith(".gz") else open
with opener(path, "rb") as fh:
# Truncated files from a head unit that died mid-ride still parse
# right up to the cut if you don't insist on a valid CRC.
with fitdecode.FitReader(fh, check_crc=fitdecode.CrcCheck.WARN) as fit:
for frame in fit:
yield frame
def records(path):
for frame in frames(path):
if frame.frame_type != fitdecode.FIT_FRAME_DATA:
continue
if frame.name != "record":
continue
yield {f.name: f.value for f in frame.fields if f.value is not None}
That dict comprehension over frame.fields is the whole trick. You are not asking for a hardcoded list of fields you expect, you’re taking whatever the device wrote. Native fields, expanded component fields, developer fields, the lot. Strava’s bulk export hands you .fit.gz, hence the gzip branch; you’ll want it the first time you process a year of archived rides.
Point it at a file and you get rows like this:
{'timestamp': datetime.datetime(2026, 9, 14, 8, 41, 3, tzinfo=datetime.timezone.utc),
'position_lat': 613244918, 'position_long': -23004411, 'distance': 18423.41,
'enhanced_altitude': 96.2, 'altitude': 96.2, 'power': 247, 'cadence': 88,
'heart_rate': 148, 'enhanced_speed': 8.472, 'speed': 8.472,
'temperature': 17, 'left_right_balance': 51, 'left_torque_effectiveness': 84.5}
Note what fitdecode has already done for you: altitude came out of the file as a 16-bit integer with scale 5 and offset 500 and arrived as metres, speed was millimetres per second and arrived as m/s, and timestamp was seconds since the FIT epoch (1989-12-31 00:00:00 UTC, an offset of 631,065,600 from the Unix epoch) and arrived as a timezone-aware datetime. What it has not done is convert position_lat and position_long out of semicircles. Multiply by 180 / 2**31 (8.3819e-08) to get degrees, or pass processor=fitdecode.StandardUnitsDataProcessor() to FitReader and let it do that job. Be deliberate about which you choose: that processor also converts speed to km/h, and quietly running your m/s maths on km/h values is a 3.6x error that looks almost plausible on a gravel ride.
Frames, and why you should care about the ones that aren’t records
A FIT file is self-describing. Before any data message can be read, the file emits a definition message declaring which fields, in which order, at which byte widths, that message type will use for the rest of the file (or until redefined). fitdecode surfaces four frame types: FIT_FRAME_HEADER, FIT_FRAME_DEFINITION, FIT_FRAME_DATA and FIT_FRAME_CRC.
Two consequences worth internalising. First, a single .fit on disk can contain several chained FIT files, one after another, each with its own header and CRC; Garmin and Wahoo both do this in some firmware paths. If you see more than one FIT_FRAME_HEADER frame, you’re looking at a chained file, and naively concatenating everything can interleave two activities. Second, a definition message can redefine a message type mid-file, which is exactly what happens when a sensor drops and reconnects. Code that caches “the record message has these 12 fields” from the first definition it sees will silently lose data forty minutes in.
The data messages that aren’t record are where the real arguments live. session (global message 18) carries total_elapsed_time, total_timer_time, total_distance, normalized_power, training_stress_score. lap (19) carries the same per lap. event (21) carries the timer starts and stops. device_info (23) tells you which power meter, which firmware, and which battery status, which is how you prove a dropout was the left pedal and not your legs.
Developer fields: the data nobody else will show you
This is the part that justifies local parsing on its own. The FIT spec lets any application define its own fields, and those definitions travel inside the file: a developer_data_id message (207) registering the application UUID, then one field_description message (206) per custom field. A Moxy or Train.Red sensor writes SmO2 and THb. A CORE sensor writes core_temperature and skin_temperature. Xert writes its own strain fields. Stryd writes form power and leg spring stiffness.
Strava will show you none of it. Garmin Connect will show you some of it, badly. Your parser can show you all of it, with correct units, because the file tells you the units:
def developer_field_map(path):
"""(developer_data_index, field_definition_number) -> {name, units}"""
out = {}
for frame in frames(path):
if frame.frame_type != fitdecode.FIT_FRAME_DATA:
continue
if frame.name != "field_description":
continue
name = frame.get_value("field_name")
units = frame.get_value("units") if frame.has_field("units") else None
# Both are FIT string arrays; some writers emit a list, some a str.
norm = lambda v: "".join(v) if isinstance(v, (list, tuple)) else v
key = (frame.get_value("developer_data_index"),
frame.get_value("field_definition_number"))
out[key] = {"name": norm(name), "units": norm(units)}
return out
Run that over a ride recorded with a muscle oxygen sensor and you get something like:
(0, 0) {'name': 'SmO2', 'units': 'percent'}
(0, 1) {'name': 'THb', 'units': 'g/dl'}
(1, 0) {'name': 'core_temperature', 'units': 'C'}
Because the {f.name: f.value} comprehension above keys on the resolved name, those fields land in your DataFrame as SmO2 and core_temperature without any extra work. You now have a per-second muscle oxygenation trace you can align against power and interval boundaries, which is a thing you literally cannot do inside any mainstream platform’s UI. If you’re building this into something you’ll reuse, the build-your-own-tools pillar covers how to structure it as a small library rather than a growing pile of one-off scripts.
Three durations, one ride
Open any file and you will find at least three different answers to “how long was that ride”, and they disagree for good reasons.
total_elapsed_time on the session message is wall clock from first to last record. total_timer_time is elapsed minus whatever the timer was stopped for. And the span you get from df.index[-1] - df.index[0] is neither, because it depends on whether your device kept writing records while paused. Garmin Edge units with Auto Pause generally stop writing record messages, leaving a hole in the timestamps. Some Wahoo and Hammerhead configurations keep writing records at zero speed. Same ride, same rider, different files.
Reconstruct the moving windows from the timer events rather than guessing:
def moving_intervals(path):
"""List of (start, end) datetimes when the timer was running."""
spans, open_at = [], None
for frame in frames(path):
if frame.frame_type != fitdecode.FIT_FRAME_DATA or frame.name != "event":
continue
if frame.get_value("event") != "timer":
continue
etype = frame.get_value("event_type")
ts = frame.get_value("timestamp")
if etype == "start":
open_at = ts
elif etype in MOVING_STOP and open_at is not None:
spans.append((open_at, ts))
open_at = None
return spans
On the file I’ve been using as a test case, a 2h48 gravel ride with two café stops, this matters more than it sounds. Elapsed was 3h07m14s, timer time 2h48m09s, and the record timestamps spanned 3h07m11s with a 14-minute hole and a 5-minute hole in the middle. Normalized Power computed three ways came out at 228 W (every timestamp on a 1 Hz grid, gaps filled with zero), 241 W (moving windows only, concatenated), and 239 W (moving windows only, but with the 30-second rolling window allowed to straddle the join). Strava’s Weighted Average Power for the same file read 236 W. None of those is wrong. They’re four different definitions, and a 13 W spread is easily the difference between “that was a tempo ride” and “that was threshold work”.
Here’s the moving-time version, which is the one I’d defend:
import pandas as pd
def normalized_power(df, spans):
chunks = [df.loc[a:b, "power"].resample("1s").mean().ffill(limit=3)
for a, b in spans]
out = []
for c in chunks:
if len(c) >= 30:
out.append(c.rolling(30, min_periods=30).mean())
if not out:
return None
rolled = pd.concat(out).dropna()
return float((rolled.pow(4).mean()) ** 0.25)
resample("1s") is doing real work there. Garmin’s Smart Recording writes samples on change detection, not on a clock, and on steady flat riding it will happily stretch to one sample every 4 to 8 seconds. A rolling(30) over raw samples is then a rolling 30-sample window covering four minutes, and your NP comes out absurdly low. Always put the series on a real time grid first.
Zeros that mean coasting, zeros that mean nothing
A missing field and a zero field are different facts, and platforms flatten them into each other. When an ANT+ or BLE power meter drops out, most head units write a record message with no power field at all, so frame.has_field("power") is False and the dict comprehension above simply omits the key. Pandas turns that into NaN, which is correct. When you coast downhill, the meter reports 0 W, which is also correct and should be averaged in.
Telling them apart is a three-line heuristic that works surprisingly well:
df["power_missing"] = df["power"].isna()
suspect = df["power_missing"] & (df["speed"].fillna(0) > 2.0) & (df["cadence"].fillna(0) > 40)
print(f"{int(suspect.sum())} s of pedalling with no power reading "
f"({suspect.sum() / len(df):.1%} of the file)")
On a Favero Assioma file recorded next to a phone in a jersey pocket, I’ve seen that print 214 seconds, 2.1% of the file. Two percent sounds like nothing. Those 214 seconds were clustered in two dropouts of roughly 90 and 120 seconds, both during a 5-minute VO2 effort, and the platform’s average power for that interval was computed as if the missing seconds didn’t exist. The interval looked 9 W easier than it was.
device_info messages confirm the diagnosis. Filter for them, look at battery_status and manufacturer, and a dropout that coincides with battery_status == 'low' is a maintenance job, not a data-cleaning job.
Handing clean data to an AI coach
Once you have a DataFrame and a set of moving intervals, the useful AI workflow changes shape. Pasting 10,000 rows of CSV into Claude or ChatGPT and asking for Normalized Power is the worst version of this: long-context arithmetic over thousands of rows is exactly where models produce confident, plausible, wrong numbers, and you have no way to audit the result.
Two patterns that hold up. Ask the model to write the analysis code, then run it yourself against the file, so the logic is reviewable and the arithmetic is Python’s. Or compute the metrics locally and hand over a compact, labelled summary, maybe 40 lines of JSON per ride, with explicit definitions attached:
Here is a ride summary computed locally from the FIT file with fitdecode.
Definitions: np_moving = 4th-root-of-mean-4th-power of a 30s rolling mean of
power, computed only over timer-running windows, resampled to 1 Hz.
power_missing_s = seconds with cadence > 40 and no power field present.
{"duration_elapsed_s": 11234, "duration_timer_s": 10089, "np_moving": 241,
"avg_power_moving": 213, "ftp": 285, "power_missing_s": 214,
"missing_clusters": [[3410, 3500], [3980, 4100]], ...}
Two of my intervals overlap a power dropout cluster. Treat those intervals as
unreliable and say so explicitly rather than estimating through the gap.
Quote the fields you used by name.
The constraint at the end is the part people skip. A model told to quote the fields it used will visibly fail when it doesn’t have the data, instead of interpolating a number that matches the shape of your question. The models are good at the coaching read: pattern across weeks, whether your intervals are decoupling, whether your easy rides are actually easy. They’re bad at being a calculator over a wall of text. Give them the calculator’s output.
Try this tonight on your last three files: print total_elapsed_time, total_timer_time, and the span of your record timestamps side by side, then compare all three against the moving time your platform shows. Whichever of those four numbers is the odd one out tells you precisely which layer has been quietly editing your training history.