Your First Strava API Project: Pulling A Year Of Rides Into A DataFrame
Every AI coaching tool you’re going to evaluate in this series has the same weakness: it sits between you and your data. You paste a question, it answers from whatever slice it can see, and you have no way to check the arithmetic. The fix is unglamorous. Get your own activity history onto your own disk, in a format pandas can open in 40 milliseconds, and suddenly every claim a model makes about your training is falsifiable in one line of code.
That’s the whole argument for spending an evening on strava api python training analysis plumbing. Not because Strava’s analytics are bad, but because a local Parquet file is the thing you can join, filter, resample and hand to a model in whatever shape the question demands. intervals.icu will already do a lot of this for you, and it’s excellent, but it’s still someone else’s schema and someone else’s uptime. Owning the raw JSON means you’re never blocked.
There’s a practical urgency too. Strava tightened its developer terms in late 2024, restricting how third-party apps can use athlete data (including for training machine-learning models) and limiting what they can display about other athletes. Read that as a signal: the pipe from Strava to clever third-party tools is narrower than it was in 2020, and the tools you liked may lose access. Your own copy, pulled under your own API application for your own activities, is the durable version.
One-time setup: get an app and a refresh token
Go to strava.com/settings/api and create an application. You need three things from that page: the Client ID, the Client Secret, and the Authorization Callback Domain, which you should set to localhost. That’s it. No approval process, no review queue.
Now authorise yourself. Paste this into a browser with your client ID substituted:
https://www.strava.com/oauth/authorize?client_id=YOUR_ID&response_type=code&redirect_uri=http://localhost/exchange_token&approval_prompt=force&scope=activity:read_all
The scope matters more than anything else on this page. activity:read silently omits activities you’ve marked private or hidden from your feed, which for most UK riders means your commutes and any ride starting from your front door with a privacy zone. activity:read_all gets everything. If you get this wrong you won’t see an error, you’ll just quietly have a 15% hole in your dataset.
Your browser will fail to load localhost, which is fine. Copy the code=... value out of the address bar and exchange it once:
import requests, json, os
from pathlib import Path
r = requests.post("https://www.strava.com/oauth/token", data={
"client_id": os.environ["STRAVA_CLIENT_ID"],
"client_secret": os.environ["STRAVA_CLIENT_SECRET"],
"code": "PASTE_THE_CODE_HERE",
"grant_type": "authorization_code",
}, timeout=30)
r.raise_for_status()
tokens = r.json()
path = Path.home() / ".strava" / "tokens.json"
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(json.dumps(tokens, indent=2))
print(tokens["expires_at"], tokens["scope"])
You now have a refresh_token that lasts indefinitely and an access_token that dies in six hours. The refresh dance is where most first attempts break, so it’s worth handling properly rather than re-pasting codes every afternoon.
The refresh handler that actually holds up
Two things trip people here. First, Strava may return a new refresh token in the refresh response, and if you don’t persist it your script works fine for weeks and then hard-fails. Second, checking expires_at > time.time() with no buffer gives you a 401 roughly once every few hundred runs when the token expires mid-pagination.
import time, json, os
from pathlib import Path
import requests
TOKEN_URL = "https://www.strava.com/oauth/token"
TOKEN_FILE = Path.home() / ".strava" / "tokens.json"
def get_access_token():
t = json.loads(TOKEN_FILE.read_text())
if t["expires_at"] - time.time() > 300: # 5 min safety margin
return t["access_token"]
r = requests.post(TOKEN_URL, data={
"client_id": os.environ["STRAVA_CLIENT_ID"],
"client_secret": os.environ["STRAVA_CLIENT_SECRET"],
"grant_type": "refresh_token",
"refresh_token": t["refresh_token"],
}, timeout=30)
r.raise_for_status()
fresh = r.json()
t.update({
"access_token": fresh["access_token"],
"refresh_token": fresh["refresh_token"], # may have rotated
"expires_at": fresh["expires_at"],
})
TOKEN_FILE.write_text(json.dumps(t, indent=2))
return t["access_token"]
Set STRAVA_CLIENT_ID and STRAVA_CLIENT_SECRET in your environment, not in the file. You will eventually push this to a repo.
Paginating a year of activities
The summary endpoint is /api/v3/athlete/activities, capped at 200 results per page. For a rider with 412 activities in the last twelve months that’s three requests, which is nothing against the standard read limit of 100 requests per 15 minutes and 1,000 per day. Check the X-RateLimit-Limit and X-RateLimit-Usage headers on any response to see where you actually stand; they come back as comma-separated pairs like 100,1000 and 7,143.
API = "https://www.strava.com/api/v3"
def fetch_summaries(after_epoch=0, per_page=200):
token = get_access_token()
headers = {"Authorization": f"Bearer {token}"}
page, rows = 1, []
while True:
r = requests.get(f"{API}/athlete/activities", headers=headers,
params={"after": after_epoch, "per_page": per_page,
"page": page}, timeout=30)
if r.status_code == 429:
sleep_for = 900 - (time.time() % 900) + 5 # to next quarter-hour
print(f"rate limited, sleeping {sleep_for:.0f}s")
time.sleep(sleep_for)
continue
r.raise_for_status()
batch = r.json()
print(f"page {page}: {len(batch)} activities "
f"(usage {r.headers.get('X-RateLimit-Usage')})")
if not batch:
break
rows.extend(batch)
page += 1
time.sleep(0.3)
return rows
Note the quirk: when you pass after, results come back oldest-first, and when you don’t, they come back newest-first. If you’re building incremental sync on top of this, don’t assume ordering, sort explicitly.
Flattening to a DataFrame without losing the plot
The summary payload has about 50 fields. Six of them matter for training analysis, and one of them is a trap.
import pandas as pd
KEEP = ["id", "name", "sport_type", "workout_type", "start_date",
"distance", "moving_time", "elapsed_time", "total_elevation_gain",
"average_watts", "weighted_average_watts", "max_watts", "kilojoules",
"device_watts", "trainer", "average_heartrate", "max_heartrate",
"average_cadence", "suffer_score"]
def to_frame(rows):
df = pd.DataFrame(rows)[KEEP].copy()
# start_date_local carries a bogus Z suffix. Parse true UTC, then convert.
df["start"] = (pd.to_datetime(df["start_date"], utc=True)
.dt.tz_convert("Europe/London"))
df["km"] = df["distance"] / 1000
df["hours"] = df["moving_time"] / 3600
df["elev_m"] = df["total_elevation_gain"]
# device_watts False means Strava ESTIMATED the power. Bin it.
df.loc[~df["device_watts"].fillna(False),
["average_watts", "weighted_average_watts", "max_watts", "kilojoules"]] = pd.NA
return df.drop(columns=["distance", "moving_time", "start_date"]) \
.sort_values("start").reset_index(drop=True)
That device_watts filter is the single most important line in the script. Strava estimates power for any ride without a meter, and those estimates land in the same average_watts field as your Quarq or 4iiii numbers. Feed the unfiltered column to a model and ask it about your aerobic progression, and it will confidently narrate a trend built partly on guesses from your phone-only café rides. Same story with trainer: Zwift rides arrive as sport_type == "VirtualRide" with genuine power, but their NP-to-heart-rate relationship is nothing like a Cotswold gravel loop, and mixing them unlabelled will flatten any decoupling analysis you try later.
workout_type is worth keeping too. For rides, 10 is the default, 11 flags a race and 12 a workout. Check it against one ride you know you raced, because the mapping isn’t well documented and Strava’s app doesn’t always write it.
Cache it, then sync incrementally
Write to Parquet, not CSV. Dtypes survive, the timezone survives, and 324 rides compress to roughly 90 KB.
CACHE = Path("data/activities.parquet")
def sync():
if CACHE.exists():
old = pd.read_parquet(CACHE)
after = int(old["start"].max().timestamp()) - 3600 # 1h overlap
else:
old, after = None, 0
new = to_frame(fetch_summaries(after_epoch=after))
df = new if old is None else pd.concat([old, new], ignore_index=True)
df = df.drop_duplicates(subset="id", keep="last").sort_values("start")
CACHE.parent.mkdir(exist_ok=True)
df.to_parquet(CACHE, index=False)
print(f"{len(df)} activities, {df['start'].min():%Y-%m-%d} to "
f"{df['start'].max():%Y-%m-%d}")
return df
The one-hour overlap on after handles activities you edited or uploaded late, and drop_duplicates(keep="last") means the fresher version wins. Run this from cron or Task Scheduler once a day and the whole thing costs one or two API calls.
What you can now answer in four lines
Here’s the payoff. A monthly rollup with an approximate TSS from weighted average power, FTP set at 285 W:
rides = df[df["sport_type"].isin(["Ride", "GravelRide", "VirtualRide"])].copy()
FTP = 285
rides["if_"] = rides["weighted_average_watts"] / FTP
rides["tss"] = rides["if_"] ** 2 * rides["hours"] * 100
monthly = rides.set_index("start").resample("MS").agg(
rides=("id", "count"), km=("km", "sum"), hours=("hours", "sum"),
elev_m=("elev_m", "sum"), np_mean=("weighted_average_watts", "mean"),
tss=("tss", "sum")).round(1)
print(monthly)
rides km hours elev_m np_mean tss
start
2025-10-01 24 712.4 26.1 7845 214.0 1402.3
2025-11-01 22 588.1 23.7 5912 208.0 1265.1
2025-12-01 19 441.6 19.8 3744 201.0 1019.4
2026-01-01 26 502.3 24.5 3088 219.0 1387.6
2026-02-01 25 611.7 25.9 4512 224.0 1509.2
2026-03-01 28 794.2 30.2 7106 221.0 1730.8
2026-04-01 31 938.5 34.8 9744 218.0 1962.4
2026-05-01 33 1042.9 37.1 11208 226.0 2147.9
2026-06-01 30 989.4 34.2 10455 231.0 2035.3
2026-07-01 34 1118.6 38.6 12011 229.0 2265.7
2026-08-01 27 851.2 29.4 8902 222.0 1686.0
2026-09-01 25 766.8 27.5 7988 217.0 1548.1
That’s 324 rides, 9,358 km, 352 hours and about 19,960 TSS over twelve months. Twelve rows, and you can paste them into Claude or any other model with a question like “my A race was 12 July, does this build look like it peaked in the right place” and actually check the answer against the frame. Compare that to asking the same question of a chat interface with no file attached and seeing what it invents.
What the summary endpoint won’t give you
No streams. No power curve, no lap splits, no second-by-second data. For that you need /activities/{id}/streams?keys=time,watts,heartrate,altitude,cadence&key_by_type=true, and it’s one request per activity. Pulling streams for all 324 rides means 324 requests, which at 100 per 15 minutes takes just over an hour of wall-clock time with sleeps. Cache each one to data/streams/{id}.parquet and never fetch it twice; a three-hour ride at 1 Hz is roughly 11,000 rows across six channels, about 200 KB compressed, so the full year lands near 65 MB. That’s the raw material for mean-maximal power curves and aerobic decoupling, and it’s the subject of the next build in the Build Your Own Training Tools series.
No FTP history either. /athlete returns your current FTP if you’ve set one, but Strava keeps no record of when it changed, so if you want honest historical IF you’ll need to maintain a small CSV of FTP by date and merge it on. Ten lines, and it fixes the single biggest source of nonsense in retrospective load analysis.
One last thing before you close the terminal: open Strava in a browser, go to your training log, and count your rides for a month you know well. If your DataFrame says 27 and the website says 31, you almost certainly authorised with activity:read instead of activity:read_all. Revoke the app at strava.com/settings/apps, run the authorisation URL again with the right scope, delete data/activities.parquet, and resync from zero. It takes four minutes and it’s the difference between a dataset you can trust and one that will quietly mislead every tool you point at it.