AI Cycling
§6.1 Build Your Own Training Tools 1,898 words · 9 min

Strava API With Python: Auth, Rate Limits And Your First Dataset

Most people who try to pull their own rides out of Strava hit the same wall in the same order. OAuth works on the third attempt, the activity list comes down fine, then somewhere around ride 180 of a two-season backfill the script starts returning 429 and they discover the hard way that there is no Retry-After header to lean on. This page is about getting past that properly: the token flow you set up once, the four headers that tell you exactly how much budget you have left, and a request plan that gets a season of power files onto your disk without babysitting.

If you want the wider map of what to build once you have the data (notebooks, intervals.icu comparisons, LLM-assisted analysis), that lives on Build Your Own Training Tools. This page is the plumbing.

The limits, in actual numbers

Strava applies two separate budgets to your client_id, and each has a 15-minute component and a daily component:

Budget15 minutesPer day
Overall requests2002,000
Read requests1001,000

Nearly everything you will do is a read, so treat 100 per quarter-hour and 1,000 per day as your real ceiling. Two details matter more than the numbers themselves.

First, the 15-minute windows are fixed to the clock in UTC, not rolling from your first request. They reset at :00, :15, :30 and :45. If you fire 100 reads at 14:44:30 you are unblocked 30 seconds later, and if you fire them at 14:30:05 you wait nearly fifteen minutes. The daily counter resets at midnight UTC, which for UK riders in summer means 01:00 local.

Second, the budget belongs to the application, not the athlete. If you have a cron job syncing overnight and you open a notebook to poke at something using the same client_id, they are eating the same 100.

Your first dataset costs three requests, not six hundred

Before writing any rate-limit machinery, check whether you need the expensive endpoints at all. GET /athlete/activities returns a SummaryActivity per ride and accepts per_page=200. Three requests gets you 600 activities, which for most self-coached riders is two to three full seasons.

Each summary already carries the fields that answer a lot of real questions:

moving_time, elapsed_time, distance, total_elevation_gain,
average_watts, weighted_average_watts, max_watts, kilojoules,
average_heartrate, max_heartrate, average_cadence, average_speed,
device_watts, trainer, commute, gear_id, suffer_score, start_date_local

weighted_average_watts is Strava’s Weighted Average Power, close enough to Normalised Power for trend work. device_watts is the flag that tells you whether the numbers came from a power meter or Strava’s estimate, and you should filter on it before you compare anything. Season-over-season TSS proxies, weekly kJ, cadence drift on your TT bike: all of that is three requests and a pd.json_normalize.

You only need streams when you care about what happened inside a ride.

OAuth, once, then never again

Register an app at strava.com/settings/api. The field labelled Authorization Callback Domain wants a domain, not a URL, and localhost is accepted. The scope you want is activity:read_all: plain read will not return your activities, and without _all you lose anything you marked private, which for a lot of people means their indoor sessions.

Run this once to capture a refresh token:

import http.server, json, pathlib, secrets, urllib.parse, webbrowser, requests

CLIENT_ID, CLIENT_SECRET = "123456", "your-secret"
REDIRECT = "http://localhost:8721/exchange"
STATE = secrets.token_urlsafe(16)

auth = "https://www.strava.com/oauth/authorize?" + urllib.parse.urlencode({
    "client_id": CLIENT_ID, "redirect_uri": REDIRECT, "response_type": "code",
    "approval_prompt": "force", "state": STATE,
    "scope": "read,activity:read_all,profile:read_all",
})

class Catch(http.server.BaseHTTPRequestHandler):
    def do_GET(self):
        q = urllib.parse.parse_qs(urllib.parse.urlparse(self.path).query)
        assert q.get("state", [""])[0] == STATE, "state mismatch"
        tok = requests.post("https://www.strava.com/oauth/token", data={
            "client_id": CLIENT_ID, "client_secret": CLIENT_SECRET,
            "code": q["code"][0], "grant_type": "authorization_code",
        }, timeout=30).json()
        pathlib.Path("strava_token.json").write_text(json.dumps(tok, indent=2))
        self.send_response(200); self.end_headers()
        self.wfile.write(b"Token stored. Close this tab.")
        print("granted scope:", q.get("scope"))

    def log_message(self, *a): pass

webbrowser.open(auth)
http.server.HTTPServer(("localhost", 8721), Catch).handle_request()

Check the printed scope against what you asked for. Strava lets the athlete tick fewer boxes than you requested, and the failure mode is a cheerful empty list rather than an error.

Access tokens last six hours (expires_in: 21600). The refresh call is where people lose a weekend, because Strava rotates the refresh token on some responses and a script that only saves access_token will be permanently locked out the next time it runs:

import time, json, pathlib, requests

STORE = pathlib.Path("strava_token.json")

def access_token():
    tok = json.loads(STORE.read_text())
    if tok["expires_at"] - time.time() > 300:
        return tok["access_token"]
    fresh = requests.post("https://www.strava.com/oauth/token", data={
        "client_id": CLIENT_ID, "client_secret": CLIENT_SECRET,
        "grant_type": "refresh_token", "refresh_token": tok["refresh_token"],
    }, timeout=30)
    fresh.raise_for_status()
    tok.update(fresh.json())          # keeps the NEW refresh_token
    STORE.write_text(json.dumps(tok, indent=2))
    return tok["access_token"]

Write the whole payload back. Every time.

The four headers that make throttling a non-event

Every response carries your current usage. Log them once and the guesswork disappears:

X-RateLimit-Limit:        200,2000
X-RateLimit-Usage:        87,415
X-ReadRateLimit-Limit:    100,1000
X-ReadRateLimit-Usage:    87,415

Format is fifteen_minute,daily. A limiter that reads those and sleeps to the next clock boundary beats exponential backoff, because backoff guesses and the header knows:

from datetime import datetime, timedelta, timezone
import time

def seconds_to_window_edge():
    now = datetime.now(timezone.utc)
    edge = now.replace(minute=0, second=0, microsecond=0) \
             + timedelta(minutes=(now.minute // 15 + 1) * 15)
    return (edge - now).total_seconds() + 2

class Budget:
    short = day = 0
    short_cap, day_cap = 100, 1000

    def observe(self, headers):
        cap = headers.get("X-ReadRateLimit-Limit")
        use = headers.get("X-ReadRateLimit-Usage")
        if cap: self.short_cap, self.day_cap = (int(x) for x in cap.split(","))
        if use: self.short, self.day = (int(x) for x in use.split(","))

    def gate(self, reserve=3):
        if self.day >= self.day_cap - reserve:
            raise SystemExit(f"daily read cap reached ({self.day}/{self.day_cap})")
        if self.short >= self.short_cap - reserve:
            wait = seconds_to_window_edge()
            print(f"  {self.short}/{self.short_cap} used, sleeping {wait:.0f}s")
            time.sleep(wait)
            self.short = 0

The reserve=3 matters if anything else shares the client. Leaving a few requests unspent is cheaper than an unhandled 429 in the middle of a long backfill.

Cache to disk before you parse anything

Streams are the expensive part of the job, and they never change once a ride is uploaded. Cache raw JSON keyed on the request, and a re-run of your script costs zero requests:

import hashlib, json, pathlib, requests

CACHE = pathlib.Path("strava_cache"); CACHE.mkdir(exist_ok=True)
budget = Budget()

def get(path, params=None):
    key = hashlib.sha1(
        (path + json.dumps(params or {}, sort_keys=True)).encode()
    ).hexdigest()
    hit = CACHE / f"{key}.json"
    if hit.exists():
        return json.loads(hit.read_text())

    budget.gate()
    r = requests.get("https://www.strava.com/api/v3" + path, params=params,
                     headers={"Authorization": f"Bearer {access_token()}"},
                     timeout=30)
    budget.observe(r.headers)
    if r.status_code == 429:
        time.sleep(seconds_to_window_edge())
        return get(path, params)
    if r.status_code == 404:
        return None                      # deleted or hidden, don't retry forever
    r.raise_for_status()
    hit.write_text(json.dumps(r.json()))
    return r.json()

Two seasons of streams runs to a few hundred megabytes of JSON on disk. That is nothing next to the cost of re-fetching it, and it means your parsing logic can be wrong nine times without burning budget.

Budget arithmetic for a two-season backfill

Say 600 rides since January 2024. Here is what each level of detail costs:

What you pullRequestsDays at 1,000/day
Activity index only3same afternoon
Index + detailed activity (laps, segment efforts, device_watts)6031
Index + streams6031
Index + detail + streams1,2032

Running flat out at 100 reads per quarter-hour, 1,000 requests takes about two and a half hours of wall clock. A sensible pattern: one night for streams on rides where device_watts is true, a second night for detailed activities if you actually need segment efforts. Filter before you fetch. A commuter ride with no power meter is not worth a stream request, and trainer: true rides have no useful latlng or altitude.

stravalib will handle the OAuth refresh and a basic limiter for you if you would rather not write the above, and it gives you typed model objects. The trade-off is that you lose the raw-JSON cache layer unless you add it back, and on a large backfill that layer is the thing that saves you.

Streams are not a tidy table yet

Request them keyed by type, then expect gaps:

KEYS = "time,watts,heartrate,cadence,velocity_smooth,altitude,distance,moving,temp"

def ride_frame(activity_id):
    s = get(f"/activities/{activity_id}/streams",
            {"keys": KEYS, "key_by_type": "true"})
    if not s or "time" not in s:
        return None
    df = pd.DataFrame({k: v["data"] for k, v in s.items()})
    # Garmin smart recording leaves holes: reindex onto 1 Hz
    df = df.set_index("time").reindex(range(0, int(df["time"].max()) + 1))
    df["watts"] = df["watts"].ffill(limit=10).fillna(0) if "watts" in df else 0
    for col in ("heartrate", "cadence", "altitude", "velocity_smooth"):
        if col in df: df[col] = df[col].interpolate(limit=30)
    return df.reset_index(names="t")

That reindex step is the one that quietly ruins numbers if you skip it. Garmin’s smart recording writes a sample every few seconds when nothing is changing, so a 3-hour ride might arrive as 6,400 rows rather than 10,800. Any rolling window you apply in row-space is then measuring the wrong duration, and your 30-second rolling average for Normalised Power becomes a 30-sample average covering anywhere from 30 to 120 real seconds. Zwift and most head units on 1-second recording are fine; the moment you mix devices across a season, you need the time index.

A few more things the API will not warn you about. Missing keys are simply absent from the response rather than present-and-null, so guard every column access. The distance stream is cumulative metres and resets to zero on some multi-file uploads. moving is a boolean stream that disagrees with moving_time on the summary often enough that you should pick one and be consistent. And elapsed_time includes every coffee stop, which makes it useless as a training-load input but the only honest number for a TT effort where you want start-line to finish-line.

The cheaper routes, and when to take them

Strava’s bulk export (Settings, My Account, Download or Delete Your Account, Request Your Archive) hands you every original .fit and .gpx file with no rate limit at all. Files land within a few hours. If your goal is one-off analysis of your own history rather than an app that stays in sync, this is strictly better: FIT files contain left/right balance, core temperature, per-lap data and every developer field your head unit recorded, none of which survives the trip through Strava’s stream endpoints. Parse with fitdecode or python-fitparse.

For ongoing sync, intervals.icu is worth a look before you write a single Strava request. Its API uses HTTP basic auth with an API key (username API_KEY, password the key itself), and GET /api/v1/athlete/{id}/activities returns a rich per-ride summary including its own eFTP estimates, load and HRV fields. If your rides already flow Strava to intervals.icu, you can read from intervals.icu and never touch a Strava quota. You also get power curves and custom stream data that Strava does not expose.

One last thing to read rather than assume: Strava’s API Agreement has tightened in recent years, notably around displaying other athletes’ data in your own interface and around using Strava data to train machine learning models. Analysing your own files locally, or pasting your own numbers into an LLM prompt, is a different thing from building a product on it. Check the current terms against what you are actually planning, especially if the plan involves anything that looks like model training.

Next practical step: run the index pull, write the 600 summaries to a Parquet file, and count how many rows have device_watts == True. That number is your real dataset size, and it is usually smaller than people expect.