Files
bike-app/docs/PLAN.md
T
BBergleandClaude Opus 5 7f33cb1593
CI / Repo hygiene (pull_request) Successful in 2s
CI / Web (lint, typecheck, build) (pull_request) Successful in 14s
CI / Migrations reversible (pull_request) Successful in 6s
CI / API (lint, types, tests) (pull_request) Successful in 53s
docs: correct the false Wi-Fi premise; add UI and live-tracking phases
The plan's headline section claimed the Rider 650 has on-device Wi-Fi and a
Data Sync menu that uploads to Bryton's cloud with no phone involved. It
does not. The unit has ANT+ and Bluetooth only; its sole sync route is BLE
to the Bryton Active app. Confirmed on the physical device, corroborated by
BikeRadar's hands-on. Likely origin: conflation with the Rider 750 / S800,
which do have Wi-Fi.

Impact is narrower than it first appears and no built code is invalidated:
everything downstream of Bryton's cloud never depended on how a ride got
into that cloud, so the poller, ingestion, schema, wear engine and all of
Phase 0 stand. What was invalidated is the product promise — Phase 1 was
called "Zero-touch ride history" and claimed to fix the original complaint
(having to remember to open the Active app). It does not; it is one-tap.
Renamed accordingly rather than leaving the doc overclaiming.

Corrections propagated everywhere the premise had spread: PLAN.md's opening
sections, Phase 1, top risks (the chain is now longer and has a human link
that fails silently — earns a "nothing ingested in N days" nudge), and the
verification checklist; DECISIONS.md D3's justification; README.md, which
was additionally stale on nearly every other point (claimed no code written,
Postgres, compose, four containers); and RESEARCH.md, where the claim
originated under a "verified" header it had not earned. RESEARCH.md is
annotated rather than rewritten — it is a record of what was found, and the
correction is part of that record. USB path facts are marked unverified too,
since they came from the same unverified batch.

D20 records the process lesson: the plan contained the right check ("first
action before writing any code"), it was never run, and nothing downstream
required it to have been. Device capabilities get confirmed on the device
before being written as fact.

Also adds the two phases requested before this came up, both grounded in
feasibility research rather than assumption:
- Phase 1A, an open-ended UI pass done together, including the verbose field
  surface driven off activity_field_inventory.
- Phase 1B, live tracking. Constrained hard by reality: iOS suspends
  backgrounded PWAs and implements no Web Bluetooth, and Bryton's own Live
  Track needs the phone relaying over BLE, so the tracking client cannot be
  our PWA. Shape that works is OwnTracks POSTing to our API for position,
  with the server deriving distance/pace/elevation; HR and power need BLE and
  are explicitly a second-class opt-in, not a blocker. Two rules written in:
  live positions must never become activities (invariant #6), and "no privacy
  zones, ever" does not extend to a public live link.

Inserted as 1A/1B rather than renumbering Phases 2-5, whose numbers are
referenced from DECISIONS.md, deploy/README.md and code comments.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R2ZKeWkZV7ehf7fivrAkkG
2026-09-21 23:20:37 -04:00

61 KiB
Raw Blame History

Self-Hosted Cycling App — Implementation Plan

Context

You want a self-hosted Strava replacement that syncs rides from your Bryton Rider 650, tracks your cumulative mileage, and adds two things Strava does badly or not at all: a spare-parts inventory and a maintenance record. It runs on your own Linux server, with source control in self-hosted Gitea and an act_runner on the same box for automated builds.

The specific pain driving this: today your Rider 650 syncs over Bluetooth to the Bryton Active phone app, which forwards to Strava — but Active has no background sync, so you have to remember to open the app. This app is not, on its own, able to fix that last part — see "How rides actually reach the app" below for why, and what it does fix.

You also want mileage-milestone notifications — "every 200 miles, clean and lube the drivetrain" — which makes the maintenance side push-based rather than something you have to remember to go look at.

Directory is empty; this is greenfield. Decisions already made:

  • Build fresh (borrow data models from Endurain and strava-gear, don't fork either)
  • Multi-user — you plus friends/family, invite-only
  • PWA, not a native iOS app — installed to the iPhone home screen
  • Imperial units by default — you think in miles; storage stays SI, display is miles
  • Phased roadmap with the self-hosting-only features staged in deliberately

How rides actually reach the app

Corrected 2026-09-22, after checking the actual device. An earlier version of this plan opened with a "sync breakthrough": the claim that the Rider 650 has on-device Wi-Fi and a Data Sync menu that uploads to Bryton's cloud with no phone involved. That is false. The Rider 650 has ANT+ and Bluetooth only — no Wi-Fi — and the only sync route it offers is Bluetooth to the Bryton Active app. The unit's menu has no Data Sync entry, and BikeRadar's hands-on confirms connectivity is "ANT+ and Bluetooth" with syncing via "Bryton's Active App." The likely source of the error is conflation with the Rider 750 / S800, which do have Wi-Fi. This mattered: it was the headline premise of the whole plan and it survived into a written roadmap unverified. The lesson is recorded in docs/DECISIONS.md — verify device capabilities against the physical device before building a plan on them, not against model-adjacent sources.

The real chain, which is what everything downstream is built on:

ride ends → BLE → Bryton Active app on your phone → Bryton cloud
          → your server's poller fetches the original FIT → app

The one fact that still holds, and is the load-bearing one: Bryton's cloud API is fully reverse-engineered and returns the original, unmodified FIT bytes. Working MIT reference implementation: github.com/jorge-huxley/intervalssync (Python, updated 2026-09-17). Everything this app does downstream of Bryton's cloud — ingestion, dedupe, the schema, the wear engine, the garage, notifications — never depended on how a ride got into that cloud, which is why losing the Wi-Fi premise costs far less than it first appears.

What this does and doesn't fix. It does not fix the original complaint. You still have to open the Active app for a ride to leave the head unit; nothing in this app can reach across that gap (see "Why not Bluetooth direct" below). What it does fix is everything after that: one tap and the ride is permanently yours — full-resolution, original bytes, in your own database, with wear recalculated and maintenance reminders armed, and no third party able to change the terms later.

Worth trying, costs nothing, no code: an iOS Shortcuts personal automation to open Bryton Active for you — triggered on the Rider 650's Bluetooth disconnecting, or on arriving home. If iOS honours it reliably, most of the hands-free behaviour comes back without any architectural change. Try this before concluding the one-tap step is permanent.

Why not Strava as the source: Strava's API has no export_original endpoint — you get decoded, smoothed streams, never the original file. Its June 2026 tier restructure also caps new apps at 10 users and requires the developer to hold a paid Strava subscription. Dead end; skip it.

Why not Bluetooth direct: Bryton's BLE sync protocol is not reverse-engineered by anyone — no Gadgetbridge support, no ANT-FS, no published UUIDs. Independently of the protocol, iOS gives web apps no Bluetooth at all (Web Bluetooth is unimplemented in WebKit, with no public Apple position), so the installed PWA structurally cannot talk to the head unit even if the protocol were known. Pulling rides off the unit directly would mean reverse-engineering an undocumented protocol and running it on non-iOS hardware in the house (a Pi, an old Android phone). That's a research project of unknown size, not a schedulable phase. Out of scope — but it is the only route that would truly remove the phone, so it is the thing to revisit if the one-tap step ever becomes intolerable.


Stack

Layer Choice Why
API Python 3.12, FastAPI, Pydantic v2, SQLAlchemy 2.0 async, Alembic fitdecode is Python-only, so Python is forced; FastAPI emits OpenAPI 3.1 free
DB SQLite (single file, inside the app container) Chosen over Postgres+PostGIS specifically to keep the whole deploy to one container — see docs/DECISIONS.md D15 for the full reasoning and what it costs (no database-level RLS, no PostGIS, procrastinate needs replacing)
Jobs TBD before Phase 1procrastinate no longer fits (Postgres-only) Nothing enqueues a job yet; pick this when the ingestion pipeline actually needs it, not before
Frontend SvelteKit adapter-static SPA + PWA — no Node process in prod Static files served by Caddy; enforces API-first by construction
Maps MapLibre GL JS from day one Renders raster tiles now, self-hosted PMTiles vector later — a config change, not a rewrite
FIT parsing fitdecode Thread-safe, preserves header+CRC, correct developer-field handling; python-fitparse's own maintainers point here
Auth Argon2id + opaque bearer tokens, no JWT Instant revocation, device list, no key rotation. Statelessness buys nothing at 15 users
Notifications apprise library (ntfy default) A library, not another container

Rejected: Redis (nothing needs it), TimescaleDB (see streams decision), SSR, Celery, native iOS.

The PWA decision

A home-screen-installed PWA gets you the icon, full-screen chrome-free display, offline caching, and — since iOS 16.4 — real push notifications, which is the only thing that used to force native. Going native would cost $99/yr for an Apple Developer account, TestFlight or sideloading to get it onto family phones, and a second codebase forever. The two real PWA gaps on iOS (Web Bluetooth, Background Sync) don't matter here: Bryton BLE is a dead end anyway, and the server does all syncing.

Design consequences — the iOS-specific traps, all of which have cheap answers if you know them upfront:

Limitation Design response
Push requires "Add to Home Screen" Silently fails from a plain Safari tab. Treat install as a mandatory onboarding flow with a persistent card (iOS has no beforeinstallprompt, so there's no programmatic install). Only offer "Enable reminders" once display-mode: standalone is true.
A denied notification permission is sticky The user must delete and reinstall the PWA to be asked again. So: request only from a direct user gesture, only in standalone, only after an explanatory screen. Never burn the prompt.
An installed PWA has its own storage partition, separate from Safari The user will be logged in in Safari and logged out in the installed app and think it's broken. Document it in onboarding; 90-day cookie; "keep me signed in" on by default.
No Web NFC Works anyway with no app: an NTAG sticker encoding https://host/b/<tag_uid> — iOS background tag reading fires from the lock screen and opens the URL, and the PWA's scope claims it so it opens in the app. Server-side that's one route.
No BarcodeDetector getUserMediazxing-wasm, lazy-loaded so the ~300KB WASM isn't in the shell bundle.
No Background Sync Irrelevant — all syncing is server-side. Offline writes go to a small IndexedDB outbox flushed on online/visibilitychange, keyed by the client-generated UUIDv7 that is already the row's PK, so replay is idempotent with no server-side idempotency table.
Storage eviction Bounded caches with explicit ExpirationPlugin limits (unbounded caches are what trigger eviction of the whole origin). Never treat client storage as durable; mark unsaved outbox items visibly rather than optimistically pretending they saved.
Aggressive service-worker termination The SW does two things only: caching and displaying pushes. Nothing load-bearing.

Service worker via @vite-pwa/sveltekit in injectManifest mode. Caching: precache the hashed shell; cache-first on index.html for navigations (so it opens instantly and offline); NetworkFirst with a 3s timeout on /garage and /maintenance/due so the garage with no signal still shows "chain: 4,100 / 4,800 mi" — the highest-value offline surface in the app; CacheFirst on tiles; network-only on all mutations, never silently cached.

Keep the backend API-first anyway (versioned /api/v1, bearer tokens, committed OpenAPI schema). It costs almost nothing, and if Apple ever makes the PWA route untenable, a native client becomes a code-generation exercise against a stable contract rather than a rewrite.


Service topology

Revised by D15 — single container, not the multi-container compose topology this section originally described. Caddy and the FastAPI app run together in one image via a lightweight process supervisor; SQLite is a file inside the same container's persistent volume, not a separate service. See docs/DECISIONS.md D15 for why, and deploy/'s own README for the actual supervisor config once it exists.

Volumes: one persistent volume holding the SQLite file, blobstore (content-addressed raw FIT + attachments), and import_inbox (USB watch folder bind-mount).

Later phases that would have been "add a container" under the old topologytileserver, topodata, ollama, grafana, valhalla/photon/overpass — still make sense as genuinely separate containers even under a single-container-for-the-app model (they're independent services with their own resource profiles, not part of "the app"). Whether the app container talks to them over a shared Docker network or they stay fully optional add-ons is a decision for whichever phase actually needs the first one — not resolved here speculatively.

Explicitly NOT on day one: Valhalla, Photon, Overpass, Nominatim (>1TB/128GB RAM — never), Ollama, Redis, MinIO, Grafana.


Schema

Note on what's below vs. what's actually built: the Identity tables (§ Tables > Identity) are real, built, and running on SQLite — apps/api/velodrome/models/identity.py and apps/api/alembic/versions/0001_baseline.py are the source of truth for those, not this prose. Everything past Identity (activities, streams, bikes/components, service rules, notifications, weather) was designed against Postgres/PostGIS conventions — geometry(...) columns, ARRAY, INET, GIST indexes, RLS policies — before D15 moved the database to SQLite. None of it is built yet, so none of it is broken; it just needs a real pass for SQLite compatibility (TEXT/BLOB for geometry, JSON-encoded TEXT for arrays, plain TEXT for IP addresses, application-layer exclusion checks instead of EXCLUDE USING gist) when each phase actually builds it, informed by whatever's learned finishing the SQLite migration on Identity first — not a mechanical find-and-replace on speculative schema now.

Conventions: UUIDv7 PKs (time-ordered, URL-safe, client-generatable). All timestamps UTC (SQLite has no native timezone-aware timestamp type — see apps/api/velodrome/models/base.py for how Identity stores them; the same convention applies going forward). All physical quantities as SI integers — distance in metres, time in seconds, speed in mm/s, altitude in cm, money in minor units. Units are a presentation concern.

Two architectural rules that everything else depends on

Rule 1 — raw bytes are the only truth. Every ingested file is written to a content-addressed blob store before parsing and is never mutated or deleted (ON DELETE RESTRICT). Every table below is a rebuildable projection: delete the projection, re-parse, and you must get the same result. This is what turns "parser bug corrupted 4,000 elevations" and "retroactively reprocess all history against a better DEM" from disasters/migrations into routine batch jobs. It is the single most important rule here.

Rule 2 — there is no odometer column anywhere. Component wear is derived by replaying the activity stream against time-ranged install records (the strava-gear insight, made relational). Correcting "I actually swapped that chain a week earlier" is one UPDATE, and every downstream number self-corrects. A stored counter cannot do that, nor handle parts moving between bikes, without a reconciliation nightmare.

Tables

Identity: users (carries timezone and unit_system, defaulting to imperial for you — it drives display only, never storage), invites (only sha256(code) stored), sessions (opaque token hashes), api_tokens (scoped, for Home Assistant/Grafana).

Raw storage: raw_filescontent_sha256, storage_path (blobstore/ab/cd/<sha>.fit), source (upload|usb|bryton_cloud|gpx_import), source_ref, parse_state (pending|parsed|failed|quarantined|not_an_activity), parser_version, fit_type. UNIQUE (user_id, content_sha256).

Activities: activities (+ activity_laps, activity_stats). Notable columns:

  • ascent_device_m and ascent_dem_m kept separately with an elevation_source flag, so Phase 3 DEM reprocessing never destroys the barometric original.
  • track geometry(LineStringZM, 4326) full-res, plus track_simplified (~10m) for list maps, plus a generated bbox. GIST index on the simplified track.
  • wet_fraction (denormalised from weather, so the wear query needs no join), is_indoor, counts_for_wear (user override).
  • fit_time_created + device_serial → partial unique index. This is the natural dedupe key.

Capture everything the head unit emits, not a whitelist. A hard requirement, not a nice-to-have: whatever fields a Rider 650 puts in a FIT file should end up queryable and displayable, including fields that aren't in the standard FIT profile. fitdecode surfaces all of it — every message type (including ones it doesn't recognise, by message number), every field (unrecognised ones as unknown_<n>), and developer fields with their definition metadata. The parser must therefore be field-agnostic by construction: iterate the messages that are actually present and persist what is found, rather than reading a fixed list of known field names and silently dropping the rest.

Concretely, three things this implies beyond the tables above:

  • activity_streams takes any channel. channel is already a free string, so a new or unknown per-record field (unknown_61, a developer field, a Bryton-specific extension) becomes a stream row with no schema change. No whitelist anywhere in the parse path.
  • activity_fit_messages — the non-time-series long tail, which the current tables have nowhere to put: device_info (firmware, battery, every paired sensor), event (start/stop/lap triggers, battery and sensor warnings), hrv, zones_target, workout/workout_step, sport, plus any message type we don't recognise. Stored as (activity_id, message_type, message_index, fields json) — JSON is correct here precisely because the shape is unknown and variable, which is the opposite of the streams case where it's uniform and huge.
  • activity_field_inventory — per activity, which channels and message types actually turned up, with units and value ranges. This is what lets the UI render "everything we got from this ride" dynamically instead of hardcoding a field list that goes stale the moment Bryton's firmware adds something. It also makes "what does this head unit actually record?" answerable without scanning every stream.

None of this risks anything, because of invariant #1: the raw bytes are retained forever, so a parser that learns to understand more fields later is a parser_version bump and a reparse, not a migration or a data-loss event. Verbosity is a projection-widening exercise and can be iterated on safely.

Streams — columnar arrays, one row per channel (activity_streams: channel, n, scale, values_i32[]). Decision and justification:

  • Size. A 3h ride at 1Hz × ~9 channels: normalized per-sample rows ≈ 1.2MB + 0.3MB index; scaled-int arrays ≈ 150250KB after TOAST/LZ4, zero index cost. ~5× multiplier, compounding forever under a full-retention requirement.
  • Read pattern. Every real query is "give me the whole stream to draw a chart" — one TOAST fetch, and it maps 1:1 onto the JSON the client wants with no row-to-column pivot.
  • Lat/lon are free. FIT stores position as int32 semicircles; values_i32 holds them losslessly.
  • Not JSONB (untyped, 35× larger, slow to deserialise). Not TimescaleDB — it solves cross-entity scans over a firehose; we do per-entity blob reads. Adopting it means abandoning the postgis/postgis base image and taking on extension-version coupling at every Postgres upgrade.
  • Aggregates (HR zone totals, power curve) are precomputed at ingest into activity_stats, never scanned from streams.

Bikes and parts — the part of the schema most people get wrong:

  • bikes — includes initial_distance_m (km ridden before this system existed) and nfc_tag_uid.
  • component_models — shared catalogue (kind, manufacturer, model, spec jsonb, gtin barcode).
  • components — a tracked individual physical object with identity and history. Has purchase_cost_minor, initial_distance_m, inventory_item_id provenance.
  • component_installstime-ranged association, mounted to a bike XOR a parent component (cassette → wheelset → bike). A GIST EXCLUDE constraint enforces that a component can only be in one place at a time. Moving a wheelset between bikes is two rows.
  • inventory_itemsfungible shelf stock with a quantity, min_quantity low-stock threshold, location, gtin. Plus inventory_transactions, an append-only ledger whose running sum is the quantity.

Why inventory and components are separate tables: a spare chain on the shelf has no identity worth tracking — you own "3 × Shimano CN-M8100", not three named chains. A fitted chain has identity, install history, accrued km, and a cost-per-km. "Install from inventory" is the state transition: decrement quantity, create a components row carrying the cost and a provenance link, create an install row. Modelling stock as components-without-installs forces fake identities onto consumables (sealant, cables, bar tape) and makes "how many chains do I have left?" a COUNT over a table that also contains every chain you retired since 2019.

Maintenance: service_rules (scoped to a component XOR bike XOR kind), service_events, service_event_components, attachments (receipts, photos, manuals — polymorphic on entity_type/id).

Notifications: push_subscriptions, notification_log, notification_prefs, odometer_milestones — defined in the milestones section below.

Integrations: integration_credentials (AES-GCM encrypted, key from the compose .env), integration_health (consecutive failures + error taxonomy, drives alerting).

Weather: weather_observations keyed on a 0.05° grid cell + UTC hour, so nearby rides reuse cached data, plus a per-activity activity_weather rollup.

Live tracking (Phase 1B): live_sessions (user_id, started_at, ended_at, share_token_hash — only the hash, same discipline as invites and sessions — expires_at, obfuscate_endpoints_m, and a nullable activity_id reconciled after the real FIT file arrives) plus live_positions (session_id, ts, lat/lon as int32 semicircles, altitude_cm, speed_mms, accuracy_m, and a nullable JSON column for whatever optional sensor metrics a screen-on client manages to send). live_positions is append-only and high-write relative to everything else here; it is also the only table in the schema that is deliberately not permanent — once a session is reconciled to its activity, the positions are redundant against the FIT file's own record, and can be pruned on a retention window without losing anything. This is the single exception to "every table is a rebuildable projection of raw bytes," and it is an exception precisely because live telemetry has no raw file behind it.

The wear engine

service_rules carries metric (distance | ride_time | calendar), threshold, basis (since_install | since_last_service), wet_multiplier, and include_indoor. Three SQL layers:

  1. v_component_activity — joins installs to activities on the time range.
  2. component_usage(component, since, wet_mult, include_indoor) — sums distance_m * (1 + (wet_mult - 1) * wet_fraction), so a fully-wet ride on rim pads (wet_multiplier = 4.0) counts 4×, a dry ride 1×, a half-wet ride 2.5×. Raw distance is retained separately so the UI can show "480 km ridden / 1,150 km effective wear".
  3. v_component_due — percentage used, remaining, and a projected due date from your trailing 90-day rate.

One view answers every maintenance question in the product:

  • chain @ 4,800km — distance, since_install, wet 2.0
  • chain wax @ 350km — distance, since_last_service, wet 3.0
  • fork lowers @ 50 ride hoursride_time, since_last_service, include_indoor=false
  • sealant @ 90 dayscalendar, since_last_service
  • BB @ 6 months OR 4,800km — two rules on one component; whichever hits 100% first wins the badge

component_usage_cache is refreshed by a debounced job after ingest, install edits, and service events. The view stays authoritative; the dashboard reads the cache.

Cost-per-km falls straight out of purchase_cost_minor / raw_distance_m — a genuinely differentiating number no cloud service gives you, for zero incremental work once these tables exist.


Mileage milestones and notifications

This is the feature that makes the maintenance side push-based instead of something you have to remember to go and check. Two distinct kinds of milestone, sharing one delivery system.

Kind 1 — recurring maintenance intervals

"Every 200 miles, clean and lube the drivetrain" is already expressible: a service_rule with metric='distance', threshold=321869 (200 mi in metres), basis='since_last_service'. The since_last_service basis is what makes it recurring — log the service, the basis moves forward, and the counter re-arms automatically. No separate "repeating rule" concept is needed.

Seed catalogue, shipped as defaults so the app is useful the moment you add a bike, with every rule editable and dismissible. Intervals below are the consensus from mainstream cycling maintenance guides; stored in metres/seconds/days, displayed in miles.

Recurring tasks (basis='since_last_service'):

Task Interval Metric Wet × Notes
Clean & lube drivetrain 200 mi distance 2.5 your example; the headline default
Wipe & re-lube chain (dry lube) 150 mi distance 3.0 dusty conditions shorten this
Wipe & re-lube chain (wet lube) 250 mi distance 3.0
Re-wax chain (if waxing) 200 mi distance 4.0
Measure chain wear with a gauge 500 mi distance 1.0 replace at 0.5% for 11/12-speed
Inspect brake pads 500 mi distance 2.0 discs: replace under 1.5mm
Check BB & headset for play 500 mi distance 1.0
Check tyre pressure 7 days calendar
Inspect / top up tubeless sealant 90 days calendar 1.0 it evaporates
Inspect cables & housing 1,000 mi distance 1.5
Bolt torque check 90 days calendar
Service hub/BB/headset bearings 2,000 mi or 180 days two rules 2.0 whichever first
Replace cables & housing 2,500 mi distance 1.5
Full annual service 365 days calendar
Fork lowers service 50 ride hours ride_time 1.0 MTB; include_indoor=false
Fork/shock full service 150 ride hours ride_time 1.0

Replacement rules (basis='since_install', is_replacement=true, retires the component):

Part Interval Wet ×
Chain 2,000 mi 2.0
Cassette 6,000 mi 2.0
Chainrings 15,000 mi 2.0
Rear tyre 2,500 mi 1.5
Front tyre 5,000 mi 1.5
Disc brake pads 1,200 mi 4.0
Rim brake pads 1,500 mi 4.0
Bar tape 365 days

Note the rear tyre wearing 23× faster than the front is why tyre_front and tyre_rear are separate kind values rather than one "tyre" kind with a shared interval.

Kind 2 — odometer achievement milestones

The other reading of "mileage milestones": "the Ribble just passed 5,000 miles", "you've done 1,000 miles this year." Cheap to add and satisfying, which is the whole point of a mileage tracker.

CREATE TABLE odometer_milestones (
  id uuid PRIMARY KEY, user_id uuid NOT NULL,
  scope text NOT NULL,              -- 'user' | 'bike' | 'component'
  scope_id uuid,
  period text NOT NULL,             -- 'lifetime' | 'year' | 'month'
  period_key text,                  -- '2026' for yearly
  threshold_m bigint NOT NULL,
  reached_at timestamptz NOT NULL,
  activity_id uuid REFERENCES activities(id),   -- the ride that crossed it
  UNIQUE (user_id, scope, scope_id, period, period_key, threshold_m)
);

Ladders (all configurable): bikes every 500 mi lifetime; you every 1,000 mi lifetime; round numbers 100/250/500/1,000/2,500/5,000/10,000; annual goal progress at 25/50/75/100%. The UNIQUE constraint means a milestone fires exactly once, ever — and the activity_id link lets the notification say "your 5,000th mile on the Ribble was on this morning's ride."

Delivery

CREATE TABLE push_subscriptions (
  id uuid PRIMARY KEY, user_id uuid NOT NULL REFERENCES users(id) ON DELETE CASCADE,
  endpoint text NOT NULL UNIQUE,    -- e.g. https://web.push.apple.com/...
  p256dh text NOT NULL, auth text NOT NULL,
  user_agent text, is_standalone boolean,
  created_at timestamptz NOT NULL DEFAULT now(),
  last_success_at timestamptz, failure_count int NOT NULL DEFAULT 0
);

CREATE TABLE notification_log (     -- idempotency AND the in-app inbox
  id uuid PRIMARY KEY, user_id uuid NOT NULL,
  kind text NOT NULL,               -- 'service_due'|'milestone'|'low_stock'|'sync_failed'|'weekly_summary'
  dedupe_key text NOT NULL,
  title text, body text, url text,
  created_at timestamptz NOT NULL DEFAULT now(),
  pushed_at timestamptz, read_at timestamptz,
  UNIQUE (user_id, dedupe_key)
);

CREATE TABLE notification_prefs (
  user_id uuid PRIMARY KEY REFERENCES users(id) ON DELETE CASCADE,
  service_warn boolean DEFAULT true,      -- fire at warn_at_pct (80%)
  service_due boolean DEFAULT true,       -- fire at 100%
  milestones boolean DEFAULT true,
  low_stock boolean DEFAULT true,
  sync_health boolean DEFAULT true,
  digest_mode text DEFAULT 'immediate',   -- 'immediate' | 'daily' | 'weekly'
  digest_hour smallint DEFAULT 18,
  quiet_hours_start time, quiet_hours_end time   -- interpreted in users.timezone
);

The dedupe key is what stops this becoming spam, and it's the one part that's easy to get wrong. Key format: service_due:<rule_id>:<component_id>:<cycle_seq>:<threshold>, where cycle_seq is the count of service events logged against that (component, rule) so far. Consequences:

  • Within one cycle, each threshold fires exactly once — one nudge at 80%, one at 100%, then silence. It does not re-notify nightly about the same chain.
  • Logging the service increments cycle_seq, so the next 200-mile crossing is a new key and fires again. That's how "every 200 miles" repeats forever without a cron-style recurrence engine.
  • Milestones use milestone:<scope>:<scope_id>:<period_key>:<threshold_m> and are naturally once-ever.

Evaluation job evaluate_notifications runs (a) after every ingest, so a milestone or a newly-due service arrives within minutes of the ride landing, and (b) nightly, to catch calendar-based rules that no ride triggers. It reads component_usage_cache against v_component_due, inserts notification_log rows, and enqueues sends on the notify queue. Quiet hours defer rather than drop.

Channels, in order of reliability:

  1. ntfy via apprise — the primary. apprise is already in the stack for poller alerts, works on iOS through the ntfy app with no PWA-install requirement, and is the most robust option. Roughly an hour of work, which is why notifications can ship in Phase 1 rather than waiting for the PWA plumbing.
  2. Web Push (VAPID) — the nicer experience. Apple's web.push.apple.com endpoint speaks standard RFC 8291, so pywebpush reaches an iPhone with no Apple Developer account and no APNs certificate — the only requirement is that the PWA is installed to the home screen. 410 Gone/404 prunes the subscription row; 429 backs off. navigator.setAppBadge(n) puts a count of outstanding due items on the home-screen icon for free.
  3. Email via apprise, for weekly digests.

Push is never the only path. Every notification is a notification_log row rendered as an in-app inbox, so a dead push channel degrades the experience without breaking the feature. That matters because iOS silently drops push subscriptions after OS updates and long idle periods.


Ingestion pipeline

One canonical path. Every source funnels through ingest_bytes(user_id, data, source, source_ref) -> RawFile before any parsing. One parser, one dedupe implementation, one set of side effects.

manual upload ─┐
USB watcher   ─┼→ ingest_bytes() → raw_files row + blob (same txn) → enqueue parse_raw_file
Bryton poller ─┘                                                              │
                                                                              ▼
                                        fitdecode → discriminate → rebuild projection in ONE txn
                                        → enqueue compute_stats, enrich_weather, recompute_wear

Idempotent by construction: INSERT ... ON CONFLICT (user_id, content_sha256) DO NOTHING RETURNING id. Nothing returned ⇒ already have it ⇒ enqueue nothing. Uploading the same file 100 times costs 100 hashes and zero rows.

Dedupe, three layers in order:

  1. Content hash — catches re-uploads and USB rescans. Filename is never consulted (Bryton reuses names).
  2. FIT natural key (device_serial + file_id.time_created) — catches the same ride arriving via USB and the cloud where the bytes differ. This is what makes dual-source operation safe.
  3. Temporal overlap heuristic — same user, start within 90s, duration within 5%, distance within 2% ⇒ flag duplicate_of_id, keep both, offer a UI action. Never auto-delete, only auto-hide.

Activity vs. course discrimination — a real bug waiting to happen. Bryton writes routes as .fit too. Check file_id.type (4 = activity, 6 = course, 32 = monitoring). Bryton's encoder is not Garmin's, so if the type is absent or nonstandard, fall back to message-shape inspection: ≥1 session and ≥1 timestamped record ⇒ activity; course/course_point messages or positions without timestamps ⇒ course. Anything unclassifiable ⇒ quarantined, blob retained, one alert, surfaced in an admin list. Never silently drop.

Other Bryton hardening: map nonstandard manufacturer/product IDs via serial prefix and store the raw values; store laps verbatim but never derive session totals by summing laps (use the session message); compute moving_time from records with speed > 0.5 m/s if absent.

The parser reads what's there, not what it expects. Per the "capture everything" requirement in the Schema section: walk every message and every field fitdecode yields, persist unrecognised ones under their raw identifiers (unknown_<n>, developer fields with their definition metadata) rather than skipping them, and record what was found in activity_field_inventory. A field the parser doesn't have a name for is still worth storing and still worth showing — Bryton's encoder is not Garmin's, and the whole point is to see everything the head unit actually recorded. A parser change that narrows what gets captured is a regression, and the golden-fixture corpus should catch it: assert on the field inventory of a known file, not just on the handful of summary numbers.

Sources implement a common ActivitySource protocol:

  • UploadPOST /api/v1/uploads, multipart, 50MB cap, accepts .fit, .fit.gz, .gpx, .tcx.
  • USB watcher — 60s periodic scan of /import/inbox/**/*.fit. A host udev rule on volume label Bryton mounts the device read-only and rsyncs into the inbox. Discover the subfolder at runtime by recursive glob — sources disagree on whether it's Activities/, Actives/, or root, so don't hardcode it (your 650 is documented as Bryton/Activities/, but verify).
  • Bryton cloud (Phase 1 — the primary path) — Meteor DDP over SockJS to m3.brytonactive.com: login with the SHA-256 digest → subscribe("activityList") → read userActivities (filter _deleted tombstones) → diff against raw_files.source_refGET /api/activity?id=<id> with X-User-Id, X-Auth-Token, x-api-key, User-Agent: okhttp/4.12.0 → raw original FIT bytes. Poll every 20 min, jittered. Vendor the intervalssync protocol logic with the upstream commit SHA in a header comment rather than taking a runtime dependency on a reverse-engineering project.

Credential warning: the Bryton SHA-256 digest is the credential — it replays as a password. Store it AES-GCM encrypted with a key from the compose .env (not in the DB), never return it from any endpoint (the Pydantic response model simply doesn't contain the field), and redact it in logs. Say so plainly in the setup UI.

Job runner — procrastinate, chosen over Celery (needs Redis, poor async, Postgres broker is second-class), arq (Redis-only — a whole container for ~50 tasks/day), and APScheduler (a scheduler, not a durable queue — no retries, no dead-lettering). The decisive property is transactional enqueue: the raw_files INSERT and the parse-job enqueue commit atomically on one connection. No orphaned blobs, no jobs referencing rolled-back rows. Impossible with a Redis broker without inventing an outbox. Queues: ingest (2), enrich (4), maintenance (1).

Poller failure alertingintegration_health tracks consecutive_failures and an error taxonomy (auth|protocol|network|ratelimit). Via apprise/ntfy, max once per 24h: 3 consecutive failures; no success in 48h while enabled; auth or protocol errors alert immediately on the first failureprotocol is the API-changed-under-us signal. A nightly canary fetches the activity list only and asserts it parses, so breakage surfaces on rest days rather than three weeks later. Persistent UI banner while unhealthy; /api/v1/health/integrations feeds Home Assistant in Phase 4.


Auth

Opaque bearer tokens against a sessions table. No JWT. At your scale, verification is one indexed PK lookup (~0.1ms), and you get instant revocation, an "active devices" list, and no key-rotation or clock-skew bugs. JWT's only advantage is stateless horizontal scale, which will never arrive here.

One token, two transports, one verification path: web gets Set-Cookie: HttpOnly; Secure; SameSite=Lax, any future non-browser client gets Authorization: Bearer. A single FastAPI dependency reads bearer first, then cookie, then sets app.user_id for RLS. The web client gets no capability another client lacks — the cookie is transport convenience only.

  • CSRF: cookie-authenticated mutations require Origin to match the configured public URL; bearer requests skip it (attackers can't set that header). SameSite=Lax as defence in depth.
  • Isolation, enforced twice: a repository layer where every query starts from a scoped(User) base, and Postgres RLS enabled from day one with SET LOCAL app.user_id per request transaction. Migrations run as a BYPASSRLS owner; the app connects as a non-owner. RLS is cheap now and means re-auditing every query later. Friends/family visibility in Phase 2 is an additive widened policy, never removal of the default scope.
  • Invites: open signup does not exist as a setting. Admin generates a code; only sha256(code) is stored; registration validates and increments used_count in the same transaction with SELECT ... FOR UPDATE so a shared link can't be used twice concurrently. First user is bootstrapped by CLI (docker compose run api velodrome create-admin), not a web setup wizard a scanner could race.
  • Argon2id t=3, m=64MiB, p=4; login rate-limited per-IP and per-account in Postgres.

Unique self-hosting features (staged by value/effort)

These are the payoff for self-hosting — things Strava structurally cannot do.

Free, because they're schema properties:

  • No privacy zones on your own archive. No third party holds your data, so show real door-to-door routes. (This reasoning covers the private archive only — a publicly shareable live-tracking link is a different risk and gets its own treatment; see the live tracking phase.)
  • Every field the head unit recorded, not the handful a platform chose to keep. Strava's API gives you decoded, smoothed streams for a fixed set of channels; upload a FIT file there and the unrecognised and vendor-specific fields are simply gone. Here the original bytes are retained forever and the parser stores unknown and developer fields under their raw identifiers, so the UI can show everything the Rider 650 actually wrote — including fields nobody has named yet.
  • Unlimited full-resolution retention, forever, of the original files.
  • Cost-per-km on every component, and per-kind averages ("my chains cost £0.019/km").
  • Receipts and photos attached to parts, service events, and bikes.
  • Wet-weighted wear — rim pads genuinely wear ~4× faster in the rain, and you have the weather data.

Cheap and high value:

  • Live tracking on your own terms (Phase 1B) — a share link your family opens with no Bryton account, no third party holding the trace, that keeps working if Bryton's service dies, and whose history lands in your own database next to the ride it belongs to. Note honestly that Bryton's own Live Track already does the live-map part and the phone has to be present either way; what self-hosting buys is ownership, not capability.
  • Mileage-milestone and maintenance pushes straight to your phone — the reason the garage data is worth keeping. Strava's gear tracking can't do interval reminders at all.
  • Grafana pointed straight at Postgres — roughly an afternoon, the cheapest analytics in the plan.
  • Barcode-scan parts into inventory at purchase/install time (zxing-wasm).
  • Low-stock alerts — "you're down to your last chain and the fitted one is at 87%".
  • Bulk-import your entire GPX archive regardless of file count — no API quotas.

Phase 34, genuinely differentiated:

  • Retroactive elevation reprocessing of every historical ride against a better DEM — the direct payoff for the immutable-bytes rule.
  • Personal segment matching against your own ride archive — your own PRs, no third-party segment database, nothing made public.
  • Overnight batch compute on idle server time: heatmap tiles, segment PRs, power curves.
  • Home Assistant entities — "km until chain due", "days since last ride", "spare chains in stock".
  • NFC tag per bike — tap the frame, its maintenance page opens (iOS Shortcuts, no app needed).
  • Local LLM ride summaries via Ollama, behind a compose profile. Nothing leaves the house.

Roadmap

Estimates assume one developer working evenings and weekends.

Phase 0 — Scaffolding. Done, with two deliberate deviations from this original description — both recorded in docs/DECISIONS.md, not silent drift. Monorepo, uv/ruff/mypy --strict, FastAPI skeleton with /healthz and OpenAPI, Alembic baseline (users/invites/sessions), SvelteKit static SPA shell with login, manifest + service worker, Caddy, Gitea Actions green, image in the registry, deployed to a real Unraid host behind real HTTPS.

  • SQLite, not Postgres+RLS. D15 reversed D4 mid-Phase-0, after the RLS version was already shipped and merged. Isolation is now enforced entirely at the repository layer (db.py's Scope), not database RLS. See D15 for the full cost/benefit record.
  • CI builds and pushes on every push to main, but does not auto-redeploy the running container. D16/D17 made this deliberate: the runner shares this Unraid host's own dockerd (DooD), and an unattended redeploy of a container on a personal server with no human gate was judged the wrong default. A push to main gets you a new image in the registry within minutes; getting it onto the running container is still a manual step (deploy/README.md). An auto-updater (Watchtower or similar) was attempted and deferred — see "Deliberately deferred" in docs/DECISIONS.md.
  • VAPID/push notifications were never started. Correctly so — per this doc's own PWA-decision table, that's gated behind standalone-mode detection and belongs to a later phase, not Phase 0.

Done when — status: Logging in at the real URL (https://bike.bbergle.com) and seeing an authenticated view of your own account is verified, including the session actually persisting (scripts/smoke-test.sh, added after a real Secure-cookie-over-HTTP bug on the first deploy — see that script's header comment). Not yet tried: adding it to an iPhone home screen and confirming a standalone launch — nobody has actually done this yet, so it isn't checked off, even though the manifest and service worker are in place.

Phase 1 — One-tap ride history (68 weeks). Ingestion core (all three dedupe layers, course discrimination, quarantine, and the capture-everything field handling from the Schema section); the Bryton cloud poller as the primary path, polling every 15 min with the full integration_health alerting stack; USB watcher for historical backfill and as the break-glass path; manual upload; activity list/detail with MapLibre and stream charts; totals and trends by week/month/year and per bike, in miles; bikes CRUD; odometer milestone notifications (they only need activities, so they ship here); invites; nightly SQLite snapshot + restic (not pg_dump — see D15). Done when: you finish a ride, open the Active app once, put the bike away, and within 15 minutes it's in your app — full-resolution original bytes, every field the head unit recorded, wear recalculated — with no further interaction, and you stop opening Strava to look at your own data.

Renamed from "Zero-touch" deliberately. The original criterion said "tap Data Sync on the 650… zero further interaction," which the device cannot do (see "How rides actually reach the app"). The honest bar is one tap in the Active app, not zero. This phase therefore does not fully fix the original complaint — it fixes everything downstream of it. Removing that last tap needs either the iOS Shortcuts automation trick (free, unproven, try it) or reverse-engineering Bryton's BLE on non-iOS hardware (unbounded, out of scope). Don't let this phase quietly grow to chase it. Protect it from scope creep.

Phase 1A — Make it yours: the UI pass (open-ended, done together). Phase 1 deliberately ships a plain, functional UI — correctness of the data first, because a beautiful page over wrong numbers is worse than an ugly page over right ones. This phase is the opposite: no new data, no new pipeline, just making the thing feel like yours, working through it together rather than against a spec written in advance.

What it covers: the ride list and ride detail layout; which numbers are hero numbers and which are buried; chart design for the stream data; the verbose field surface — everything the head unit recorded, driven off activity_field_inventory so unknown and vendor-specific fields appear rather than being silently hidden; dark mode; the mobile layout, since the real reading device is a phone on a home screen; and the empty/loading/error states that a plain Phase 1 will have done crudely.

Why it's a phase and not a task: UI taste isn't specifiable up front by either of us — it needs real rides on a real screen and a few rounds of "no, bigger / not that / what if the map was the whole page." Budgeting it as its own phase makes that iteration legitimate rather than scope creep against Phase 1. Done when: you'd rather open this than Strava to look at a ride you just did — and you can find every field the Rider 650 recorded without asking where it went.

Phase 1B — Live tracking (35 weeks). A self-hosted equivalent of Bryton's Live Track: someone at home opens a link and watches your position and live metrics move on a map.

Read the constraints before designing anything here — they are hard, and they shape the feature:

  • The phone must be in your pocket. Bryton's own Live Track requires the Active app running and relaying over BLE (Rider 650 manual, "LIVE TRACK"); the head unit has no independent uplink. No self-hosted design changes that.
  • An installed iOS PWA cannot do this. iOS suspends JS when backgrounded or screen-locked, so a PWA can only track foreground with the screen on (Wake Lock, iOS 18.4+, helps but doesn't lift the restriction). And WebKit implements no Web Bluetooth at all, so the PWA cannot read HR/power/cadence sensors under any circumstances.
  • Therefore: the tracking client is not our PWA. It is an existing, backgrounded app POSTing to our API.

The shape that actually works: OwnTracks (free, open-source, App Store) in HTTP mode, POSTing to an authenticated endpoint on our server — genuinely backgrounded, screen off, phone in pocket, ~30sfew-minute fixes in "move" mode. That gives position + GPS speed, and the server derives the rest from the position stream: distance, elapsed and moving time, current/average pace, elevation gain (via the Phase 3 DEM), and progress against the route if one is loaded. That is a genuinely useful live metric set with no BLE at all.

What is not achievable backgrounded: heart rate, power, cadence. Those need BLE, which means either a screen-on phone mounted on the bars running a BLE-capable browser (a second-class, opt-in mode — not the installed PWA), a companion Android device or LTE tracker in a jersey pocket, or a native app and a $99/yr Apple Developer account. Ship the location-first version; treat sensor metrics as a separate, explicitly optional follow-on, and don't let them block the useful 80%.

Two design rules this phase must not break:

  1. Live positions are telemetry, not a ride. They must never become an activity — the real activity still arrives as original FIT bytes via Bryton (invariant #6, one ingestion path). A live session is linked to the activity it corresponds to after the fact; it is not a second, lower-fidelity source of truth. Getting this wrong produces duplicate, worse rides.
  2. "No privacy zones, ever" does not extend to a public live link. That stance is sound for your own archive on your own server; it is not sound for a URL that shows strangers your current location, or your home, in real time. This phase needs: expiring share tokens, explicit start/stop (plus an auto-end on inactivity so a forgotten session doesn't broadcast indefinitely), and the option to blur the first and last N metres.

Done when: your family can open a link while you're out, see where you are and how far you've gone, and the link stops working when the ride does.

On the phase numbering: 1A and 1B are inserted rather than renumbering Phases 25, because "Phase 2"/"Phase 3" are referenced from docs/DECISIONS.md, deploy/README.md and code comments, and silently shifting their meaning is exactly the kind of stale cross-reference D20 is about.

Phase 2 — The garage (57 weeks). Components and time-ranged installs; inventory with the stock ledger and install-from-stock; service events with photo/receipt attachments; the seeded service-rule catalogue across all three metrics with due/warn badges and projected-due dates; component_usage_cache and its recompute job; maintenance notifications end-to-end (ntfy first, Web Push second, in-app inbox always) with the cycle-seq dedupe; low-stock alerts; Open-Meteo weather enrichment; wet-ride weighting switched on. Done when: you get a push saying "Drivetrain clean & lube due on the Ribble — 205 miles since last time. You have 2 chains on shelf B," and can log the job from the garage floor.

Phase 3 — Own the map + cheap differentiators (46 weeks). Regional PMTiles + tileserver-gl, Open Topo Data with retroactive re-elevation of the whole archive, overnight heatmap generation, personal segment matching, cost-per-km dashboards, barcode scanning, NFC deep-link routes, nested installs (wheelsets). Done when: no third-party network request is needed in normal use.

Phase 4 — Home and household (35 weeks). Offline write outbox, friends/family visibility and a household feed, Home Assistant entities via scoped API tokens, Grafana on the obs profile, Ollama ride summaries on the ai profile, weekly summary push. Done when: a display in the hallway shows the next thing that needs doing to a bike.

Phase 5 — Routing (optional). Valhalla, Photon, Overpass surface tags, route planning with .fit course export back to the Rider 650. Several GB of RAM for something Komoot already does well — build it only if Phase 4 leaves appetite.


Repo layout and Gitea CI/CD

Monorepo — one maintainer, and API and client change together constantly.

apps/api/          pyproject.toml, velodrome/{api,models,ingest,sources,jobs,wear}/, alembic/, tests/
apps/web/          SvelteKit PWA, src/lib/api/ (generated types)
packages/openapi/  openapi.json  ← COMMITTED; the contract artefact
deploy/            docker-compose.yml, .prod.yml, Caddyfile, .env.example,
                   systemd/velodrome-backup.{service,timer}, scripts/restore-drill.sh
.gitea/workflows/  ci.yml release.yml deploy.yml renovate.yml nightly.yml

ci.yml — push/PR with paths: filters so a web change doesn't run pytest. actions/cache@v4 works (act_runner has a built-in cache server). API job runs lint/mypy/pytest against a postgis/postgis:16-3.4 service container (service containers work in Docker mode) with real Alembic migrations and a corpus of ~20 real Rider 650 .fit files as golden parser fixtures. An openapi-drift job regenerates the schema and git diff --exit-codes it. A migrations job asserts upgrade → downgrade -1 → upgrade succeeds and alembic check finds no model drift.

release.yml — on v* tags, buildx multi-stage, push to the Gitea registry. Authenticate with a PAT secret (package:write scope) — secrets.GITEA_TOKEN cannot push to the Gitea container registry. This is a documented Gitea limitation and will waste an hour if forgotten.

deploy.ymlworkflow_dispatch + on release. Runs on a [self-hosted, host]-labelled runner in host mode so it can reach the Docker socket: compose pullrun --rm api alembic upgrade headcompose up -dcurl -f /healthz. Migrations run as an explicit step before up -d, never from the container entrypoint, so failures fail the deploy visibly. Note jobs.*.environment is silently ignored by Gitea — there are no environment protection rules, so the gate is manual dispatch plus a namespaced PROD_* secret.

renovate.yml — self-hosted Renovate (Dependabot is GitHub-only), weekly cron plus workflow_dispatch, because Gitea's cron scheduler has shipped flaky releases and you need a manual trigger. Exclude fitdecode and anything Bryton-adjacent from auto-merge.

nightly.yml — Bryton canary, projection-integrity check, restic check --read-data-subset=5%, heatmap/segment recompute.

Security note: the runner holds the Docker socket, which is root-equivalent on the host. Acceptable for a private single-maintainer instance — but never make this repo public or add untrusted collaborators without disabling Actions on fork PRs.

Backups are a systemd timer on the host, not a Gitea Action — a consistent SQLite snapshot (sqlite3 .backup or VACUUM INTO, not a raw file copy of a live database; pg_dump no longer applies, see D15) + restic to B2/S3 with append-only repo credentials (so a compromised app host can't delete history), plus a restic snapshot of blobstore/. Backups must not depend on CI, because CI is the thing most likely to be broken when you need a restore. Quarterly restore-drill.sh restores into a throwaway stack and asserts activity counts match.


Top risks

1. The Bryton chain breaks silently — and it is now a longer chain than originally planned. With the Wi-Fi premise gone, every automatically-ingested ride passes through head unit → BLE → Active app → Bryton cloud → poller, and only the last link is ours. Two of those links can fail quietly: the app not being opened (rides simply never leave the unit — invisible to the server, which cannot distinguish "no rides uploaded" from "you didn't ride"), and Bryton's private API itself (hardcoded key, undocumented protocol, zero stability guarantee). Mitigations: the poller is one ActivitySource among several, and the USB path ships in Phase 1 so the system is never dependent on the cloud; nightly canary; immediate alerting on auth/protocol errors; vendored protocol pinned to an upstream SHA so fixes are a diff, not a re-derivation; dual-source dedupe on the FIT natural key means you can fall back to USB mid-week and lose nothing and duplicate nothing. New mitigation the longer chain earns: a "nothing ingested in N days" nudge, so the silent failure mode of simply forgetting to open Active surfaces as a notification rather than as a gap you notice months later.

2. Losing or corrupting years of ride history. The realistic threats are mundane — a parser bug writes wrong elevation to 4,000 rides, a migration drops a column, a disk dies. The raw-bytes-are-truth rule is the control: every derived table rebuilds from the blob store, so a parser bug is a parser_version bump and a requeue, not data loss. Plus offsite append-only backups independent of CI, migration up/down/up testing in CI, golden FIT fixtures, and quarterly restore drills — an untested backup is a hypothesis.

3. Never shipping. This scope is multiple person-years if attacked at once, and the failure mode is a half-built system where you're still syncing manually. Mitigations: every phase has a behavioural done criterion, not a feature list; v1 is capped at four containers; the heavy geo/AI services are behind opt-in profiles in Phases 35; the highest-value differentiators (cost-per-km, wet-weighted wear) are schema properties that cost nothing; and Phase 2 is scheduled early and protected, because if the project stalls right after it, it has still succeeded.


Verification

Before coding — and this section is the reason the Wi-Fi premise survived as long as it did, so treat it as load-bearing, not boilerplate. The original version of this checklist said to verify Main Menu → Data Sync on the Rider 650. That check was never run, and the feature does not exist; a false premise sat at the top of this plan through an entire phase of work. Any capability of a physical device that a phase depends on gets confirmed on the device before it is written down as a fact.

Still worth doing before Phase 1:

  • Plug the Rider 650 in over USB and ls -R the mounted volume to confirm the actual .fit path (sources disagree: Activities/, Actives/, or root).
  • Ride, open the Active app, and confirm the activity reaches Bryton's cloud — then confirm the intervalssync protocol actually retrieves it, before building a poller on the assumption.
  • Test the iOS Shortcuts automation (open Active on Rider 650 BLE disconnect, or on arriving home) and see whether it fires reliably. This determines whether the one-tap step is permanent.

Phase 0 — done, verified for real, not just assumed from CI going green: curl https://host/healthz returns 200; a push to main produces a new registry image (docs/DECISIONS.md D17's release workflow) — redeploying the running container from it is a manual step (D16/D17), not automatic; docker exec velodrome velodrome create-admin bootstraps the first user (D18); logging in at the real HTTPS URL and staying logged in on the next request is checked by scripts/smoke-test.sh, not eyeballed in a browser, after that exact failure mode (a Secure cookie silently dropped when tested over plain HTTP) actually happened on the first deploy.

Phase 1 — ingestion:

  • Upload a real Rider 650 .fit → activity appears with correct distance, elevation, and map track.
  • Upload the same file twice → exactly one activity, one raw_files row.
  • Upload a Bryton route/course .fit → does not become an activity; classified correctly.
  • Run the poller against your real Bryton account → new rides ingest with byte-identical content to the USB copy (compare content_sha256).
  • Import the same ride via both USB and cloud → one activity; dedupe layer 2 catches it.
  • Break the credential deliberately → an auth alert fires immediately and the UI banner appears.
  • Confirm the stored Bryton credential appears in no API response and no log line.
  • pytest green against the golden-fixture corpus in CI.
  • Log in as a second invited user → sees none of your data. Verify RLS directly: SET app.user_id to user B, SELECT * FROM activities, expect zero of user A's rows.
  • Cross 1,000 miles on a bike → exactly one milestone notification, naming the ride that crossed it.

Phase 1A — UI: open a ride you actually did on your actual phone and check you can find every field the head unit recorded without hunting; confirm a file containing an unknown or developer field still surfaces it (feed a golden fixture with a deliberately nonstandard field and check it renders rather than vanishing).

Phase 1B — live tracking:

  • Start a session, lock the phone, put it in a jersey pocket, ride — confirm positions keep arriving with the screen off. This is the whole feature; if it only works screen-on, it has failed.
  • Open the share link on a device that has never logged in → the map moves.
  • Let the token expire (or end the ride) → the same link stops working.
  • Confirm a live session never produces an activity row, and that once the real FIT file lands via Bryton, the session reconciles to it rather than sitting alongside as a duplicate.
  • Enable endpoint obfuscation, then check the public link genuinely does not reveal your house.

Phase 2 — garage and notifications:

  • Create a bike, install a chain, import 3 rides → chain shows summed distance. Edit the install date backwards → the number self-corrects with no manual recomputation.
  • 90-day sealant rule → due badge at the right date. 50-ride-hour fork rule → tracks hours not miles, and ignores indoor rides.
  • Install a part from inventory → quantity decrements, a components row appears carrying the cost, ledger balances.
  • The recurrence test, which is the one that matters: set the 200-mile drivetrain rule, ride past 160 mi → one "80%" notification. Ride past 200 mi → one "due" notification. Run the evaluator ten more times → no further notifications. Log the service → ride another 200 mi → it fires again.
  • Import a wet ride (wet_fraction near 1.0) on rim pads → effective wear advances ~4× raw distance, and the UI shows both numbers.
  • Install the PWA on a second family member's iPhone, enable reminders, trigger a due rule → push arrives on the lock screen and the app badge shows the count.
  • Disable push and re-run → the notification still appears in the in-app inbox and via ntfy.

End-to-end, the real test: finish a ride, put the bike away, don't touch your phone. Within 15 minutes the ride is in the app with weather, every fitted component's wear has moved, and if the drivetrain crossed 200 miles your phone has already told you.