Merge pull request 'refactor(api): move from Postgres+RLS to single-engine SQLite' (#5) from refactor/sqlite-single-container into main
CI / Repo hygiene (push) Successful in 2s
CI / Web (lint, typecheck, build) (push) Successful in 14s
CI / Migrations reversible (push) Successful in 5s
CI / API (lint, types, tests) (push) Successful in 53s

Reviewed-on: #5
This commit was merged in pull request #5.
This commit is contained in:
2026-09-21 15:44:54 -04:00
17 changed files with 548 additions and 544 deletions
+5 -34
View File
@@ -54,27 +54,11 @@ jobs:
api:
name: API (lint, types, tests)
runs-on: ubuntu-latest
services:
postgres:
image: postgis/postgis:16-3.4
env:
# Superuser — this is the "owner" role migrations run as (see alembic/env.py); it's
# what CREATEs the two runtime roles below, which is why it isn't one of them.
POSTGRES_USER: postgres
POSTGRES_PASSWORD: postgres
POSTGRES_DB: velodrome_test
options: >-
--health-cmd pg_isready
--health-interval 10s
--health-timeout 5s
--health-retries 10
# No postgres service container — SQLite (docs/DECISIONS.md D15) needs no server to talk to;
# the test suite creates its own temp file (see tests/conftest.py). This also makes the job
# meaningfully faster: no service container to start and health-check before tests can run.
env:
VELODROME_ENVIRONMENT: test
VELODROME_DATABASE_URL_MIGRATE: postgresql+asyncpg://postgres:postgres@postgres:5432/velodrome_test
VELODROME_DB_APP_PASSWORD: ci-only-app-password
VELODROME_DB_AUTH_PASSWORD: ci-only-auth-password
VELODROME_DATABASE_URL_APP: postgresql+asyncpg://velodrome_app:ci-only-app-password@postgres:5432/velodrome_test
VELODROME_DATABASE_URL_AUTH: postgresql+asyncpg://velodrome_auth:ci-only-auth-password@postgres:5432/velodrome_test
steps:
- uses: actions/checkout@v4
@@ -165,22 +149,9 @@ jobs:
migrations:
name: Migrations reversible
runs-on: ubuntu-latest
services:
postgres:
image: postgis/postgis:16-3.4
env:
POSTGRES_USER: postgres
POSTGRES_PASSWORD: postgres
POSTGRES_DB: velodrome_mig
options: >-
--health-cmd pg_isready
--health-interval 10s
--health-timeout 5s
--health-retries 10
# No postgres service container — see the api job's comment; same reasoning.
env:
VELODROME_DATABASE_URL_MIGRATE: postgresql+asyncpg://postgres:postgres@postgres:5432/velodrome_mig
VELODROME_DB_APP_PASSWORD: ci-only-app-password
VELODROME_DB_AUTH_PASSWORD: ci-only-auth-password
VELODROME_DATABASE_URL: sqlite+aiosqlite:///./velodrome-ci-migrations.db
steps:
- uses: actions/checkout@v4
+15 -7
View File
@@ -17,7 +17,8 @@ say so and argue it — but don't silently contradict it.
apps/api/ Python 3.12 / FastAPI / SQLAlchemy async / Alembic
apps/web/ SvelteKit static SPA (installable PWA)
packages/openapi/ openapi.json — COMMITTED contract artefact, CI enforces it matches the code
deploy/ docker-compose, Caddyfile, systemd units, backup scripts
deploy/ single-container Dockerfile, Caddyfile, systemd units, backup scripts
see docs/DECISIONS.md D15 for why this isn't docker-compose
docs/ plan, decisions, research
scripts/ repo tooling (PR helpers, etc.)
.gitea/workflows/ CI
@@ -34,8 +35,12 @@ These are load-bearing. Breaking one is a correctness bug, not a style choice.
time-ranged `component_installs`. Never add a stored running total to a component.
3. **All physical quantities are SI integers** in storage — metres, seconds, mm/s, centimetres,
grams, minor currency units. Imperial is display-only. Never store a float mile.
4. **Every user-owned table has `user_id`, an RLS policy, and a repository-layer scope.** Both
layers, always. Never rely on the query alone.
4. **Every user-owned table has `user_id`, and every query against it goes through the
repository-layer scope helper — never a raw query filtered by hand.** This used to be backed
by Postgres RLS as a second, database-enforced layer (see `docs/DECISIONS.md` D4/D15); SQLite
has no equivalent, so the repository-layer scope is now the *only* enforcement, which makes it
non-negotiable rather than defense-in-depth. A new domain table without a passing isolation
test (see `tests/test_auth.py`'s pattern) is not done.
5. **Secrets never leave the server.** The Bryton credential is password-equivalent. It must not
appear in any API response model, any log line, or any error message.
6. **One ingestion path.** All sources funnel through `ingest_bytes()`. Never add a second parse
@@ -47,8 +52,8 @@ These are load-bearing. Breaking one is a correctness bug, not a style choice.
calls in request handlers.
- **SQL:** migrations via Alembic only, never manual DDL. Every migration must survive
`upgrade -> downgrade -1 -> upgrade`.
- **Tests:** pytest against a real Postgres service container, never mocks for DB behaviour.
Parser changes need a golden fixture in `apps/api/tests/fixtures/fit/`.
- **Tests:** pytest against a real SQLite file, never mocks for DB behaviour. Parser changes need
a golden fixture in `apps/api/tests/fixtures/fit/`.
- **Commits:** imperative mood, explain *why* in the body. Conventional-commit prefixes
(`feat:`, `fix:`, `refactor:`, `test:`, `docs:`, `chore:`, `ci:`).
- Match surrounding code. Don't introduce a new pattern when one exists.
@@ -100,7 +105,10 @@ silently corrupts data for six months is expensive work.
are silent and corrupt the archive.
- `wear/` — the wear SQL. Wrong numbers that still look plausible are the worst failure mode in the
product, because nobody notices.
- `auth/`, RLS policies — security, and a mistake exposes another user's data.
- `auth/`, any repository-layer user-scoping code — security, and a mistake exposes another user's
data. This carries more weight than it used to: there is no database-enforced RLS backstop
anymore (see invariant #4 and `docs/DECISIONS.md` D15), so this code *is* the isolation
boundary, not one layer of it.
- `sources/bryton/` — a reverse-engineered protocol with no spec to check against.
- Schema migrations that alter or drop existing columns.
- Debugging anything that two Sonnet attempts have already failed to fix.
@@ -113,7 +121,7 @@ careful transcription plus ordinary judgement, and CI catches the rest:
- CRUD endpoints, Pydantic models, repository methods
- SvelteKit components, routes, styling, the service worker
- Tests against an already-decided behaviour
- Additive migrations, docker-compose and Caddy config, CI workflows
- Additive migrations, the deploy Dockerfile/Caddy config, CI workflows
- Documentation
### Haiku — mechanical work
+73 -49
View File
@@ -1,7 +1,8 @@
# apps/api
Python 3.12 / FastAPI / SQLAlchemy 2.0 async / Alembic. See `docs/PLAN.md` for the overall
design and `docs/DECISIONS.md` for why things are built this way.
Python 3.12 / FastAPI / SQLAlchemy 2.0 async / Alembic / SQLite. See `docs/PLAN.md` for the
overall design and `docs/DECISIONS.md` (D15 especially) for why the database is SQLite and not
the Postgres+PostGIS setup Phase 0 originally shipped with.
## Layout
@@ -9,47 +10,78 @@ design and `docs/DECISIONS.md` for why things are built this way.
velodrome/
app.py FastAPI app factory
config.py Settings (env-driven, see below)
db.py Two database engines — read this first, it's the load-bearing module
db.py Single engine + the repository-layer scope — read this first, it's the
load-bearing module for user isolation now that there's no RLS
ids.py UUIDv7 generation
models/ SQLAlchemy models
auth/ Password hashing, session service, FastAPI auth dependencies
api/v1/ Route handlers
schemas/ Pydantic request/response models
alembic/ Migrations. 0001_baseline.py creates the identity tables, the two
runtime DB roles, and their RLS policies — read its module docstring.
tests/ pytest, against a real Postgres, never mocked
alembic/ Migrations. 0001_baseline.py creates the identity tables — no roles, no
RLS, no GRANTs, none of those concepts exist in SQLite.
tests/ pytest, against a real SQLite file, never mocked
```
## Why two database connections
## Why isolation is enforced in Python now, not the database
The app connects to Postgres as **two different roles**, not one — this is the single most
important thing to understand before touching `auth/` or `db.py`:
Phase 0 originally ran two Postgres roles (`velodrome_app`/`velodrome_auth`) with row-level
security as a database-enforced isolation layer. SQLite has no roles, no session variables, and no
policy engine — there is no database-level backstop anymore. This is the single most important
thing to understand before touching `auth/` or `db.py`:
- **`velodrome_app`** — `NOBYPASSRLS`. Every request that already knows who's calling uses this,
via `db.scoped_session(user_id)`, which sets `app.user_id` for the transaction so Postgres RLS
policies can key off it.
- **`velodrome_auth`** — `BYPASSRLS`. Used *only* by `auth/service.py`, and only for the narrow
set of lookups that must happen *before* identity is known: login by email, a session by its
token hash, an invite by its code hash — plus the inserts that create those rows. Every one of
those queries is an exact match on a unique key, never an unfiltered scan, which is what makes
bypassing RLS safe there.
- **`scoped_session(user_id)`** — for every query against a user-owned table once identity is
known. Yields a `Scope`, whose `select()` is the *only* way to build a query through it, and
every query it builds is pre-filtered to that `user_id` on any model that declares one. There is
no method on `Scope` that returns an unfiltered query — see `tests/test_auth.py`'s
`test_scope_select_rejects_models_without_user_id` for what happens if you try it on a model
that isn't user-owned (`Invite`, scoped by `created_by` rather than `user_id`, is the real
example used there).
- **`auth_session()`** — used *only* by `auth/service.py`, for the narrow set of lookups that must
happen *before* identity is known: login by email, a session by its token hash, an invite by its
code hash, plus the inserts that create those rows. Every one of those queries is an exact match
on a unique key, never an unfiltered scan — that discipline is what makes it safe to use a plain
session here instead of `Scope`.
- **`unscoped_session()`** — a plain session with no scoping applied at all, for health checks and
anything that never touches a user-owned table. `tests/test_auth.py`'s
`test_unscoped_session_can_see_every_user_when_misused` demonstrates, deliberately, what happens
if this gets used on a user-owned table instead of `scoped_session` — it sees *everyone's* rows.
That test exists to make the point vivid: reaching for `unscoped_session()` (or a raw
`auth_session()` query) against a user-owned table is a review-blocking mistake now, not a style
preference, because there's nothing else standing behind it.
A third connection — `VELODROME_DATABASE_URL_MIGRATE` — is used only by Alembic. It needs enough
privilege to `CREATE ROLE` and `GRANT`, so in practice it's the Postgres superuser (or a
schema-owning role); the app itself never connects with it.
If you're adding a new table with per-user data: give it a `user_id` column and query it *only*
through `scoped_session`. See CLAUDE.md's invariant #4.
If you're adding a new table with per-user data: give it a `user_id` column, enable RLS on it in
a migration, and query it only through `scoped_session`. See CLAUDE.md's invariants.
## Two SQLite behaviours that don't match the defaults you'd expect
Both confirmed empirically against real `aiosqlite`, not assumed from docs — see `db.py`'s
`_configure_sqlite_for_concurrent_writers` for the fixes:
1. **Foreign keys, `ON DELETE CASCADE` included, are OFF by default per connection.** Without
`PRAGMA foreign_keys=ON`, deleting a user silently leaves its sessions/api_tokens behind
instead of cascading — no error either way.
2. **Transactions default to DEFERRED**, which only takes a write lock on the first actual write —
leaving a real check-then-act race window (e.g. two concurrent redemptions of the same invite
code both reading `used_count < max_uses` as true before either commits). Every transaction on
this engine issues `BEGIN IMMEDIATE` instead, which takes the write lock up front. One
consequence worth knowing if you're writing tests: a session that autobegins a transaction via
a bare read and never explicitly commits/rolls back holds that write lock until the session
closes — see the comment above `await db_auth.commit()` in
`test_scoped_session_blocks_cross_user_reads` for a real example of this biting a long-lived
test fixture.
Also worth knowing: `Uuid(as_uuid=True)` stores as 32-char hex **with no hyphens** on SQLite, not
`str(uuid)`'s hyphenated form. This only matters if you ever write a UUID into this schema via raw
SQL instead of the ORM (as the test fixtures do, to set up state without going through the API) —
use `.hex`, not `str()`, or the ORM's own later queries against that row won't match it. See the
docstring on `_seed_invite` in `tests/test_auth.py` for the failure this caused when it was
gotten wrong.
## Environment variables
| Variable | Used by | Notes |
|---|---|---|
| `VELODROME_DATABASE_URL_MIGRATE` | Alembic only | Superuser/owner DSN. Never used by the app itself. |
| `VELODROME_DB_APP_PASSWORD` | Alembic (creates the role) | No default — the migration fails loudly if unset. |
| `VELODROME_DB_AUTH_PASSWORD` | Alembic (creates the role) | Same. |
| `VELODROME_DATABASE_URL_APP` | The app | Connects as `velodrome_app`. |
| `VELODROME_DATABASE_URL_AUTH` | The app | Connects as `velodrome_auth`. |
| `VELODROME_DATABASE_URL` | The app and Alembic, both | e.g. `sqlite+aiosqlite:////data/velodrome.db`. One DSN — there's no separate migration role anymore since SQLite has no roles to separate. |
| `VELODROME_SECRET_KEY` | The app | Not yet used (arrives with the Bryton credential encryption in a later phase); declared now so the settings shape is stable. |
| `VELODROME_ENVIRONMENT` | The app | `development` / `test` / `production`. Gates the session cookie's `Secure` flag — see the comment in `api/v1/auth.py` before changing this condition. |
| `VELODROME_PUBLIC_URL` | The app | Used for the CSRF `Origin` check on cookie-authenticated mutations. |
@@ -59,15 +91,7 @@ a migration, and query it only through `scoped_session`. See CLAUDE.md's invaria
```bash
uv sync --all-extras
# a throwaway Postgres
docker run -d --name velodrome-dev-pg -e POSTGRES_PASSWORD=postgres -e POSTGRES_DB=velodrome \
-p 5432:5432 postgis/postgis:16-3.4
export VELODROME_DATABASE_URL_MIGRATE=postgresql+asyncpg://postgres:postgres@localhost:5432/velodrome
export VELODROME_DB_APP_PASSWORD=devpassword1
export VELODROME_DB_AUTH_PASSWORD=devpassword2
export VELODROME_DATABASE_URL_APP=postgresql+asyncpg://velodrome_app:devpassword1@localhost:5432/velodrome
export VELODROME_DATABASE_URL_AUTH=postgresql+asyncpg://velodrome_auth:devpassword2@localhost:5432/velodrome
export VELODROME_DATABASE_URL=sqlite+aiosqlite:///./velodrome-dev.db
uv run alembic upgrade head
uv run uvicorn velodrome.app:app --reload
@@ -84,21 +108,21 @@ uv run mypy --strict velodrome
uv run pytest
```
Tests run against a real Postgres (a service container in CI, or point `tests/conftest.py`'s
defaults at your own) — never mocks for DB behaviour, per CLAUDE.md. The test suite runs the real
Alembic migration at session start, not a parallel schema-creation shortcut, so it's exercising
the exact same path a real deploy uses.
Tests run against a real SQLite file in a temp directory (never `:memory:`, which gives each
separate connection its own isolated database rather than one shared one, and never mocks for DB
behaviour, per CLAUDE.md). The suite runs the real Alembic migration at session start, not a
parallel schema-creation shortcut, so it's exercising the exact same path a real deploy uses.
The test worth reading if you're new to this codebase is
`tests/test_auth.py::test_rls_blocks_cross_user_session_reads` — it doesn't trust that the RLS
policy SQL is correct because it reads correctly; it proves it by actually trying to read another
user's row through the scoped role and asserting zero come back.
The test worth reading first if you're new to this codebase is
`tests/test_auth.py::test_scoped_session_blocks_cross_user_reads` — it doesn't trust that
`Scope.select()` filters correctly because the code reads correctly; it proves it by registering
two real users and confirming a scoped read for one never returns the other's row, even though
both exist in the same table.
## A note on `alembic check`
CI's `migrations` job runs `alembic check` to catch drift between the ORM models and the actual
migrations. `alembic/env.py`'s `include_object` filter restricts that comparison to tables our own
metadata declares — **deliberately an allowlist, not a denylist of PostGIS/TIGER tables**, because
reflected foreign tables can come back with `schema=None` regardless of which schema they actually
live in (confirmed against a real `postgis/postgis:16-3.4` container), which makes a schema-based
denylist unreliable. If you add a new model, it's automatically covered — no filter to update.
migrations. There's no `include_object` filter in `alembic/env.py` anymore — Phase 0's version
needed one to exclude PostGIS/TIGER's own pre-installed tables from the comparison, but SQLite
starts with nothing but what this app's own migrations create, so there's no foreign-table noise
to filter out in the first place.
+12 -39
View File
@@ -1,16 +1,15 @@
"""Alembic environment.
Deliberately independent of velodrome.config.Settings: migrations run as a privileged/owner
connection (VELODROME_DATABASE_URL_MIGRATE — a superuser or schema-owning role, which trivially
satisfies "BYPASSRLS" since superusers always bypass RLS), never as either of the two runtime
roles the app itself uses. Reading env vars directly here, rather than importing the app's
settings, keeps a migration-only CI job from needing every runtime env var the app requires.
Reads VELODROME_DATABASE_URL directly from the environment rather than importing
velodrome.config.Settings — keeps a migration-only invocation from needing every other runtime
env var the app requires, even though today they'd resolve to the same value. There's no more
separate migration/owner role to reason about here (docs/DECISIONS.md D15): SQLite has no roles,
so migrations run against the exact same file and connection the app itself uses.
"""
import asyncio
import os
import sqlalchemy as sa
from sqlalchemy.engine import Connection
from sqlalchemy.ext.asyncio import async_engine_from_config
@@ -21,60 +20,34 @@ config = context.config
target_metadata = Base.metadata
# postgis/postgis ships an entire pre-installed schema of its own (PostGIS core tables plus the
# TIGER geocoder's tiger/topology schemas — dozens of tables) that Alembic never created and
# doesn't manage. Without this filter, `alembic check`/autogenerate sees every single one as
# "should be dropped" simply because it's not in our SQLAlchemy metadata — which would make CI's
# `alembic check` step permanently useless (always red, for reasons that have nothing to do with
# an actual drift).
#
# A denylist keyed on schema name is NOT reliable here: reflected foreign tables can come back
# with `schema=None` on their Table object regardless of which schema they actually live in on
# the server (confirmed against a real postgis/postgis:16-3.4 container — tables that `\dt`
# clearly shows under the `tiger` schema still reflect with schema=None). An allowlist is the
# robust version of the same idea: only ever compare tables OUR metadata declares, so a future
# PostGIS/TIGER version adding more foreign tables can never cause a false positive here.
def include_object(
object: sa.schema.SchemaItem, name: str | None, type_: str, reflected: bool, compare_to: object
) -> bool:
if type_ == "table":
return name in target_metadata.tables
return True
def _migrate_url() -> str:
url = os.environ.get("VELODROME_DATABASE_URL_MIGRATE")
def _database_url() -> str:
url = os.environ.get("VELODROME_DATABASE_URL")
if not url:
# Local-dev convenience only — every real environment (CI, deploy/) sets this explicitly.
url = "postgresql+asyncpg://postgres:postgres@localhost:5432/velodrome"
# Local-dev convenience only — CI and deploy/ both set this explicitly.
url = "sqlite+aiosqlite:///./velodrome.db"
return url
def run_migrations_offline() -> None:
context.configure(
url=_migrate_url(),
url=_database_url(),
target_metadata=target_metadata,
literal_binds=True,
dialect_opts={"paramstyle": "named"},
include_object=include_object,
)
with context.begin_transaction():
context.run_migrations()
def _do_run_migrations(connection: Connection) -> None:
context.configure(
connection=connection,
target_metadata=target_metadata,
include_object=include_object,
)
context.configure(connection=connection, target_metadata=target_metadata)
with context.begin_transaction():
context.run_migrations()
async def run_migrations_online() -> None:
configuration = config.get_section(config.config_ini_section) or {}
configuration["sqlalchemy.url"] = _migrate_url()
configuration["sqlalchemy.url"] = _database_url()
connectable = async_engine_from_config(configuration, prefix="sqlalchemy.")
async with connectable.connect() as connection:
+22 -155
View File
@@ -1,21 +1,18 @@
"""Baseline: identity tables, two runtime DB roles, RLS policies.
"""Baseline: identity tables (users, invites, sessions, api_tokens).
This migration creates the schema AND the two application-facing Postgres roles it depends on
(velodrome_app, velodrome_auth) — see db.py and auth/service.py module docstrings for the full
rationale. It runs as a privileged/owner connection (VELODROME_DATABASE_URL_MIGRATE), which is
why it's able to CREATE ROLE and GRANT at all; neither of the two roles it creates could do this
to itself.
No roles, no RLS, no GRANTs — SQLite has none of those concepts. Isolation between users is
enforced entirely at the application layer now; see db.py and CLAUDE.md's invariant #4, and
docs/DECISIONS.md D15 for the full story of why this migration looks nothing like the Postgres
version it replaced.
Revision ID: 0001
Revises:
Create Date: 2026-09-21
"""
import os
from collections.abc import Sequence
import sqlalchemy as sa
from sqlalchemy.dialects import postgresql as pg
from alembic import op
@@ -25,77 +22,10 @@ branch_labels: str | Sequence[str] | None = None
depends_on: str | Sequence[str] | None = None
def _require_password(env_var: str) -> str:
value = os.environ.get(env_var)
if not value:
raise RuntimeError(
f"{env_var} must be set before running this migration — see deploy/.env.example. "
"There is no default: these are the passwords for real runtime database roles."
)
return value
def _dollar_quoted(password: str) -> str:
"""Safely embed a password literal in DDL.
CREATE ROLE's PASSWORD clause is DDL, not DML — it does NOT accept bind parameters over the
wire (Postgres rejects `PASSWORD $1` with a syntax error; ask how we know). Dollar-quoting
sidesteps manual escaping entirely rather than hand-rolling quote-doubling, which is easy to
get subtly wrong for a password an operator chose.
"""
tag = "$velodrome_pw$"
if tag in password:
raise RuntimeError(
f"password must not contain the literal sequence {tag!r} — pick a different one"
)
return f"{tag}{password}{tag}"
def upgrade() -> None:
app_password = _require_password("VELODROME_DB_APP_PASSWORD")
auth_password = _require_password("VELODROME_DB_AUTH_PASSWORD")
# --- Runtime roles -------------------------------------------------------------------
# velodrome_app: NOBYPASSRLS — every request-scoped query after identity is established.
# velodrome_auth: BYPASSRLS — ONLY auth/service.py's pre-identity lookups. See db.py.
# Roles are cluster-wide in Postgres, not per-database — if this migration ever runs
# against a second database sharing the same cluster (exactly what a local dev + test setup
# commonly looks like), a plain CREATE ROLE fails with "role already exists" even though
# THIS database has never seen the migration before. Postgres has no CREATE ROLE IF NOT
# EXISTS, so the standard idiom is catching duplicate_object in a DO block. ALTER ROLE
# afterwards keeps the password in sync with the current env var either way, rather than
# silently keeping whatever password the role happened to be created with previously.
op.execute(
f"""
DO $$
BEGIN
CREATE ROLE velodrome_app LOGIN PASSWORD {_dollar_quoted(app_password)}
NOSUPERUSER NOCREATEDB NOCREATEROLE NOREPLICATION NOBYPASSRLS;
EXCEPTION WHEN duplicate_object THEN
ALTER ROLE velodrome_app WITH LOGIN PASSWORD {_dollar_quoted(app_password)}
NOSUPERUSER NOCREATEDB NOCREATEROLE NOREPLICATION NOBYPASSRLS;
END
$$;
"""
)
op.execute(
f"""
DO $$
BEGIN
CREATE ROLE velodrome_auth LOGIN PASSWORD {_dollar_quoted(auth_password)}
NOSUPERUSER NOCREATEDB NOCREATEROLE NOREPLICATION BYPASSRLS;
EXCEPTION WHEN duplicate_object THEN
ALTER ROLE velodrome_auth WITH LOGIN PASSWORD {_dollar_quoted(auth_password)}
NOSUPERUSER NOCREATEDB NOCREATEROLE NOREPLICATION BYPASSRLS;
END
$$;
"""
)
# --- Tables ----------------------------------------------------------------------------
op.create_table(
"users",
sa.Column("id", pg.UUID(as_uuid=True), primary_key=True),
sa.Column("id", sa.Uuid(as_uuid=True), primary_key=True),
sa.Column("email", sa.String(320), nullable=False, unique=True),
sa.Column("display_name", sa.String(200), nullable=False),
sa.Column("password_hash", sa.Text(), nullable=False),
@@ -103,18 +33,16 @@ def upgrade() -> None:
sa.Column("timezone", sa.String(64), nullable=False, server_default="UTC"),
sa.Column("unit_system", sa.String(10), nullable=False, server_default="imperial"),
sa.Column("is_active", sa.Boolean(), nullable=False, server_default=sa.true()),
sa.Column(
"created_at", sa.DateTime(timezone=True), nullable=False, server_default=sa.func.now()
),
sa.Column("created_at", sa.DateTime(timezone=True), nullable=False),
)
op.create_table(
"invites",
sa.Column("id", pg.UUID(as_uuid=True), primary_key=True),
sa.Column("id", sa.Uuid(as_uuid=True), primary_key=True),
sa.Column("code_hash", sa.LargeBinary(32), nullable=False, unique=True),
sa.Column(
"created_by",
pg.UUID(as_uuid=True),
sa.Uuid(as_uuid=True),
sa.ForeignKey("users.id"),
nullable=False,
),
@@ -128,23 +56,20 @@ def upgrade() -> None:
op.create_table(
"sessions",
sa.Column("id", pg.UUID(as_uuid=True), primary_key=True),
sa.Column("id", sa.Uuid(as_uuid=True), primary_key=True),
sa.Column(
"user_id",
pg.UUID(as_uuid=True),
sa.Uuid(as_uuid=True),
sa.ForeignKey("users.id", ondelete="CASCADE"),
nullable=False,
),
sa.Column("token_hash", sa.LargeBinary(32), nullable=False, unique=True),
sa.Column("client", sa.String(20), nullable=False, server_default="web"),
sa.Column("user_agent", sa.Text(), nullable=True),
sa.Column("ip", pg.INET(), nullable=True),
sa.Column(
"created_at", sa.DateTime(timezone=True), nullable=False, server_default=sa.func.now()
),
sa.Column(
"last_seen_at", sa.DateTime(timezone=True), nullable=False, server_default=sa.func.now()
),
# Plain string — SQLite has no INET type, and nothing queries the address's structure.
sa.Column("ip", sa.String(45), nullable=True),
sa.Column("created_at", sa.DateTime(timezone=True), nullable=False),
sa.Column("last_seen_at", sa.DateTime(timezone=True), nullable=False),
sa.Column("expires_at", sa.DateTime(timezone=True), nullable=False),
sa.Column("revoked_at", sa.DateTime(timezone=True), nullable=True),
)
@@ -152,87 +77,29 @@ def upgrade() -> None:
op.create_table(
"api_tokens",
sa.Column("id", pg.UUID(as_uuid=True), primary_key=True),
sa.Column("id", sa.Uuid(as_uuid=True), primary_key=True),
sa.Column(
"user_id",
pg.UUID(as_uuid=True),
sa.Uuid(as_uuid=True),
sa.ForeignKey("users.id", ondelete="CASCADE"),
nullable=False,
),
sa.Column("name", sa.String(200), nullable=False),
sa.Column("token_hash", sa.LargeBinary(32), nullable=False, unique=True),
sa.Column(
"scopes",
pg.ARRAY(sa.String()),
nullable=False,
server_default="{}",
),
# JSON-encoded TEXT — SQLite has no array type; see models/identity.py.
sa.Column("scopes", sa.JSON(), nullable=False, server_default="[]"),
sa.Column("last_used_at", sa.DateTime(timezone=True), nullable=True),
sa.Column("expires_at", sa.DateTime(timezone=True), nullable=True),
sa.Column("revoked_at", sa.DateTime(timezone=True), nullable=True),
)
op.create_index("ix_api_tokens_user_id", "api_tokens", ["user_id"])
# --- Grants ------------------------------------------------------------------------------
# velodrome_auth needs full read/write on these four tables — it's the only thing that ever
# creates a user, a session, or redeems an invite. velodrome_app needs the same grants
# because RLS policies (below) restrict WHICH ROWS it sees, not whether the underlying
# privilege exists — GRANT and POLICY are two independent layers, both required.
# Deliberately NOT granting TRUNCATE: it's a whole-table operation that RLS cannot filter
# (Postgres RLS policies do not apply to TRUNCATE at all), so granting it to velodrome_app
# would let any bug in ordinary request-handling code wipe an entire table in one statement,
# defeating the isolation these policies exist to provide. Neither runtime role needs it —
# tests use DELETE for fixture cleanup instead (see tests/conftest.py).
for table in ("users", "invites", "sessions", "api_tokens"):
op.execute(f"GRANT SELECT, INSERT, UPDATE, DELETE ON {table} TO velodrome_app")
op.execute(f"GRANT SELECT, INSERT, UPDATE, DELETE ON {table} TO velodrome_auth")
# --- Row-level security ------------------------------------------------------------------
# Policies below apply to velodrome_app only — velodrome_auth has BYPASSRLS and ignores them
# entirely, by design (see module docstring). current_setting('app.user_id', true) returns
# NULL when unset, which makes every policy below deny-by-default for an unscoped connection.
op.execute("ALTER TABLE users ENABLE ROW LEVEL SECURITY")
op.execute(
"CREATE POLICY own_row ON users FOR ALL "
"USING (id = current_setting('app.user_id', true)::uuid) "
"WITH CHECK (id = current_setting('app.user_id', true)::uuid)"
)
op.execute("ALTER TABLE sessions ENABLE ROW LEVEL SECURITY")
op.execute(
"CREATE POLICY own_rows ON sessions FOR ALL "
"USING (user_id = current_setting('app.user_id', true)::uuid) "
"WITH CHECK (user_id = current_setting('app.user_id', true)::uuid)"
)
op.execute("ALTER TABLE api_tokens ENABLE ROW LEVEL SECURITY")
op.execute(
"CREATE POLICY own_rows ON api_tokens FOR ALL "
"USING (user_id = current_setting('app.user_id', true)::uuid) "
"WITH CHECK (user_id = current_setting('app.user_id', true)::uuid)"
)
op.execute("ALTER TABLE invites ENABLE ROW LEVEL SECURITY")
op.execute(
"CREATE POLICY own_rows ON invites FOR ALL "
"USING (created_by = current_setting('app.user_id', true)::uuid) "
"WITH CHECK (created_by = current_setting('app.user_id', true)::uuid)"
)
def downgrade() -> None:
for table in ("invites", "api_tokens", "sessions", "users"):
op.execute(f"DROP POLICY IF EXISTS own_rows ON {table}")
op.execute(f"DROP POLICY IF EXISTS own_row ON {table}")
# Children before parents (FK order).
# Children before parents (FK order) — matters even with ON DELETE CASCADE enforced (see
# db.py's PRAGMA foreign_keys=ON), since dropping a table Postgres-style doesn't rely on that
# at all; SQLite's DROP TABLE has no dependency ordering of its own to lean on.
op.drop_table("api_tokens")
op.drop_table("sessions")
op.drop_table("invites")
op.drop_table("users")
# Dropping the tables drops the GRANTs that referenced them; the roles themselves remain
# until explicitly dropped here.
op.execute("DROP ROLE IF EXISTS velodrome_app")
op.execute("DROP ROLE IF EXISTS velodrome_auth")
+1 -1
View File
@@ -7,7 +7,7 @@ dependencies = [
"fastapi>=0.115",
"uvicorn[standard]>=0.32",
"sqlalchemy[asyncio]>=2.0.35",
"asyncpg>=0.30",
"aiosqlite>=0.20",
"alembic>=1.13",
"pydantic[email]>=2.9",
"pydantic-settings>=2.6",
+13 -30
View File
@@ -1,13 +1,16 @@
"""Test fixtures.
Runs the real Alembic migrations once per session against a real Postgres (never mocked — see
CLAUDE.md's test policy), then truncates the identity tables between tests for isolation. This
deliberately exercises the exact same migration path CI's `migrations` job and a real deploy use,
not a parallel test-only schema-creation shortcut.
Runs the real Alembic migrations once per session against a real SQLite file (never mocked — see
CLAUDE.md's test policy — and never `:memory:`, which gives each separate connection its own
isolated database rather than one shared one), then deletes all rows between tests for isolation.
This deliberately exercises the exact same migration path CI's `migrations` job and a real deploy
use, not a parallel test-only schema-creation shortcut.
"""
import os
import tempfile
from collections.abc import AsyncIterator
from pathlib import Path
import pytest
from alembic.config import Config
@@ -18,29 +21,12 @@ from alembic import command
os.environ.setdefault("VELODROME_ENVIRONMENT", "test")
# Test-only role passwords. Never used outside this process; the migration requires them to be
# set explicitly (see alembic/versions/0001_baseline.py::_require_password) rather than default
# to anything, on purpose — that's the same rule for a real deploy, just satisfied differently.
_APP_PW = "test-only-app-password"
_AUTH_PW = "test-only-auth-password"
_tmp_dir = tempfile.mkdtemp(prefix="velodrome-test-")
_db_path = Path(_tmp_dir) / "test.db"
os.environ.setdefault("VELODROME_DATABASE_URL", f"sqlite+aiosqlite:///{_db_path}")
os.environ.setdefault(
"VELODROME_DATABASE_URL_MIGRATE",
"postgresql+asyncpg://postgres:postgres@localhost:5432/velodrome_test",
)
os.environ.setdefault("VELODROME_DB_APP_PASSWORD", _APP_PW)
os.environ.setdefault("VELODROME_DB_AUTH_PASSWORD", _AUTH_PW)
os.environ.setdefault(
"VELODROME_DATABASE_URL_APP",
f"postgresql+asyncpg://velodrome_app:{_APP_PW}@localhost:5432/velodrome_test",
)
os.environ.setdefault(
"VELODROME_DATABASE_URL_AUTH",
f"postgresql+asyncpg://velodrome_auth:{_AUTH_PW}@localhost:5432/velodrome_test",
)
# Settings/engines must not be constructed before the env vars above are set, so these imports
# are deliberately below the os.environ.setdefault block, not at module top.
# Settings/engines must not be constructed before the env var above is set, so these imports are
# deliberately below it, not at module top.
import httpx # noqa: E402
from velodrome.app import app # noqa: E402
@@ -51,7 +37,6 @@ from velodrome.db import auth_session # noqa: E402
def _run_migrations() -> None:
cfg = Config(os.path.join(os.path.dirname(__file__), "..", "alembic.ini"))
cfg.set_main_option("script_location", os.path.join(os.path.dirname(__file__), "..", "alembic"))
command.downgrade(cfg, "base") # clean slate even on a reused test database
command.upgrade(cfg, "head")
@@ -60,9 +45,7 @@ async def _clean_tables() -> AsyncIterator[None]:
yield
async with auth_session() as db:
async with db.begin():
# DELETE, not TRUNCATE — velodrome_auth deliberately isn't granted TRUNCATE in
# production (see the migration's comment), and using the same privilege level in
# tests as in prod is the point. Children before parents for the FK constraints.
# Children before parents for the FK constraints.
for table in ("sessions", "api_tokens", "invites", "users"):
await db.execute(text(f"DELETE FROM {table}"))
+107 -41
View File
@@ -1,44 +1,73 @@
"""Auth endpoint tests, plus — the one that matters most — an empirical proof of RLS isolation.
"""Auth endpoint tests, plus — the one that matters most — an empirical proof that the
repository-layer scope actually enforces isolation between users.
Per CLAUDE.md: 'auth/, RLS policies — security, and a mistake exposes another user's data.' The
whole point of test_rls_blocks_cross_user_session_reads below is that it doesn't trust the SQL in
the migration is correct because it reads correctly it proves it by actually trying to read
another user's row through the RLS-scoped role and confirming zero rows come back.
Per CLAUDE.md invariant #4: since D15 removed Postgres RLS, `Scope` in db.py is not one layer of
isolation among two — it's the only one. test_scoped_session_blocks_cross_user_reads below doesn't
trust that `Scope.select()` filters correctly because the code reads correctly; it proves it by
registering two real users and confirming a scoped read for user A never returns user B's row,
even when both exist in the same table.
"""
from datetime import UTC, datetime, timedelta
from uuid import UUID
from uuid import UUID, uuid4
import httpx
from sqlalchemy import text
import pytest
from sqlalchemy import select, text
from sqlalchemy.ext.asyncio import AsyncSession
from velodrome.auth.security import hash_invite_code, hash_password
from velodrome.db import scoped_session, unscoped_session
from velodrome.models import Invite, Session
_REGISTER_PASSWORD = "correct horse battery staple"
async def _seed_invite(db_auth: AsyncSession, *, code: str = "TESTCODE123") -> UUID:
"""Insert a usable invite directly, bypassing the API — this is fixture setup, not the
thing under test."""
thing under test.
UUID values are passed as `.hex` (32 hex chars, no hyphens), not `str()` (36 chars, hyphenated)
or a raw UUID object — a `text()` query has no ORM-level column-type awareness, so this is
fixture setup working around two separate things confirmed empirically, not assumed:
(1) aiosqlite's driver has no built-in adapter for a Python `UUID` object at all (raises
"type 'UUID' is not supported"); (2) SQLAlchemy's `Uuid(as_uuid=True)` column type stores as
the 32-char hex form on SQLite, NOT the hyphenated `str()` form — inserting the hyphenated
form via raw SQL left rows the ORM's own later `UPDATE ... WHERE id = ?` (implicitly issued
by `invite.used_count += 1` in auth/service.py) could never match, silently updating zero
rows. `.hex` is what the ORM itself writes, so raw-SQL-inserted rows are indistinguishable
from ORM-inserted ones.
`creator_id` is a fixed sentinel, not a fresh `uuid4()` per call — confirmed the hard way: a
test calling this twice (two invite codes) with a fresh id each time but the same hardcoded
seed email hit the `ON CONFLICT DO NOTHING` on email on the second call, which silently
no-ops, leaving that call's fresh id never actually inserted — so its invite's `created_by`
pointed at a user row that was never created, and the FK constraint on the invites insert
failed. A stable id makes repeat calls genuinely idempotent instead of just quiet about it.
"""
# datetime binds use .isoformat() explicitly, not a raw datetime object — Python's sqlite3
# module has its own implicit datetime adapter, which is deprecated as of 3.12 and warns on
# every use (confirmed against this exact fixture); the ORM's own UTCDateTime type never hits
# this because SQLAlchemy's dialect handles the conversion itself rather than delegating to
# sqlite3's adapter, but a bare text() bind param has no such handling.
creator_id = UUID(int=0)
await db_auth.execute(
text(
"INSERT INTO users (id, email, display_name, password_hash) "
"VALUES (:id, 'seed@example.com', 'Seed', :ph) ON CONFLICT DO NOTHING"
"INSERT INTO users (id, email, display_name, password_hash, created_at) "
"VALUES (:id, 'seed@example.com', 'Seed', :ph, :now) ON CONFLICT DO NOTHING"
),
{"id": str(creator_id), "ph": hash_password("unused")},
{"id": creator_id.hex, "ph": hash_password("unused"), "now": datetime.now(UTC).isoformat()},
)
await db_auth.execute(
text(
"INSERT INTO invites (id, code_hash, created_by, expires_at, max_uses, used_count) "
"VALUES (gen_random_uuid(), :hash, :creator, :expires, 1, 0)"
"VALUES (:id, :hash, :creator, :expires, 1, 0)"
),
{
"id": uuid4().hex,
"hash": hash_invite_code(code),
"creator": str(creator_id),
"expires": datetime.now(UTC) + timedelta(days=1),
"creator": creator_id.hex,
"expires": (datetime.now(UTC) + timedelta(days=1)).isoformat(),
},
)
await db_auth.commit()
@@ -195,17 +224,18 @@ async def test_logout_without_matching_origin_is_rejected(
assert resp.status_code == 403
async def test_rls_blocks_cross_user_session_reads(
async def test_scoped_session_blocks_cross_user_reads(
client: httpx.AsyncClient, db_auth: AsyncSession
) -> None:
"""The load-bearing test: prove RLS actually enforces isolation, not just that the migration
ran without a syntax error.
"""The load-bearing test: prove the repository-layer scope actually enforces isolation, not
just that `Scope.select()` reads like it should.
Two users register (creating two session rows). We then open a `scoped_session` as user A —
the exact code path every real request handler will use once Phase 1 adds user-owned domain
tables — and confirm the raw SQL result set contains ONLY user A's session, even though it
runs no WHERE clause on user_id at all. If this test ever passes with more than one row, RLS
is not doing its job and every other invariant in this codebase is resting on nothing.
Two users register (creating two session rows). We then open `scoped_session` as user A — the
exact code path every real request handler will use once Phase 1 adds user-owned domain
tables — and confirm `Scope.select(Session)` returns ONLY user A's row, even though both exist
in the same table. If this test ever passes with more than one row, the scope is not isolating
users and every other invariant in this codebase is resting on nothing (see CLAUDE.md
invariant #4 — there is no database-level backstop anymore; this IS the isolation boundary).
"""
await _seed_invite(db_auth, code="CODE-FOR-A")
await _seed_invite(db_auth, code="CODE-FOR-B")
@@ -233,43 +263,79 @@ async def test_rls_blocks_cross_user_session_reads(
user_b_id = UUID(resp_b.json()["id"])
assert user_a_id != user_b_id
# Sanity check first: BOTH sessions genuinely exist, via the bypass role.
# Sanity check first: BOTH sessions genuinely exist, via the pre-identity bootstrap path.
total = (await db_auth.execute(text("SELECT count(*) FROM sessions"))).scalar_one()
assert total == 2
# Close out the transaction this read just autobegan. db.py's BEGIN IMMEDIATE (see its
# docstring) takes SQLite's exclusive write lock the instant ANY transaction starts, read
# included — confirmed the hard way: without this commit, the db_auth fixture's session
# (which the pytest fixture keeps open for the whole test, not just this block) holds that
# lock indefinitely, and the scoped_session() calls below then fail with "database is
# locked" trying to acquire their own. Every write in this file already commits promptly;
# a read needs the same discipline once BEGIN IMMEDIATE is in play.
await db_auth.commit()
# Now the real test: as user A, scoped through the RLS-subject role, with NO WHERE clause.
async with scoped_session(user_a_id) as scoped_db:
rows = (await scoped_db.execute(text("SELECT user_id FROM sessions"))).scalars().all()
# Now the real test: Scope.select(), which is the only way application code is meant to
# query a user-owned table once identity is known.
async with scoped_session(user_a_id) as scope:
rows = (await scope.session.execute(scope.select(Session))).scalars().all()
assert len(rows) == 1, (
f"expected exactly 1 row (user A's own session) via RLS, got {len(rows)}"
"RLS is not isolating users"
f"expected exactly 1 row (user A's own session), got {len(rows)}"
"the repository-layer scope is not isolating users"
)
assert UUID(str(rows[0])) == user_a_id
assert rows[0].user_id == user_a_id
# And the mirror image, as user B, proving this isn't a coincidence of row ordering.
async with scoped_session(user_b_id) as scoped_db:
rows_b = (await scoped_db.execute(text("SELECT user_id FROM sessions"))).scalars().all()
async with scoped_session(user_b_id) as scope:
rows_b = (await scope.session.execute(scope.select(Session))).scalars().all()
assert len(rows_b) == 1
assert UUID(str(rows_b[0])) == user_b_id
assert rows_b[0].user_id == user_b_id
async def test_unscoped_session_sees_zero_rows_of_user_owned_tables(
async def test_scope_select_rejects_models_without_user_id() -> None:
"""`Scope.select()` refuses to build a query for a model that has no `user_id` column,
instead of silently returning an unfiltered (and therefore unsafe) query. `Invite` is a real
example, not a contrived one: it's scoped by `created_by`, not `user_id`."""
async with scoped_session(uuid4()) as scope:
with pytest.raises(TypeError, match="has no user_id column"):
scope.select(Invite)
async def test_unscoped_session_can_see_every_user_when_misused(
client: httpx.AsyncClient, db_auth: AsyncSession
) -> None:
"""The default-deny half of the same proof: with app.user_id unset entirely (the state the
health check and any other unauthenticated code path runs in), RLS denies everything."""
await _seed_invite(db_auth)
"""The deliberately alarming counterpart to the isolation test above: `unscoped_session()`
plus a raw `select(Session)` returns EVERY user's rows, with no filtering at all — there is no
database-level backstop to catch this mistake anymore (docs/DECISIONS.md D15). This is exactly
why CLAUDE.md invariant #4 treats reaching for `unscoped_session()` on a user-owned table as a
review-blocking mistake, not a style preference: the code review IS the isolation boundary
for this specific failure mode, `Scope` is the boundary for the query-construction one.
"""
await _seed_invite(db_auth, code="CODE-FOR-A")
await _seed_invite(db_auth, code="CODE-FOR-B")
await client.post(
"/api/v1/auth/register",
json={
"email": "rider@example.com",
"email": "user-a@example.com",
"password": _REGISTER_PASSWORD,
"display_name": "Rider",
"invite_code": "TESTCODE123",
"display_name": "User A",
"invite_code": "CODE-FOR-A",
},
)
await client.post(
"/api/v1/auth/register",
json={
"email": "user-b@example.com",
"password": _REGISTER_PASSWORD,
"display_name": "User B",
"invite_code": "CODE-FOR-B",
},
)
async with unscoped_session() as db:
rows = (await db.execute(text("SELECT * FROM sessions"))).all()
assert rows == []
rows = (await db.execute(select(Session))).scalars().all()
assert len(rows) == 2, (
"unscoped_session with a raw query sees every user's rows — by design, this is the "
"unsafe path Scope exists to replace"
)
+11 -42
View File
@@ -7,6 +7,15 @@ resolution-markers = [
"python_full_version < '3.14'",
]
[[package]]
name = "aiosqlite"
version = "0.22.1"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/4e/8a/64761f4005f17809769d23e518d915db74e6310474e733e3593cfc854ef1/aiosqlite-0.22.1.tar.gz", hash = "sha256:043e0bd78d32888c0a9ca90fc788b38796843360c855a7262a532813133a0650", size = 14821, upload-time = "2025-12-23T19:25:43.997Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/00/b7/e3bf5133d697a08128598c8d0abc5e16377b51465a33756de24fa7dee953/aiosqlite-0.22.1-py3-none-any.whl", hash = "sha256:21c002eb13823fad740196c5a2e9d8e62f6243bd9e7e4a1f87fb5e44ecb4fceb", size = 17405, upload-time = "2025-12-23T19:25:42.139Z" },
]
[[package]]
name = "alembic"
version = "1.20.0"
@@ -171,46 +180,6 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/91/a7/c8bbb2173f7a7131b3b2412035b2d814ab5ef2ce9799bd06f07c451640e4/ast_serialize-0.11.2-cp39-abi3-win_arm64.whl", hash = "sha256:dab599cbdcb7b45b18c41fad746645580b3a24357082b7f0e8921cd373804f27", size = 1136031, upload-time = "2026-09-13T18:48:54.04Z" },
]
[[package]]
name = "asyncpg"
version = "0.31.0"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/fe/cc/d18065ce2380d80b1bcce927c24a2642efd38918e33fd724bc4bca904877/asyncpg-0.31.0.tar.gz", hash = "sha256:c989386c83940bfbd787180f2b1519415e2d3d6277a70d9d0f0145ac73500735", size = 993667, upload-time = "2025-11-24T23:27:00.812Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/2a/a6/59d0a146e61d20e18db7396583242e32e0f120693b67a8de43f1557033e2/asyncpg-0.31.0-cp312-cp312-macosx_10_13_x86_64.whl", hash = "sha256:b44c31e1efc1c15188ef183f287c728e2046abb1d26af4d20858215d50d91fad", size = 662042, upload-time = "2025-11-24T23:25:49.578Z" },
{ url = "https://files.pythonhosted.org/packages/36/01/ffaa189dcb63a2471720615e60185c3f6327716fdc0fc04334436fbb7c65/asyncpg-0.31.0-cp312-cp312-macosx_11_0_arm64.whl", hash = "sha256:0c89ccf741c067614c9b5fc7f1fc6f3b61ab05ae4aaa966e6fd6b93097c7d20d", size = 638504, upload-time = "2025-11-24T23:25:51.501Z" },
{ url = "https://files.pythonhosted.org/packages/9f/62/3f699ba45d8bd24c5d65392190d19656d74ff0185f42e19d0bbd973bb371/asyncpg-0.31.0-cp312-cp312-manylinux_2_28_aarch64.whl", hash = "sha256:12b3b2e39dc5470abd5e98c8d3373e4b1d1234d9fbdedf538798b2c13c64460a", size = 3426241, upload-time = "2025-11-24T23:25:53.278Z" },
{ url = "https://files.pythonhosted.org/packages/8c/d1/a867c2150f9c6e7af6462637f613ba67f78a314b00db220cd26ff559d532/asyncpg-0.31.0-cp312-cp312-manylinux_2_28_x86_64.whl", hash = "sha256:aad7a33913fb8bcb5454313377cc330fbb19a0cd5faa7272407d8a0c4257b671", size = 3520321, upload-time = "2025-11-24T23:25:54.982Z" },
{ url = "https://files.pythonhosted.org/packages/7a/1a/cce4c3f246805ecd285a3591222a2611141f1669d002163abef999b60f98/asyncpg-0.31.0-cp312-cp312-musllinux_1_2_aarch64.whl", hash = "sha256:3df118d94f46d85b2e434fd62c84cb66d5834d5a890725fe625f498e72e4d5ec", size = 3316685, upload-time = "2025-11-24T23:25:57.43Z" },
{ url = "https://files.pythonhosted.org/packages/40/ae/0fc961179e78cc579e138fad6eb580448ecae64908f95b8cb8ee2f241f67/asyncpg-0.31.0-cp312-cp312-musllinux_1_2_x86_64.whl", hash = "sha256:bd5b6efff3c17c3202d4b37189969acf8927438a238c6257f66be3c426beba20", size = 3471858, upload-time = "2025-11-24T23:25:59.636Z" },
{ url = "https://files.pythonhosted.org/packages/52/b2/b20e09670be031afa4cbfabd645caece7f85ec62d69c312239de568e058e/asyncpg-0.31.0-cp312-cp312-win32.whl", hash = "sha256:027eaa61361ec735926566f995d959ade4796f6a49d3bde17e5134b9964f9ba8", size = 527852, upload-time = "2025-11-24T23:26:01.084Z" },
{ url = "https://files.pythonhosted.org/packages/b5/f0/f2ed1de154e15b107dc692262395b3c17fc34eafe2a78fc2115931561730/asyncpg-0.31.0-cp312-cp312-win_amd64.whl", hash = "sha256:72d6bdcbc93d608a1158f17932de2321f68b1a967a13e014998db87a72ed3186", size = 597175, upload-time = "2025-11-24T23:26:02.564Z" },
{ url = "https://files.pythonhosted.org/packages/95/11/97b5c2af72a5d0b9bc3fa30cd4b9ce22284a9a943a150fdc768763caf035/asyncpg-0.31.0-cp313-cp313-macosx_10_13_x86_64.whl", hash = "sha256:c204fab1b91e08b0f47e90a75d1b3c62174dab21f670ad6c5d0f243a228f015b", size = 661111, upload-time = "2025-11-24T23:26:04.467Z" },
{ url = "https://files.pythonhosted.org/packages/1b/71/157d611c791a5e2d0423f09f027bd499935f0906e0c2a416ce712ba51ef3/asyncpg-0.31.0-cp313-cp313-macosx_11_0_arm64.whl", hash = "sha256:54a64f91839ba59008eccf7aad2e93d6e3de688d796f35803235ea1c4898ae1e", size = 636928, upload-time = "2025-11-24T23:26:05.944Z" },
{ url = "https://files.pythonhosted.org/packages/2e/fc/9e3486fb2bbe69d4a867c0b76d68542650a7ff1574ca40e84c3111bb0c6e/asyncpg-0.31.0-cp313-cp313-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:c0e0822b1038dc7253b337b0f3f676cadc4ac31b126c5d42691c39691962e403", size = 3424067, upload-time = "2025-11-24T23:26:07.957Z" },
{ url = "https://files.pythonhosted.org/packages/12/c6/8c9d076f73f07f995013c791e018a1cd5f31823c2a3187fc8581706aa00f/asyncpg-0.31.0-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:bef056aa502ee34204c161c72ca1f3c274917596877f825968368b2c33f585f4", size = 3518156, upload-time = "2025-11-24T23:26:09.591Z" },
{ url = "https://files.pythonhosted.org/packages/ae/3b/60683a0baf50fbc546499cfb53132cb6835b92b529a05f6a81471ab60d0c/asyncpg-0.31.0-cp313-cp313-musllinux_1_2_aarch64.whl", hash = "sha256:0bfbcc5b7ffcd9b75ab1558f00db2ae07db9c80637ad1b2469c43df79d7a5ae2", size = 3319636, upload-time = "2025-11-24T23:26:11.168Z" },
{ url = "https://files.pythonhosted.org/packages/50/dc/8487df0f69bd398a61e1792b3cba0e47477f214eff085ba0efa7eac9ce87/asyncpg-0.31.0-cp313-cp313-musllinux_1_2_x86_64.whl", hash = "sha256:22bc525ebbdc24d1261ecbf6f504998244d4e3be1721784b5f64664d61fbe602", size = 3472079, upload-time = "2025-11-24T23:26:13.164Z" },
{ url = "https://files.pythonhosted.org/packages/13/a1/c5bbeeb8531c05c89135cb8b28575ac2fac618bcb60119ee9696c3faf71c/asyncpg-0.31.0-cp313-cp313-win32.whl", hash = "sha256:f890de5e1e4f7e14023619399a471ce4b71f5418cd67a51853b9910fdfa73696", size = 527606, upload-time = "2025-11-24T23:26:14.78Z" },
{ url = "https://files.pythonhosted.org/packages/91/66/b25ccb84a246b470eb943b0107c07edcae51804912b824054b3413995a10/asyncpg-0.31.0-cp313-cp313-win_amd64.whl", hash = "sha256:dc5f2fa9916f292e5c5c8b2ac2813763bcd7f58e130055b4ad8a0531314201ab", size = 596569, upload-time = "2025-11-24T23:26:16.189Z" },
{ url = "https://files.pythonhosted.org/packages/3c/36/e9450d62e84a13aea6580c83a47a437f26c7ca6fa0f0fd40b6670793ea30/asyncpg-0.31.0-cp314-cp314-macosx_10_15_x86_64.whl", hash = "sha256:f6b56b91bb0ffc328c4e3ed113136cddd9deefdf5f79ab448598b9772831df44", size = 660867, upload-time = "2025-11-24T23:26:17.631Z" },
{ url = "https://files.pythonhosted.org/packages/82/4b/1d0a2b33b3102d210439338e1beea616a6122267c0df459ff0265cd5807a/asyncpg-0.31.0-cp314-cp314-macosx_11_0_arm64.whl", hash = "sha256:334dec28cf20d7f5bb9e45b39546ddf247f8042a690bff9b9573d00086e69cb5", size = 638349, upload-time = "2025-11-24T23:26:19.689Z" },
{ url = "https://files.pythonhosted.org/packages/41/aa/e7f7ac9a7974f08eff9183e392b2d62516f90412686532d27e196c0f0eeb/asyncpg-0.31.0-cp314-cp314-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:98cc158c53f46de7bb677fd20c417e264fc02b36d901cc2a43bd6cb0dc6dbfd2", size = 3410428, upload-time = "2025-11-24T23:26:21.275Z" },
{ url = "https://files.pythonhosted.org/packages/6f/de/bf1b60de3dede5c2731e6788617a512bc0ebd9693eac297ee74086f101d7/asyncpg-0.31.0-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:9322b563e2661a52e3cdbc93eed3be7748b289f792e0011cb2720d278b366ce2", size = 3471678, upload-time = "2025-11-24T23:26:23.627Z" },
{ url = "https://files.pythonhosted.org/packages/46/78/fc3ade003e22d8bd53aaf8f75f4be48f0b460fa73738f0391b9c856a9147/asyncpg-0.31.0-cp314-cp314-musllinux_1_2_aarch64.whl", hash = "sha256:19857a358fc811d82227449b7ca40afb46e75b33eb8897240c3839dd8b744218", size = 3313505, upload-time = "2025-11-24T23:26:25.235Z" },
{ url = "https://files.pythonhosted.org/packages/bf/e9/73eb8a6789e927816f4705291be21f2225687bfa97321e40cd23055e903a/asyncpg-0.31.0-cp314-cp314-musllinux_1_2_x86_64.whl", hash = "sha256:ba5f8886e850882ff2c2ace5732300e99193823e8107e2c53ef01c1ebfa1e85d", size = 3434744, upload-time = "2025-11-24T23:26:26.944Z" },
{ url = "https://files.pythonhosted.org/packages/08/4b/f10b880534413c65c5b5862f79b8e81553a8f364e5238832ad4c0af71b7f/asyncpg-0.31.0-cp314-cp314-win32.whl", hash = "sha256:cea3a0b2a14f95834cee29432e4ddc399b95700eb1d51bbc5bfee8f31fa07b2b", size = 532251, upload-time = "2025-11-24T23:26:28.404Z" },
{ url = "https://files.pythonhosted.org/packages/d3/2d/7aa40750b7a19efa5d66e67fc06008ca0f27ba1bd082e457ad82f59aba49/asyncpg-0.31.0-cp314-cp314-win_amd64.whl", hash = "sha256:04d19392716af6b029411a0264d92093b6e5e8285ae97a39957b9a9c14ea72be", size = 604901, upload-time = "2025-11-24T23:26:30.34Z" },
{ url = "https://files.pythonhosted.org/packages/ce/fe/b9dfe349b83b9dee28cc42360d2c86b2cdce4cb551a2c2d27e156bcac84d/asyncpg-0.31.0-cp314-cp314t-macosx_10_15_x86_64.whl", hash = "sha256:bdb957706da132e982cc6856bb2f7b740603472b54c3ebc77fe60ea3e57e1bd2", size = 702280, upload-time = "2025-11-24T23:26:32Z" },
{ url = "https://files.pythonhosted.org/packages/6a/81/e6be6e37e560bd91e6c23ea8a6138a04fd057b08cf63d3c5055c98e81c1d/asyncpg-0.31.0-cp314-cp314t-macosx_11_0_arm64.whl", hash = "sha256:6d11b198111a72f47154fa03b85799f9be63701e068b43f84ac25da0bda9cb31", size = 682931, upload-time = "2025-11-24T23:26:33.572Z" },
{ url = "https://files.pythonhosted.org/packages/a6/45/6009040da85a1648dd5bc75b3b0a062081c483e75a1a29041ae63a0bf0dc/asyncpg-0.31.0-cp314-cp314t-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl", hash = "sha256:18c83b03bc0d1b23e6230f5bf8d4f217dc9bc08644ce0502a9d91dc9e634a9c7", size = 3581608, upload-time = "2025-11-24T23:26:35.638Z" },
{ url = "https://files.pythonhosted.org/packages/7e/06/2e3d4d7608b0b2b3adbee0d0bd6a2d29ca0fc4d8a78f8277df04e2d1fd7b/asyncpg-0.31.0-cp314-cp314t-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl", hash = "sha256:e009abc333464ff18b8f6fd146addffd9aaf63e79aa3bb40ab7a4c332d0c5e9e", size = 3498738, upload-time = "2025-11-24T23:26:37.275Z" },
{ url = "https://files.pythonhosted.org/packages/7d/aa/7d75ede780033141c51d83577ea23236ba7d3a23593929b32b49db8ed36e/asyncpg-0.31.0-cp314-cp314t-musllinux_1_2_aarch64.whl", hash = "sha256:3b1fbcb0e396a5ca435a8826a87e5c2c2cc0c8c68eb6fadf82168056b0e53a8c", size = 3401026, upload-time = "2025-11-24T23:26:39.423Z" },
{ url = "https://files.pythonhosted.org/packages/ba/7a/15e37d45e7f7c94facc1e9148c0e455e8f33c08f0b8a0b1deb2c5171771b/asyncpg-0.31.0-cp314-cp314t-musllinux_1_2_x86_64.whl", hash = "sha256:8df714dba348efcc162d2adf02d213e5fab1bd9f557e1305633e851a61814a7a", size = 3429426, upload-time = "2025-11-24T23:26:41.032Z" },
{ url = "https://files.pythonhosted.org/packages/13/d5/71437c5f6ae5f307828710efbe62163974e71237d5d46ebd2869ea052d10/asyncpg-0.31.0-cp314-cp314t-win32.whl", hash = "sha256:1b41f1afb1033f2b44f3234993b15096ddc9cd71b21a42dbd87fc6a57b43d65d", size = 614495, upload-time = "2025-11-24T23:26:42.659Z" },
{ url = "https://files.pythonhosted.org/packages/3c/d7/8fb3044eaef08a310acfe23dae9a8e2e07d305edc29a53497e52bc76eca7/asyncpg-0.31.0-cp314-cp314t-win_amd64.whl", hash = "sha256:bd4107bb7cdd0e9e65fae66a62afd3a249663b844fa34d479f6d5b3bef9c04c3", size = 706062, upload-time = "2025-11-24T23:26:44.086Z" },
]
[[package]]
name = "certifi"
version = "2026.7.22"
@@ -1170,9 +1139,9 @@ name = "velodrome"
version = "0.1.0"
source = { editable = "." }
dependencies = [
{ name = "aiosqlite" },
{ name = "alembic" },
{ name = "argon2-cffi" },
{ name = "asyncpg" },
{ name = "fastapi" },
{ name = "pydantic", extra = ["email"] },
{ name = "pydantic-settings" },
@@ -1192,9 +1161,9 @@ dev = [
[package.metadata]
requires-dist = [
{ name = "aiosqlite", specifier = ">=0.20" },
{ name = "alembic", specifier = ">=1.13" },
{ name = "argon2-cffi", specifier = ">=23.1" },
{ name = "asyncpg", specifier = ">=0.30" },
{ name = "fastapi", specifier = ">=0.115" },
{ name = "httpx", marker = "extra == 'dev'", specifier = ">=0.27" },
{ name = "mypy", marker = "extra == 'dev'", specifier = ">=1.13" },
+2 -2
View File
@@ -10,8 +10,8 @@ router = APIRouter(tags=["health"])
async def healthz() -> dict[str, str]:
"""Liveness + a real database round trip.
Not RLS-scoped (there's no user yet at this point) — see db.unscoped_session's docstring for
why that's safe: it can see zero rows of any user-owned table regardless.
Uses `unscoped_session()`, not `scoped_session()` — there's no user yet at this point, and
`SELECT 1` never touches a user-owned table anyway. See db.py's module docstring.
"""
async with unscoped_session() as db:
await db.execute(text("SELECT 1"))
+12 -10
View File
@@ -1,9 +1,8 @@
"""Auth bootstrap logic: register, login, session validation, logout.
Every function here runs against the BYPASSRLS `auth` database role (see db.py's module
docstring for why) and every query is an exact match on a unique key — email, token_hash, or
code_hash — never an unfiltered scan. That's what makes bypassing RLS safe here: there's no
"list everything" code path for these functions to accidentally expose.
Every function here runs against `db.auth_session()` (see its docstring for why these specific
pre-identity lookups need it) and every query is an exact match on a unique key — email,
token_hash, or code_hash — never an unfiltered scan.
Nothing outside this module should import `db.auth_session` — if a new feature needs it, that's
a sign the feature belongs here, not that the import should spread.
@@ -59,17 +58,20 @@ async def register(
) -> AuthenticatedSession:
"""Validate an invite and create a user, atomically.
Open signup does not exist — see docs/PLAN.md "Auth". The `SELECT ... FOR UPDATE` on the
invite row is what stops a shared invite link being redeemed twice concurrently; without it,
two requests could both read `used_count < max_uses` as true before either commits.
Open signup does not exist — see docs/PLAN.md "Auth". What stops a shared invite link being
redeemed twice concurrently is db.py's `BEGIN IMMEDIATE` setup, not a row lock on this
SELECT — SQLite has no `SELECT ... FOR UPDATE` (SQLAlchemy's SQLite dialect silently no-ops
`.with_for_update()`, confirmed empirically; it used to appear here when this ran against
Postgres — see git history). `BEGIN IMMEDIATE` takes SQLite's write lock for the whole
transaction up front, so two concurrent redemptions can't both read `used_count < max_uses`
as true before either commits — the second one simply waits for the first transaction to
finish, then sees the incremented count.
"""
code_hash = hash_invite_code(invite_code)
async with auth_session() as db:
async with db.begin():
invite = (
await db.execute(
select(Invite).where(Invite.code_hash == code_hash).with_for_update()
)
await db.execute(select(Invite).where(Invite.code_hash == code_hash))
).scalar_one_or_none()
if invite is None:
+8 -9
View File
@@ -1,24 +1,23 @@
"""Application settings, read from environment variables.
Two separate database DSNs are deliberate, not an oversight — see db.py for why: one connects
as a role with BYPASSRLS (used only by the auth bootstrap path, which must look identity up
*before* it can be scoped), the other as a normal RLS-subject role (used for every other query).
Single SQLite database (see docs/DECISIONS.md D15 for why this isn't the two-role Postgres+RLS
setup Phase 0 originally shipped with) — one DSN, one engine, no BYPASSRLS/NOBYPASSRLS split.
Isolation between users is enforced entirely by the repository-layer scope now; see db.py and
CLAUDE.md's invariant #4.
"""
from functools import lru_cache
from pydantic import PostgresDsn
from pydantic_settings import BaseSettings, SettingsConfigDict
class Settings(BaseSettings):
model_config = SettingsConfigDict(env_prefix="VELODROME_", extra="ignore")
# Scoped role: NOBYPASSRLS. Every request-handling query outside the auth bootstrap uses this.
database_url_app: PostgresDsn
# Bootstrap role: BYPASSRLS. Used ONLY by velodrome.auth.service for the pre-identity lookups
# (login by email, session-by-token, invite-by-code) and for creating new users/sessions.
database_url_auth: PostgresDsn
# A SQLAlchemy URL, e.g. sqlite+aiosqlite:////data/velodrome.db. Not typed as a stricter DSN
# (pydantic has no built-in sqlite+aiosqlite validator) — Alembic and the app both read this
# same setting, so keep it a plain string rather than inventing a validator neither needs.
database_url: str = "sqlite+aiosqlite:///./velodrome.db"
# AES-GCM key for encrypting third-party credentials (e.g. the Bryton digest, added in a later
# phase). Not used yet in Phase 0, but declared now so the settings shape is stable.
+107 -77
View File
@@ -1,113 +1,143 @@
"""Two database engines, on purpose.
"""Single-engine SQLite access.
`app` engine: connects as a role with RLS enforced (NOBYPASSRLS). Every request handler that has
already established who the caller is uses this, wrapped in `scoped_session()` below, which sets
`app.user_id` for the transaction so RLS policies can key off it.
Phase 0 originally ran two Postgres roles with row-level security as a database-enforced
isolation layer (`velodrome_app`/`velodrome_auth` — see docs/DECISIONS.md D4). D15 moved the
database to SQLite, which has no roles, no session variables, and no policy engine — there is no
database-enforced layer left. Isolation between users now rests entirely on the query builder
below. CLAUDE.md's invariant #4 treats this file as load-bearing, not a convenience wrapper: a
new user-owned table without a passing isolation test (see tests/test_auth.py's pattern) is not
done.
`auth` engine: connects as a role with BYPASSRLS. Used ONLY by velodrome.auth.service, and only
for the narrow set of queries that must run *before* identity is known — looking a user up by
email at login, a session up by its hashed token, an invite up by its hashed code — plus the
inserts that create those rows in the first place. Nothing outside auth/service.py should import
this engine; if you find yourself reaching for it elsewhere, the query almost certainly belongs
in a repository method on the scoped session instead (see CLAUDE.md invariant #4).
Two access patterns:
- `scoped_session(user_id)` — for every query against a user-owned table once identity is known.
Yields a `Scope`, whose `select()` is the ONLY way to build a query through it — every query it
builds is pre-filtered to that user_id, on any model that declares a `user_id` column. There is
no method on `Scope` that returns an unfiltered query. This is what makes "forgot the WHERE
clause" structurally harder than remembering to write one by hand.
- `auth_session()` — for velodrome.auth.service ONLY: the handful of pre-identity lookups (login
by email, a session by its token hash, an invite by its code hash) that by definition can't be
scoped to a user_id nobody has established yet. Every query on this session must be an exact
match on a unique key, never an unfiltered scan — that discipline is what made bypassing RLS
safe before, and it's what keeps this safe now that RLS is gone. Nothing outside auth/service.py
should import this.
"""
from collections.abc import AsyncIterator
from contextlib import asynccontextmanager
from typing import Any
from uuid import UUID
from sqlalchemy import text
from sqlalchemy import Select, event, select
from sqlalchemy.ext.asyncio import (
AsyncEngine,
AsyncSession,
async_sessionmaker,
create_async_engine,
)
from sqlalchemy.orm import DeclarativeBase
from sqlalchemy.pool import NullPool
from velodrome.config import get_settings
def _make_engine(url: str) -> AsyncEngine:
settings = get_settings()
# NullPool in dev/test keeps behaviour predictable across the pytest-asyncio event loop;
# production tuning (pool_size etc.) is a deploy-time concern, not a Phase-0 one.
return create_async_engine(
url,
poolclass=NullPool if settings.environment == "test" else None,
echo=False,
)
def _configure_sqlite_for_concurrent_writers(eng: AsyncEngine) -> None:
"""Two SQLite defaults that silently do the wrong thing if left alone — confirmed empirically
against a real aiosqlite connection, not assumed from docs:
1. Foreign key enforcement, `ON DELETE CASCADE` included, is OFF by default per connection.
Without `PRAGMA foreign_keys=ON`, deleting a user leaves its sessions/api_tokens rows
behind instead of cascading — verified this happens silently, no error either way.
2. pysqlite/aiosqlite default to a DEFERRED transaction, which only takes a write lock on the
first actual write statement — leaving a real check-then-act race window. Concretely: two
concurrent redemptions of the same invite code could both read `used_count < max_uses` as
true before either commits, double-spending a single-use invite. Disable the driver's own
implicit BEGIN handling and issue `BEGIN IMMEDIATE` ourselves instead, which takes the
write lock at transaction start and correctly serializes writers — SQLAlchemy's own
documented recipe for this, not a workaround improvised here.
"""
@event.listens_for(eng.sync_engine, "connect")
def _on_connect(dbapi_connection: Any, connection_record: Any) -> None:
cursor = dbapi_connection.cursor()
cursor.execute("PRAGMA foreign_keys=ON")
cursor.close()
dbapi_connection.isolation_level = None
@event.listens_for(eng.sync_engine, "begin")
def _begin_immediate(conn: Any) -> None:
conn.exec_driver_sql("BEGIN IMMEDIATE")
_app_engine: AsyncEngine | None = None
_auth_engine: AsyncEngine | None = None
_engine: AsyncEngine | None = None
_sessionmaker: async_sessionmaker[AsyncSession] | None = None
def app_engine() -> AsyncEngine:
global _app_engine
if _app_engine is None:
_app_engine = _make_engine(str(get_settings().database_url_app))
return _app_engine
def _engine_instance() -> AsyncEngine:
global _engine
if _engine is None:
settings = get_settings()
url = settings.database_url
# NullPool in dev/test keeps behaviour predictable across the pytest-asyncio event loop;
# production pool tuning is a deploy-time concern, not a Phase-0 one.
_engine = create_async_engine(
url,
poolclass=NullPool if settings.environment == "test" else None,
echo=False,
)
if url.startswith("sqlite"):
_configure_sqlite_for_concurrent_writers(_engine)
return _engine
def auth_engine() -> AsyncEngine:
global _auth_engine
if _auth_engine is None:
_auth_engine = _make_engine(str(get_settings().database_url_auth))
return _auth_engine
_app_sessionmaker: async_sessionmaker[AsyncSession] | None = None
_auth_sessionmaker: async_sessionmaker[AsyncSession] | None = None
def _app_sessions() -> async_sessionmaker[AsyncSession]:
global _app_sessionmaker
if _app_sessionmaker is None:
_app_sessionmaker = async_sessionmaker(app_engine(), expire_on_commit=False)
return _app_sessionmaker
def _auth_sessions() -> async_sessionmaker[AsyncSession]:
global _auth_sessionmaker
if _auth_sessionmaker is None:
_auth_sessionmaker = async_sessionmaker(auth_engine(), expire_on_commit=False)
return _auth_sessionmaker
def _sessions() -> async_sessionmaker[AsyncSession]:
global _sessionmaker
if _sessionmaker is None:
_sessionmaker = async_sessionmaker(_engine_instance(), expire_on_commit=False)
return _sessionmaker
@asynccontextmanager
async def auth_session() -> AsyncIterator[AsyncSession]:
"""A BYPASSRLS session. See module docstring — auth/service.py only."""
async with _auth_sessions()() as session:
"""See module docstring — auth/service.py only."""
async with _sessions()() as session:
yield session
@asynccontextmanager
async def scoped_session(user_id: UUID) -> AsyncIterator[AsyncSession]:
"""An RLS-scoped session for a known, authenticated user.
`SET LOCAL` binds to the current transaction, not the connection, so this is safe under
connection pooling — it can never leak `app.user_id` from one request into a pooled
connection reused by a different request.
"""
async with _app_sessions()() as session:
async with session.begin():
await session.execute(
# bound parameter, not string interpolation — user_id is a UUID we generated
# or validated ourselves, but there is no reason to ever risk it.
text("SELECT set_config('app.user_id', :uid, true)"),
{"uid": str(user_id)},
)
yield session
@asynccontextmanager
async def unscoped_session() -> AsyncIterator[AsyncSession]:
"""An `app`-role session with no `app.user_id` set.
RLS policies default-deny when `current_setting('app.user_id', true)` is NULL, so this sees
zero rows of any user-owned table — useful for health checks and anything that only touches
non-RLS tables. Prefer `scoped_session` whenever a user is known.
"""A plain session with no user scoping applied at all — for health checks and anything that
never touches a user-owned table. Prefer `scoped_session` whenever a user is known; reaching
for this instead of that for a query against a user-owned table is exactly the mistake
invariant #4 exists to catch in review.
"""
async with _app_sessions()() as session:
async with _sessions()() as session:
yield session
class Scope:
"""A user-scoped query builder. `select()` is the only way to build a query through this
object, and it is always pre-filtered to `user_id` — there is no method here that returns an
unfiltered query against a user-owned table.
"""
def __init__(self, session: AsyncSession, user_id: UUID) -> None:
self.session = session
self.user_id = user_id
def select(self, model: type[DeclarativeBase]) -> "Select[Any]":
if not hasattr(model, "user_id"):
raise TypeError(
f"{model.__name__} has no user_id column — it isn't a user-owned table, so "
"scoped_session() is the wrong tool here. Use auth_session() (pre-identity "
"lookups only) or unscoped_session() (health checks etc.) instead."
)
return select(model).where(model.user_id == self.user_id)
@asynccontextmanager
async def scoped_session(user_id: UUID) -> AsyncIterator[Scope]:
"""A `Scope` for a known, authenticated user. See module docstring."""
async with _sessions()() as session:
async with session.begin():
yield Scope(session, user_id)
+34 -7
View File
@@ -1,15 +1,42 @@
from datetime import datetime
from datetime import UTC, datetime
from typing import Any
from sqlalchemy import DateTime
from sqlalchemy.engine import Dialect
from sqlalchemy.orm import DeclarativeBase
from sqlalchemy.types import TypeDecorator
class UTCDateTime(TypeDecorator[datetime]):
"""`DateTime(timezone=True)` does NOT round-trip tzinfo on SQLite — confirmed empirically:
a tz-aware datetime goes in, a naive one comes back out, and every
`expires_at < datetime.now(UTC)` style comparison in auth/service.py then raises
`TypeError: can't compare offset-naive and offset-aware datetimes`. Not a quirk worth a
per-column workaround — every timestamp in this app is UTC by convention (docs/PLAN.md), so
re-attach UTC on load here, once, rather than trust the driver to preserve it.
"""
impl = DateTime(timezone=True)
cache_ok = True
def process_bind_param(self, value: datetime | None, dialect: Dialect) -> datetime | None:
if value is not None and value.tzinfo is None:
raise ValueError(
"naive datetime passed to a UTCDateTime column — always construct with "
"datetime.now(UTC), never datetime.now() or datetime.utcnow()"
)
return value
def process_result_value(self, value: Any, dialect: Dialect) -> datetime | None:
if value is not None and value.tzinfo is None:
value = value.replace(tzinfo=UTC)
return value # type: ignore[no-any-return]
class Base(DeclarativeBase):
# Every timestamp in this schema is timestamptz UTC (docs/PLAN.md's convention) — without
# this, a bare `Mapped[datetime]` infers a naive TIMESTAMP column, which then silently
# disagrees with a migration that (correctly) declares DateTime(timezone=True), and asyncpg
# rejects the mismatch at insert time. Setting it once here means every current and future
# model gets it right by default instead of each column needing to repeat it.
# Every timestamp in this schema is UTC (docs/PLAN.md's convention) — without this, a bare
# `Mapped[datetime]` infers a plain DateTime, which hits the SQLite round-trip bug above on
# every single column instead of being fixed once, here, for every current and future model.
type_annotation_map = {
datetime: DateTime(timezone=True),
datetime: UTCDateTime(),
}
+30 -22
View File
@@ -1,29 +1,33 @@
"""Identity tables: users, invites, sessions, api_tokens.
See docs/PLAN.md "Auth" and "Schema > Tables > Identity" for the design rationale, and
db.py's module docstring for why two DB roles exist. RLS policies for these tables are created
in the baseline Alembic migration (alembic/versions/0001_baseline.py), not here — SQLAlchemy
models describe columns, not database-level security policy, and keeping the policy SQL visible
and reviewable in the migration is deliberate.
See docs/PLAN.md "Auth" and "Schema > Tables > Identity" for the design rationale, and db.py's
module docstring for how isolation between users is enforced now that there's no RLS to do it at
the database level (docs/DECISIONS.md D15).
Types are deliberately dialect-generic (`sqlalchemy.Uuid`, `JSON`, plain `String` for the IP
column) rather than the `postgresql.*` variants Phase 0 originally used — this schema now targets
SQLite only, but there's no reason to hand-tie it to a Postgres-only type where a portable one
works identically. See models/base.py for why timestamps need a custom type at all.
"""
from datetime import datetime
from datetime import UTC, datetime
from uuid import UUID
from sqlalchemy import ARRAY, ForeignKey, LargeBinary, String, Text
from sqlalchemy.dialects.postgresql import INET
from sqlalchemy.dialects.postgresql import UUID as PGUUID
from sqlalchemy import JSON, ForeignKey, LargeBinary, String, Text, Uuid
from sqlalchemy.orm import Mapped, mapped_column, relationship
from sqlalchemy.sql import func
from velodrome.ids import new_id
from velodrome.models.base import Base
def _now_utc() -> datetime:
return datetime.now(UTC)
class User(Base):
__tablename__ = "users"
id: Mapped[UUID] = mapped_column(PGUUID(as_uuid=True), primary_key=True, default=new_id)
id: Mapped[UUID] = mapped_column(Uuid(as_uuid=True), primary_key=True, default=new_id)
email: Mapped[str] = mapped_column(String(320), unique=True, nullable=False)
display_name: Mapped[str] = mapped_column(String(200), nullable=False)
password_hash: Mapped[str] = mapped_column(Text, nullable=False)
@@ -32,7 +36,7 @@ class User(Base):
# Display-only, per CLAUDE.md invariant #3 — storage is always SI, this never touches a query.
unit_system: Mapped[str] = mapped_column(String(10), nullable=False, default="imperial")
is_active: Mapped[bool] = mapped_column(nullable=False, default=True)
created_at: Mapped[datetime] = mapped_column(server_default=func.now(), nullable=False)
created_at: Mapped[datetime] = mapped_column(nullable=False, default=_now_utc)
sessions: Mapped[list["Session"]] = relationship(back_populates="user")
api_tokens: Mapped[list["ApiToken"]] = relationship(back_populates="user")
@@ -41,12 +45,12 @@ class User(Base):
class Invite(Base):
__tablename__ = "invites"
id: Mapped[UUID] = mapped_column(PGUUID(as_uuid=True), primary_key=True, default=new_id)
id: Mapped[UUID] = mapped_column(Uuid(as_uuid=True), primary_key=True, default=new_id)
# sha256 digest of the invite code. The code itself is never stored anywhere — see
# auth/service.py. 32 bytes for sha256.
code_hash: Mapped[bytes] = mapped_column(LargeBinary(32), unique=True, nullable=False)
created_by: Mapped[UUID] = mapped_column(
PGUUID(as_uuid=True), ForeignKey("users.id"), nullable=False
Uuid(as_uuid=True), ForeignKey("users.id"), nullable=False
)
email: Mapped[str | None] = mapped_column(String(320), nullable=True)
role: Mapped[str] = mapped_column(String(20), nullable=False, default="member")
@@ -59,9 +63,9 @@ class Invite(Base):
class Session(Base):
__tablename__ = "sessions"
id: Mapped[UUID] = mapped_column(PGUUID(as_uuid=True), primary_key=True, default=new_id)
id: Mapped[UUID] = mapped_column(Uuid(as_uuid=True), primary_key=True, default=new_id)
user_id: Mapped[UUID] = mapped_column(
PGUUID(as_uuid=True),
Uuid(as_uuid=True),
ForeignKey("users.id", ondelete="CASCADE"),
nullable=False,
index=True,
@@ -71,9 +75,11 @@ class Session(Base):
token_hash: Mapped[bytes] = mapped_column(LargeBinary(32), unique=True, nullable=False)
client: Mapped[str] = mapped_column(String(20), nullable=False, default="web")
user_agent: Mapped[str | None] = mapped_column(Text, nullable=True)
ip: Mapped[str | None] = mapped_column(INET, nullable=True)
created_at: Mapped[datetime] = mapped_column(server_default=func.now(), nullable=False)
last_seen_at: Mapped[datetime] = mapped_column(server_default=func.now(), nullable=False)
# Plain string, not a native INET type — SQLite has no such type, and app code never queries
# or indexes on structure within the address, only stores/displays it.
ip: Mapped[str | None] = mapped_column(String(45), nullable=True)
created_at: Mapped[datetime] = mapped_column(nullable=False, default=_now_utc)
last_seen_at: Mapped[datetime] = mapped_column(nullable=False, default=_now_utc)
expires_at: Mapped[datetime] = mapped_column(nullable=False)
revoked_at: Mapped[datetime | None] = mapped_column(nullable=True)
@@ -83,16 +89,18 @@ class Session(Base):
class ApiToken(Base):
__tablename__ = "api_tokens"
id: Mapped[UUID] = mapped_column(PGUUID(as_uuid=True), primary_key=True, default=new_id)
id: Mapped[UUID] = mapped_column(Uuid(as_uuid=True), primary_key=True, default=new_id)
user_id: Mapped[UUID] = mapped_column(
PGUUID(as_uuid=True),
Uuid(as_uuid=True),
ForeignKey("users.id", ondelete="CASCADE"),
nullable=False,
index=True,
)
name: Mapped[str] = mapped_column(String(200), nullable=False)
token_hash: Mapped[bytes] = mapped_column(LargeBinary(32), unique=True, nullable=False)
scopes: Mapped[list[str]] = mapped_column(ARRAY(String), nullable=False, default=list)
# JSON, not ARRAY(String) — SQLite has no array type. Stored as a JSON-encoded TEXT column;
# SQLAlchemy handles the (de)serialization transparently.
scopes: Mapped[list[str]] = mapped_column(JSON, nullable=False, default=list)
last_used_at: Mapped[datetime | None] = mapped_column(nullable=True)
expires_at: Mapped[datetime | None] = mapped_column(nullable=True)
revoked_at: Mapped[datetime | None] = mapped_column(nullable=True)
+63 -4
View File
@@ -38,7 +38,7 @@ zero-touch *and* higher fidelity than the current Strava route.
**Fallbacks, both built:** USB watch folder (also the historical-backfill mechanism, so it stays
exercised rather than bit-rotting) and manual upload.
### D4 — Python / FastAPI / Postgres+PostGIS
### D4 — Python / FastAPI / Postgres+PostGIS — **database choice superseded by D15**
**Chosen:** Python 3.12, FastAPI, Pydantic v2, SQLAlchemy 2.0 async, Alembic, PostgreSQL 16 + PostGIS 3.4.
**Rejected:** TypeScript full-stack, Go.
**Why:** `fitdecode` is Python-only, and the Bryton poller reference implementation is Python
@@ -47,7 +47,13 @@ looks in JS). FastAPI emits OpenAPI 3.1 for free. PostGIS is needed for heatmaps
self-segment matching. Go has the best raw performance and a fine FIT library but the weakest
data-analysis ecosystem for the later analytics work.
### D5 — procrastinate for background jobs
The Python/FastAPI half of this still stands. **The database half — Postgres+PostGIS — was
replaced with SQLite in D15**, after Phase 0 was already built and merged against Postgres. Kept
here, marked superseded rather than deleted, so the PostGIS-specific reasoning (heatmaps, bbox
queries, segment matching) is still visible as context for whatever Phase 3 ends up doing about
spatial storage without it.
### D5 — procrastinate for background jobs — **needs a replacement, see D15**
**Chosen:** `procrastinate` (Postgres-backed queue).
**Rejected:** Celery (needs Redis/RabbitMQ, heavyweight, weak async, second-class Postgres broker),
arq (Redis-only — a whole container for ~50 tasks/day), APScheduler (a scheduler, not a durable
@@ -57,6 +63,13 @@ enqueue commit atomically on one connection, so there are no orphaned blobs and
rolled-back rows. That's impossible with a Redis broker without inventing an outbox. It also has
built-in cron, which removes the scheduler container, and keeps job state inside the same `pg_dump`.
`procrastinate` is Postgres-only — no SQLite backend exists, so **D15's move to SQLite invalidates
this choice**. Nothing consumes a job queue yet (no ingestion pipeline exists), so this is a
deferred decision, not an urgent one: whatever replaces it (APScheduler for a *scheduler*, or a
hand-rolled `SELECT ... WHERE claimed_at IS NULL LIMIT 1` polling table for real job durability —
SQLite's single-writer model makes even a crude polling table viable at this app's scale) needs
picking before Phase 1's ingestion pipeline, not before.
### D6 — Opaque bearer tokens, no JWT
**Chosen:** Argon2id passwords + opaque tokens in a `sessions` table, HttpOnly cookie for the PWA.
**Rejected:** JWT.
@@ -82,6 +95,13 @@ downstream number self-corrects. Parts moving between bikes is two rows. A store
neither without a reconciliation nightmare — which is exactly where FitTrackee's flat equipment tag
falls over in year two.
The time-ranged-association *model* doesn't depend on Postgres. The `EXCLUDE USING gist` constraint
enforcing "a component is in exactly one place at a time" does — SQLite has no range types and no
exclusion constraints. `component_installs` doesn't exist yet (Phase 2), so this is another
deferred casualty of D15, not an active one: the same invariant will need enforcing at the
application layer (check-then-insert inside a transaction) instead of the database refusing an
overlapping row outright.
### D9 — Streams as columnar int32 arrays
**Chosen:** one row per channel per activity, `values_i32[]` with a scale factor.
**Rejected:** a normalized per-sample table (~5x larger with index, and every real query wants the
@@ -123,6 +143,44 @@ provide) — now every artefact builds on the same runner.
reverse-engineering, synthesis-heavy); Sonnet for the two breadth surveys (existing self-hosted
apps, Gitea CI patterns) where material is well-documented and the work is volume.
### D15 — SQLite, single container, no database-level RLS
**Chosen:** SQLite as the database, and the whole app (Caddy + API, static web build baked in) as
a single container. Made explicitly *after* Phase 0 was already built, tested, and merged against
Postgres+PostGIS with a two-role RLS architecture (D4, and the `velodrome_app`/`velodrome_auth`
split in `apps/api/velodrome/db.py`) — this is a deliberate reversal of a shipped decision, not a
greenfield choice, and it was made with the costs stated plainly first.
**Rejected, with reasons on the record:** keeping Postgres as a second container (rejected —
explicitly wanted exactly one container total); bundling Postgres+PostGIS *inside* the single
container via a process supervisor (offered as the way to get "one container" without losing RLS
or PostGIS — rejected in favour of SQLite specifically).
**What this costs, stated once here rather than re-litigated every time it's felt:**
- **Row-level security is gone.** SQLite has no roles, no session variables, no policy engine —
there is no database-enforced layer left, only the repository-layer scope. CLAUDE.md's
invariant #4 is revised accordingly (see the file) to describe app-layer scoping as the sole
mechanism rather than one of two layers. The user isolation test in `tests/test_auth.py` that
used to prove RLS itself now proves the repository-layer scope does the same job in its
absence — read it before touching any query that filters by `user_id`.
- **PostGIS is gone.** No native geometry columns, no GIST spatial indexes, no `ST_Envelope`.
Nothing in the schema uses it yet (Phase 0 has no `activities` table), so this is a live
decision for Phase 1/3 to make, not a retrofit — options include SpatiaLite, or storing tracks
as GeoJSON/WKB in a `TEXT`/`BLOB` column with spatial math done in application code.
- **`procrastinate` is gone** (D5) — Postgres-only, no SQLite backend. Also nothing consumes it
yet; a replacement gets picked before Phase 1's ingestion pipeline needs one, not now.
- **The `EXCLUDE USING gist` constraint design for `component_installs`** (D8) — doesn't exist
yet either (Phase 2); the "one place at a time" invariant will need application-layer
enforcement instead of the database refusing an overlapping row outright.
**Why proceed anyway:** raised as a concern in-session, with each cost above stated before this
decision was made; the user heard the full list and confirmed SQLite regardless. That's their call
to make about their own single-user/family-scale instance, not an oversight to correct for them.
**What's unchanged:** invariants #1 (raw bytes immutable), #3 (SI integers in storage), #5 (secret
containment), #6 (single ingestion path) — none of those were ever Postgres-specific. Auth design
(D6, opaque bearer tokens) is unaffected. FastAPI/SQLAlchemy/Alembic stay exactly as chosen in D4;
only the database engine underneath them changed.
---
## Deliberately deferred
@@ -130,7 +188,8 @@ apps, Gitea CI patterns) where material is well-documented and the work is volum
- **Routing** (Valhalla/Photon/Overpass) — Phase 5, optional. Several GB of RAM for something
Komoot already does well.
- **Local LLM ride summaries** (Ollama) — Phase 4, behind a compose profile.
- **Friends/family cross-visibility** — Phase 4, as an *additive* widened RLS policy, never as
removal of the default scope.
- **Friends/family cross-visibility** — Phase 4, as an *additive* widening of the repository-layer
scope (an RLS policy pre-D15; see D15 for why that's no longer the mechanism), never as removal
of the default per-user scope.
- **Legacy Bryton format support** — out of scope entirely. The Rider 650 writes `.fit`.
- **Reverse-engineering Bryton's BLE** — explicitly rejected. See RESEARCH.md §1.
+33 -15
View File
@@ -61,8 +61,8 @@ can pull rides off the head unit. Explicitly out of scope.
| Layer | Choice | Why |
|---|---|---|
| API | Python 3.12, FastAPI, Pydantic v2, SQLAlchemy 2.0 async, Alembic | `fitdecode` is Python-only, so Python is forced; FastAPI emits OpenAPI 3.1 free |
| DB | PostgreSQL 16 + PostGIS 3.4 (`postgis/postgis:16-3.4`) | Needed for heatmaps, bbox queries, self-segment matching |
| Jobs | **procrastinate** (Postgres-backed queue) | Transactional enqueue; no Redis; built-in cron and retries |
| DB | **SQLite** (single file, inside the app container) | Chosen over Postgres+PostGIS specifically to keep the whole deploy to one container — see `docs/DECISIONS.md` D15 for the full reasoning and what it costs (no database-level RLS, no PostGIS, `procrastinate` needs replacing) |
| Jobs | **TBD before Phase 1**`procrastinate` no longer fits (Postgres-only) | Nothing enqueues a job yet; pick this when the ingestion pipeline actually needs it, not before |
| Frontend | SvelteKit `adapter-static` **SPA + PWA** — no Node process in prod | Static files served by Caddy; enforces API-first by construction |
| Maps | MapLibre GL JS from day one | Renders raster tiles now, self-hosted PMTiles vector later — a config change, not a rewrite |
| FIT parsing | `fitdecode` | Thread-safe, preserves header+CRC, correct developer-field handling; `python-fitparse`'s own maintainers point here |
@@ -106,27 +106,45 @@ code-generation exercise against a stable contract rather than a rewrite.
## Service topology
**v1 — 4 containers, under 2GB RAM total:**
- `caddy` — serves the static SPA, proxies `/api/*`. Same-origin, so no CORS. Your existing reverse proxy terminates TLS in front.
- `api` — uvicorn/FastAPI. Runs Alembic on entrypoint.
- `worker` — same image, `procrastinate worker`. Owns periodic tasks too, so no separate scheduler.
- `db` — postgis/postgis:16-3.4.
**Revised by D15 — single container**, not the multi-container compose topology this section
originally described. Caddy and the FastAPI app run together in one image via a lightweight
process supervisor; SQLite is a file inside the same container's persistent volume, not a
separate service. See `docs/DECISIONS.md` D15 for why, and `deploy/`'s own README for the actual
supervisor config once it exists.
Volumes: `pgdata`, `blobstore` (content-addressed raw FIT + attachments), `import_inbox` (USB watch folder bind-mount).
Volumes: one persistent volume holding the SQLite file, `blobstore` (content-addressed raw FIT +
attachments), and `import_inbox` (USB watch folder bind-mount).
**Phase 2 adds zero containers** (poller is a periodic task; Open-Meteo is a public API).
**Phase 3 adds two:** `tileserver` (tileserver-gl-light + regional PMTiles, ~200MB1GB) and `topodata` (Open Topo Data + region-clipped SRTM).
**Phase 4+ behind opt-in compose profiles:** `ollama`, `grafana`, and optionally `valhalla`/`photon`/`overpass`.
**Later phases that would have been "add a container" under the old topology**`tileserver`,
`topodata`, `ollama`, `grafana`, `valhalla`/`photon`/`overpass` — still make sense as genuinely
separate containers even under a single-container-for-the-app model (they're independent services
with their own resource profiles, not part of "the app"). Whether the app container talks to them
over a shared Docker network or they stay fully optional add-ons is a decision for whichever phase
actually needs the first one — not resolved here speculatively.
**Explicitly NOT on day one:** Valhalla, Photon, Overpass, Nominatim (>1TB/128GB RAM — never), Ollama, Redis, MinIO, Grafana, a separate scheduler.
**Explicitly NOT on day one:** Valhalla, Photon, Overpass, Nominatim (>1TB/128GB RAM — never), Ollama, Redis, MinIO, Grafana.
---
## Schema
Conventions: UUIDv7 PKs (time-ordered, URL-safe, client-generatable). All `timestamptz` UTC. **All
physical quantities as SI integers** — distance in metres, time in seconds, speed in mm/s, altitude
in cm, money in minor units. Units are a presentation concern.
**Note on what's below vs. what's actually built:** the Identity tables (§ Tables > Identity) are
real, built, and running on SQLite — `apps/api/velodrome/models/identity.py` and
`apps/api/alembic/versions/0001_baseline.py` are the source of truth for those, not this
prose. Everything past Identity (activities, streams, bikes/components, service rules,
notifications, weather) was designed against Postgres/PostGIS conventions — `geometry(...)`
columns, `ARRAY`, `INET`, GIST indexes, RLS policies — before D15 moved the database to SQLite.
None of it is built yet, so none of it is broken; it just needs a real pass for SQLite
compatibility (TEXT/BLOB for geometry, JSON-encoded TEXT for arrays, plain TEXT for IP addresses,
application-layer exclusion checks instead of `EXCLUDE USING gist`) when each phase actually
builds it, informed by whatever's learned finishing the SQLite migration on Identity first —
not a mechanical find-and-replace on speculative schema now.
Conventions: UUIDv7 PKs (time-ordered, URL-safe, client-generatable). All timestamps UTC (SQLite
has no native timezone-aware timestamp type — see `apps/api/velodrome/models/base.py` for how
Identity stores them; the same convention applies going forward). **All physical quantities as SI
integers** — distance in metres, time in seconds, speed in mm/s, altitude in cm, money in minor
units. Units are a presentation concern.
### Two architectural rules that everything else depends on