refactor(api): move from Postgres+RLS to single-engine SQLite
Reverses a shipped, tested, merged decision (D4/PR #2) rather than building on it — see docs/DECISIONS.md D15 for the full record: what was rejected (Postgres as a second container; Postgres+PostGIS bundled inside the single container via a supervisor), what this costs (no database-level RLS, no PostGIS, procrastinate needs replacing — all stated as a concern before this was decided, and reaffirmed anyway, which is the user's call to make about their own instance). The one invariant-critical consequence: isolation between users now rests entirely on the repository-layer scope (db.py's `Scope.select()`), not two layers. CLAUDE.md's invariant #4 is revised accordingly. This is not a downgrade-and-hope — `Scope` is built so an unfiltered query against a user-owned table is structurally harder to write than a scoped one (there is no method on `Scope` that returns one), and tests/test_auth.py::test_scoped_session_blocks_cross_user_reads replaces the old RLS proof with the same empirical standard: it doesn't trust the query builder filters correctly because the code reads correctly, it registers two real users and checks. test_unscoped_session_can_see_every_user_when_misused is the deliberately alarming companion — it demonstrates exactly what a reviewer must now catch, since nothing else will. Six real, non-obvious SQLite behaviours found and fixed by actually running this against a real file, not assumed from docs: - Foreign keys, ON DELETE CASCADE included, are OFF by default per connection — deleting a user silently left orphaned sessions/api_tokens, no error either way. Fixed with PRAGMA foreign_keys=ON on every connect. - Transactions default to DEFERRED, which only takes a write lock on the first actual write — a real check-then-act race for invite redemption (two concurrent redemptions could both read used_count < max_uses as true before either commits). Fixed by disabling the driver's implicit BEGIN and issuing BEGIN IMMEDIATE ourselves — SQLAlchemy's own documented recipe for this, not improvised. - DateTime(timezone=True) does NOT round-trip tzinfo on SQLite — a tz-aware datetime goes in, a naive one comes back out, and every `expires_at < datetime.now(UTC)` comparison in auth/service.py then raises TypeError. Fixed once at the Base level with a UTCDateTime TypeDecorator rather than per-column. - Uuid(as_uuid=True) stores as 32-char hex with NO hyphens on SQLite, not str(uuid)'s hyphenated form. A test fixture that raw-inserted the hyphenated form left rows the ORM's own later UPDATE (via invite.used_count += 1's autoflush) could never match by primary key, updating zero rows and raising StaleDataError. Fixed by using .hex to match exactly what the ORM itself writes. - BEGIN IMMEDIATE applies to every transaction, reads included — a long-lived test fixture that autobegins a transaction via a bare read and never explicitly closes it holds SQLite's exclusive write lock for the rest of the test, and a later scoped_session() call fails with "database is locked". Not an app-code bug (every real session block closes cleanly on exit), but real enough to document since the next person writing a test against the db_auth fixture will hit it too. - Python's sqlite3 module deprecates its own implicit datetime adapter as of 3.12 — silent today, warns on every raw-SQL datetime bind. Only ever hit test fixture code (the ORM path never uses it, confirmed by running the ORM-only health test with warnings promoted to errors and it stayed clean); fixed there with an explicit .isoformat() rather than left for a future Python version to turn into a real failure. Also, since with_for_update() silently no-ops on SQLite (confirmed — SQLAlchemy emits no SQL for it, no error either) rather than actually locking anything: removed it from register()'s invite-redemption query and corrected the comment to attribute the concurrency guarantee to BEGIN IMMEDIATE, where it now actually lives. One PR, not several, for the same reason PR #2 was: the migration, the models, db.py, and the docs recording why are five views of one decision — splitting them wouldn't make review easier, just disconnected. 552 insertions / 548 deletions across 17 files, most of it necessarily touching what PR #2 shipped rather than net-new code. Deliberately deferred, not solved here: PostGIS's replacement for spatial storage, procrastinate's replacement for background jobs, and the EXCLUDE USING gist constraint's replacement for component_installs — none of those tables exist yet (Phase 1-2), so none of it is broken, and docs/DECISIONS.md D15 records exactly what each future phase needs to decide before it can be built. .gitea/workflows/deploy pipeline (PR #4, built for the old 3-container Postgres compose stack) was closed as superseded rather than merged; the single-container image build is follow-up work, not part of this change. Verified: ruff check, ruff format --check, and mypy --strict all clean. 13/13 pytest passing against a real SQLite file, including with DeprecationWarning promoted to an error (confirms the sqlite3 adapter deprecation fix actually holds, not just that it's quiet by default). Full alembic upgrade -> downgrade -1 -> upgrade cycle run clean. alembic check clean with no include_object filter needed at all now (SQLite starts with nothing but what our own migrations create — no PostGIS/TIGER noise to filter out in the first place). CI's exact migration command sequence reproduced locally end to end before touching the workflow file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -17,7 +17,8 @@ say so and argue it — but don't silently contradict it.
|
||||
apps/api/ Python 3.12 / FastAPI / SQLAlchemy async / Alembic
|
||||
apps/web/ SvelteKit static SPA (installable PWA)
|
||||
packages/openapi/ openapi.json — COMMITTED contract artefact, CI enforces it matches the code
|
||||
deploy/ docker-compose, Caddyfile, systemd units, backup scripts
|
||||
deploy/ single-container Dockerfile, Caddyfile, systemd units, backup scripts —
|
||||
see docs/DECISIONS.md D15 for why this isn't docker-compose
|
||||
docs/ plan, decisions, research
|
||||
scripts/ repo tooling (PR helpers, etc.)
|
||||
.gitea/workflows/ CI
|
||||
@@ -34,8 +35,12 @@ These are load-bearing. Breaking one is a correctness bug, not a style choice.
|
||||
time-ranged `component_installs`. Never add a stored running total to a component.
|
||||
3. **All physical quantities are SI integers** in storage — metres, seconds, mm/s, centimetres,
|
||||
grams, minor currency units. Imperial is display-only. Never store a float mile.
|
||||
4. **Every user-owned table has `user_id`, an RLS policy, and a repository-layer scope.** Both
|
||||
layers, always. Never rely on the query alone.
|
||||
4. **Every user-owned table has `user_id`, and every query against it goes through the
|
||||
repository-layer scope helper — never a raw query filtered by hand.** This used to be backed
|
||||
by Postgres RLS as a second, database-enforced layer (see `docs/DECISIONS.md` D4/D15); SQLite
|
||||
has no equivalent, so the repository-layer scope is now the *only* enforcement, which makes it
|
||||
non-negotiable rather than defense-in-depth. A new domain table without a passing isolation
|
||||
test (see `tests/test_auth.py`'s pattern) is not done.
|
||||
5. **Secrets never leave the server.** The Bryton credential is password-equivalent. It must not
|
||||
appear in any API response model, any log line, or any error message.
|
||||
6. **One ingestion path.** All sources funnel through `ingest_bytes()`. Never add a second parse
|
||||
@@ -47,8 +52,8 @@ These are load-bearing. Breaking one is a correctness bug, not a style choice.
|
||||
calls in request handlers.
|
||||
- **SQL:** migrations via Alembic only, never manual DDL. Every migration must survive
|
||||
`upgrade -> downgrade -1 -> upgrade`.
|
||||
- **Tests:** pytest against a real Postgres service container, never mocks for DB behaviour.
|
||||
Parser changes need a golden fixture in `apps/api/tests/fixtures/fit/`.
|
||||
- **Tests:** pytest against a real SQLite file, never mocks for DB behaviour. Parser changes need
|
||||
a golden fixture in `apps/api/tests/fixtures/fit/`.
|
||||
- **Commits:** imperative mood, explain *why* in the body. Conventional-commit prefixes
|
||||
(`feat:`, `fix:`, `refactor:`, `test:`, `docs:`, `chore:`, `ci:`).
|
||||
- Match surrounding code. Don't introduce a new pattern when one exists.
|
||||
@@ -100,7 +105,10 @@ silently corrupts data for six months is expensive work.
|
||||
are silent and corrupt the archive.
|
||||
- `wear/` — the wear SQL. Wrong numbers that still look plausible are the worst failure mode in the
|
||||
product, because nobody notices.
|
||||
- `auth/`, RLS policies — security, and a mistake exposes another user's data.
|
||||
- `auth/`, any repository-layer user-scoping code — security, and a mistake exposes another user's
|
||||
data. This carries more weight than it used to: there is no database-enforced RLS backstop
|
||||
anymore (see invariant #4 and `docs/DECISIONS.md` D15), so this code *is* the isolation
|
||||
boundary, not one layer of it.
|
||||
- `sources/bryton/` — a reverse-engineered protocol with no spec to check against.
|
||||
- Schema migrations that alter or drop existing columns.
|
||||
- Debugging anything that two Sonnet attempts have already failed to fix.
|
||||
@@ -113,7 +121,7 @@ careful transcription plus ordinary judgement, and CI catches the rest:
|
||||
- CRUD endpoints, Pydantic models, repository methods
|
||||
- SvelteKit components, routes, styling, the service worker
|
||||
- Tests against an already-decided behaviour
|
||||
- Additive migrations, docker-compose and Caddy config, CI workflows
|
||||
- Additive migrations, the deploy Dockerfile/Caddy config, CI workflows
|
||||
- Documentation
|
||||
|
||||
### Haiku — mechanical work
|
||||
|
||||
Reference in New Issue
Block a user