fix(deploy): push through a TLS-terminating proxy, not raw Gitea HTTP
CI / Repo hygiene (pull_request) Successful in 2s
CI / Web (lint, typecheck, build) (pull_request) Successful in 16s
CI / Migrations reversible (pull_request) Successful in 6s
CI / API (lint, types, tests) (pull_request) Successful in 54s

release.yml's first real run failed: docker/login-action against
192.168.0.3:3000 hit "server gave HTTP response to HTTPS client" — Docker
refuses any non-localhost registry over plain HTTP by default, so this was
never actually a workflow bug.

Rejected insecure-registries in daemon.json after reading this Unraid host's
own rc.docker script: applying it needs a full dockerd restart, and with
Live Restore disabled here, that stops every one of the ~40 other containers
on the box first. Also rejected a real Let's Encrypt cert on a public
bbergle.com subdomain — this host's other subdomains are Cloudflare-proxied,
which would terminate TLS at Cloudflare's edge and never reach our own cert
at all.

Chosen instead, scoped to touch nothing already working: a self-signed cert
for registry.bbergle.com behind a new NPMplus proxy host (found its real
HTTPS port, 9537, by reading `docker port NPMplus` rather than assuming 443,
which is a different nginx process on this box entirely); an /etc/hosts
entry on the Unraid host so only that host needs to resolve the name (no
DNS record, no router/NAT dependency); and its CA dropped into
/etc/docker/certs.d, which Docker's own docs confirm is read per-connection
with no daemon restart required. Also pins buildx to driver: docker instead
of setup-buildx-action's default docker-container driver, which runs an
isolated builder that doesn't see /etc/docker/certs.d and would have quietly
defeated all of the above.

Full record, including what was rejected and why, in docs/DECISIONS.md D17.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R2ZKeWkZV7ehf7fivrAkkG
This commit is contained in:
2026-09-21 20:52:56 -04:00
co-authored by Claude Sonnet 5
parent 32037b1190
commit a114a7d3d8
4 changed files with 85 additions and 10 deletions
+56
View File
@@ -233,6 +233,62 @@ image out is a manual/Unraid-side action (pull + Apply, or Unraid's own update c
something CI does unattended — consistent with treating "affects a shared, already-running system"
as something a human triggers, not automation.
### D17 — Registry TLS: self-signed cert behind NPMplus, not `insecure-registries`, not a real domain
**Problem:** `release.yml`'s first real run failed — `docker/login-action` against
`192.168.0.3:3000` (Gitea's plain-HTTP address) hit `server gave HTTP response to HTTPS client`.
Docker refuses TLS-less registries by default; this was never a workflow misconfiguration, it's
expected Docker behaviour for any non-localhost registry.
**Rejected: `insecure-registries` in `daemon.json`.** The obvious fix. Rejected after actually
reading `/etc/rc.d/rc.docker` on the Unraid host rather than assuming: applying a `daemon.json`
change requires a full `dockerd` restart, and (with `Live Restore` disabled on this host) both
Unraid's own restart path *and* a raw `kill` of `dockerd` stop every one of the ~40 other
containers running on the box first, as part of the restart/shutdown sequence — Plex, Home
Assistant, Vaultwarden, everything. Correct fix for the narrow problem, unacceptable blast radius
for this specific host.
**Rejected: a real Let's Encrypt cert on a new `bbergle.com` subdomain routed publicly.** The
user's other NPMplus-fronted subdomains resolve through Cloudflare's proxy (orange-cloud), not
directly to the home IP. A Cloudflare-proxied hostname would have terminated TLS at Cloudflare's
edge with Cloudflare's own cert, never reaching our self-signed cert or NPMplus's own TLS
config at all — the entire trust chain would depend on Cloudflare's origin SSL mode, and likely on
firewall rules restricting port 443 to Cloudflare's IP ranges, neither of which this problem
needed to involve.
**Chosen:** a small, fully self-contained fix, scoped to touch nothing already working:
- A 10-year self-signed cert for `registry.bbergle.com` (SAN-only, no real domain dependency).
- An NPMplus proxy host (`registry.bbergle.com` -> `192.168.0.3:3000` over plain HTTP internally)
terminating TLS with that cert, on NPMplus's existing HTTPS port (`9537` on this host — found by
reading `docker port NPMplus` rather than assuming 443, which is a *different* nginx process on
this box entirely).
- `/etc/hosts` on the Unraid host mapping `registry.bbergle.com` -> `192.168.0.103` (itself) —
chosen over a real DNS record specifically because the only client that ever needs to resolve
this hostname is the Unraid host's own `dockerd` (Gitea Actions runs in DooD mode against that
same host's Docker socket). This sidesteps Cloudflare, the router's NAT/hairpin behaviour, and
any port-forwarding question entirely — verified separately that hairpin NAT works by default on
this user's UniFi gateway, but it turned out to be unnecessary for this fix regardless.
- `/etc/docker/certs.d/registry.bbergle.com:9537/ca.crt` on the Unraid host, trusting that cert for
that host:port specifically. Confirmed (Docker's own docs) that `certs.d` is read per-connection,
not baked in at daemon start — no `dockerd` restart, no impact on any other container.
- `docker/setup-buildx-action@v3` pinned to `driver: docker` in `release.yml` instead of its
default `docker-container` driver — the default runs BuildKit in an isolated builder container
that does not see the host's `/etc/docker/certs.d`, which would have silently defeated the whole
point of the trust setup above. We don't build multi-platform images, so nothing the
`docker-container` driver offers is actually needed here.
**Not persisted across a reboot, deliberately, for now:** neither the `/etc/hosts` line nor the
`certs.d` file are wired into `/boot/config/go` — both live under `/`, which Unraid rebuilds fresh
from `/boot` on every boot. Raised explicitly rather than assumed: the user was (rightly) wary of
hand-editing anything under `/boot` after an earlier, unrelated discussion of what a broken `go`
script could do to boot. Persisting this is a five-minute follow-up (append two lines to `go`) once
they're ready to make that call deliberately, not bundled into this fix.
**What's unaffected:** Gitea's own web UI, git remote, and API — all still plain
`http://192.168.0.3:3000`, exactly as CLAUDE.md documents. NPMplus's existing public proxy hosts
and certs (`vaultwarden.bbergle.com` etc.) — untouched, new proxy host only. No other container on
the Unraid host was restarted, reconfigured, or otherwise touched to make this work.
---
## Deliberately deferred