SCRUM-225 — Patch: LISTEN/NOTIFY reload with poll fallback¶
Plan ref: PAT-B3/B4 (docs/11-admin-plane-plan.md). Stacked on SCRUM-224.
What changed¶
internal/manifest/watch.go — Watcher.Run(ctx) runs two loops:
| Loop | Behaviour |
|---|---|
| listener | acquires a pool connection and hijacks it (the pool never reuses a LISTENing conn), LISTEN config_release; payload dev/staging/live → Reload(channel), anything else → LoadAll. On error: close, back off 1s→30s (reset on success), reconnect, and LoadAll after every (re)connect — notifications sent while it was down are lost |
| poll | every PATCH_POLL_INTERVAL (60s): read channel, release_id from channel_head; reload any channel whose release differs from what is served |
- Readiness:
/readyz= manifests loaded and the listener has not been down for longer than the poll interval (release listener disconnected for …). - Metric:
patch_manifest_reload_total{source="notify|poll|reconnect"}. - Wired in
patch.goon the main context after the initial load.
Why both: Config's pg_notify inside its publish transaction is delivered only
after commit (right ordering), but it is fire-and-forget — the poll bounds how long
a missed notification can leave a stale manifest (≤ 60s), and the reconnect reload
covers the gap immediately.
How to verify¶
cd services/patch
export PATCH_TEST_DATABASE_URL='postgres://USER:PASS@127.0.0.1:5433/config_test?sslmode=disable'
go vet ./... && go test -race -count=2 ./... # -count=2: stable and self-cleaning
go test -race -count=1 -v -run 'Watcher' ./internal/manifest/
| Test | Proves |
|---|---|
TestWatcherReloadsOnNotify |
a committed head change + pg_notify('config_release','dev') → new dev entry, source=notify counted |
TestWatcherFallsBackToPoll |
listener disabled, 200ms poll, change without notify → reloaded via poll |
TestWatcherReconnectsAndReloadsAll |
pg_terminate_backend on the listener's backend → reconnects; a change made while it was down is picked up by the reconnect LoadAll |
TestWatcherReadyDrivesStateDirectly |
readiness flips only after being down longer than the poll interval |
The tests commit real changes to the dev head in the test DB with valid,
correctly hashed releases, and restore the original head in t.Cleanup (checked
after the run: dev/staging/live back on the bootstrap releases).
Results at time of writing¶
go vet,go test -race -count=2(5 packages) against Postgres 16: pass.
How it was built¶
DeepSeek run scoped (Landlock) to services/patch (142 s, ~25k output tokens).
Claude review: listener/backoff/reconnect logic and nil-safety read, suite run
twice, head restoration checked. No changes needed.