Skip to content

SCRUM-235 — Dashboard service health prober and GET /services

Plan ref: DSH-H1 (docs/11-admin-plane-plan.md). Builds on SCRUM-232.

What exists

Piece Notes
DASHBOARD_TARGETS comma-separated name=url (internal base URLs); names ^[a-z][a-z0-9_]{0,31}$, unique; absolute http(s) URLs; empty = no targets (valid)
DASHBOARD_PROBE_INTERVAL 15s default, must be ≥ 1s
internal/health.Prober GET <url>/readyz for every target concurrently, first round at start, then every interval; timeout = DASHBOARD_UPSTREAM_TIMEOUT; never follows redirects; reads ≤ 4 KiB
Status rules transport error → up=false, reason connection refused / connection reset / timeout after 2s / innermost error; any response → up=true; 200 → ready; otherwise reason = COM-5 error.message, else first 200 bytes of the body (cut on a rune boundary), else HTTP <code>; before the first probe not probed yet with checked_at: null; last_ready_at carried across failures
GET /api/admin/dashboard/services (viewer) {"services":[{"name","up","ready","reason","latency_ms","checked_at","last_ready_at"}]} sorted by name; [] with no targets. Other routes still 501 behind api.Handlers.For
Metrics dashboard_service_up{service}, dashboard_service_ready{service} (0/1), dashboard_probe_duration_seconds{service}

The prober is not part of the Dashboard's own readiness: an unreachable target is what it exists to report.

How to verify

cd services/dashboard
gofmt -l . && go vet ./... && go test -race -count=1 ./...
go build -o /tmp/dashboard . && \
DASHBOARD_LISTEN_ADDR=127.0.0.1:18080 DASHBOARD_METRICS_ADDR=127.0.0.1:19090 \
DASHBOARD_STAFF_JWKS_URL=http://127.0.0.1:1/jwks \
DASHBOARD_TARGETS=config=http://127.0.0.1:1,self=http://127.0.0.1:19090 \
DASHBOARD_PROBE_INTERVAL=1s /tmp/dashboard serve &
curl -s 127.0.0.1:19090/metrics | grep dashboard_service_
# config: up 0 ready 0 (closed port); self: up 1 ready 0 (its own /readyz is 503 without JWKS keys)

Tests cover: config parsing (good list, duplicate, bad name, relative URL, empty entry, empty value, interval floor); prober — 200 ready, 503 COM-5 reason, 503 plain text, closed port, timeout, redirect not followed, >4 KiB body truncated, not-probed-yet, recovery sets last_ready_at, 5 × 300 ms targets finish a round in < 1 s; JSON checked_at: null before the first probe; rune-safe reason truncation; handler through the server with a viewer token (shape, sort, []); metrics reflect up/ready.

Results at time of writing

  • gofmt, go vet, go test -race (5 packages): pass. Binary smoke-run as above.

How it was built

DeepSeek run scoped (Landlock) to services/dashboard (307 s, ~53k output tokens, reasoning effort low). Claude review found two small issues, fixed with tests: checked_at rendered as Go's zero time before the first probe (now null), and the 200-byte reason cut could split a UTF-8 rune.