Skip to content

SCRUM-234 — Dashboard overview endpoint with cache

Plan ref: DSH-C5/C7 (docs/11-admin-plane-plan.md). Stacked on SCRUM-233.

What exists

GET /api/admin/dashboard/overview (viewer):

{"generated_at":"…","services":[{"name","up","ready","reason","rps","error_ratio","p95_ms","version"}],
 "host":{"cpu_ratio","mem_ratio","disk_ratio"},"online_players":…,"degraded":[…]}
Piece Notes
Queries (internal/promql/overview.go) one fixed query per card across all services, grouped by (service): up, rps, error ratio, p95 ms, build version; node_exporter CPU/memory/disk; sum(otomo_online_players)
Execution all queries concurrently under one 2 s deadline; a failed, timed-out or empty card is null and named in degraded; the response is always 200
Services union of Prometheus targets and the health prober's targets, sorted; ready/reason from the prober
Cache (internal/cache) key = endpoint + 10 s time bucket, pre-serialised bytes, TTL 10 s, max 64 entries (oldest evicted), singleflight per key; the shared batch is detached from any one caller's cancellation. Cache-Control: no-store

No service exports otomo_online_players yet, so that card is null and degraded until one does (expected).

How to verify

cd services/dashboard
gofmt -l . && go vet ./... && go test -race -count=1 ./...

Tests (fake MetricsSource): every field mapped; one failing query → that field null and degraded, the rest present; a query past 2 s → degraded and the handler returns in < 2.5 s; 10 concurrent requests → one batch; same bucket → cache hit; rolled bucket → new batch; cache bounded at 64; service union; through the server with a viewer token; no token 401.

Results at time of writing

  • gofmt, go vet, go test -race (8 packages): pass.

How it was built

DeepSeek run scoped (Landlock) to services/dashboard (318 s, ~55k output tokens, reasoning effort low). Claude review: cache, singleflight and deadline checked; no changes needed.