SCRUM-234 — Dashboard overview endpoint with cache¶
Plan ref: DSH-C5/C7 (docs/11-admin-plane-plan.md). Stacked on SCRUM-233.
What exists¶
GET /api/admin/dashboard/overview (viewer):
{"generated_at":"…","services":[{"name","up","ready","reason","rps","error_ratio","p95_ms","version"}],
"host":{"cpu_ratio","mem_ratio","disk_ratio"},"online_players":…,"degraded":[…]}
| Piece | Notes |
|---|---|
Queries (internal/promql/overview.go) |
one fixed query per card across all services, grouped by (service): up, rps, error ratio, p95 ms, build version; node_exporter CPU/memory/disk; sum(otomo_online_players) |
| Execution | all queries concurrently under one 2 s deadline; a failed, timed-out or empty card is null and named in degraded; the response is always 200 |
| Services | union of Prometheus targets and the health prober's targets, sorted; ready/reason from the prober |
Cache (internal/cache) |
key = endpoint + 10 s time bucket, pre-serialised bytes, TTL 10 s, max 64 entries (oldest evicted), singleflight per key; the shared batch is detached from any one caller's cancellation. Cache-Control: no-store |
No service exports otomo_online_players yet, so that card is null and degraded
until one does (expected).
How to verify¶
Tests (fake MetricsSource): every field mapped; one failing query → that field null and degraded, the rest present; a query past 2 s → degraded and the handler returns in < 2.5 s; 10 concurrent requests → one batch; same bucket → cache hit; rolled bucket → new batch; cache bounded at 64; service union; through the server with a viewer token; no token 401.
Results at time of writing¶
gofmt,go vet,go test -race(8 packages): pass.
How it was built¶
DeepSeek run scoped (Landlock) to services/dashboard (318 s, ~55k output tokens,
reasoning effort low). Claude review: cache, singleflight and deadline checked; no
changes needed.