Skip to content

SCRUM-236 — Dashboard per-service time series endpoint

Plan ref: DSH-C6 (docs/11-admin-plane-plan.md). Stacked on SCRUM-234.

What exists

GET /api/admin/dashboard/services/{name}/series?metric=&from=&to=&step= (viewer):

Rule Behaviour
metric a catalogue template; unknown or missing → 400 unknown_metric
{name} allow-list + label grammar via Template.Build → else 400 unknown_service, before Prometheus is called
Range RFC3339 or unix seconds; default last hour; future to → now; from >= to → 400; more than 7 d is clamped ("clamped": true)
Step max(requested, 15 s, ceil(span/1500)), whole seconds, from aligned down to the step, and the window re-fitted so there are never more than 1500 points
Rate window max(4 × step, 60 s)
Response columnar {"metric","title","unit","service","from","to","step","clamped","t":[…],"series":[{"name","values":[…]}]}; NaN → null; series named by their single grouping label (e.g. 2xx), sorted
Errors Prometheus error/timeout → 502 upstream_error (no URL); no allow-list → 503
Cache 10 s, keyed by service, metric, aligned from/to and step; singleflight; Cache-Control: no-store

How to verify

cd services/dashboard
gofmt -l . && go vet ./... && go test -race -count=1 ./...

Tests: step table (1 h → 15 s, 24 h → 58 s, 7 d → 404 s, requested 5 m kept) and the 1500-point bound incl. after alignment; a 30 d request is clamped to 7 d; unknown metric/service and an injection attempt in {name} are 400 without a Prometheus call; from >= to 400; columnar shape, NaN → null, series naming/sorting; upstream error 502; cache hit; through the server with a viewer token.

Results at time of writing

  • gofmt, go vet, go test -race (8 packages): pass.

How it was built

DeepSeek run scoped (Landlock) to services/dashboard (266 s, ~46k output tokens, reasoning effort low). Claude review: step/alignment arithmetic and point bound checked; no changes needed.