Otomo: Allocator and player edge plan (M1)¶
For: the owner of the Allocator, service orchestration and the public player edge
(reverse proxy, TLS, Let's Encrypt): Tai Xiao Xuan.
Tickets: SCRUM-290…295 (Allocator), SCRUM-298 (public edge and TLS), plus SCRUM-271 (the
launch hand-off design, which you are best placed to write). Plan IDs are from
13-player-plane-plan.md.
Works with: Isaac Foo (Session calls the Allocator and receives its callbacks) and
Lew Zheng Song (game servers register with the Allocator; the Gameplay Proxy verifies its
join tickets).
1. Facts about the team45 host (checked 2026-09-28)¶
| Fact | Value | Why it matters |
|---|---|---|
| Public address | 51.79.241.70, DNS name team45.dp-ext8.com (the school's domain) |
The certificate's name, and who controls its DNS |
Firewall (ufw) |
Allows only SSH, TCP 5000–5050, UDP 5000–5050, UDP 5055/5056/5058, UDP 27000–27002 | Ports 80 and 443 are closed. Let's Encrypt's usual challenges need one of them (§3) |
| Already in use | 5008 Jenkins (loopback), 5009 Jenkins webhook proxy (public), 5000 local registry (loopback), 8080 player gateway (loopback), 8090 staff gateway (loopback) | Pick new public ports that don't collide |
| Stack | ~/otomo is a git checkout of staging; compose project otomo; update with a git bundle + fast-forward, then deploy/scripts/up.sh |
How your services get deployed (see handoff/HANDOVER.md) |
| Player gateway today | 127.0.0.1:8080, reached through an SSH tunnel |
GW-1 makes it public |
| Staff gateway | 127.0.0.1:8090, SSH tunnel only |
Stays private; not part of GW-1 |
The UDP ranges that are already open (5000–5050, 27000–27002) are what the Gameplay Proxy can use. Tell Lew which one you reserve for it.
2. Order of work¶
- SCRUM-271, launch hand-off design (
docs/14-launch-handoff.md). It fixes the interfaces Isaac and Lew build against, so it comes first. §5 below is a starting draft. - SCRUM-290, Allocator skeleton, and SCRUM-293, join tickets, in parallel.
- SCRUM-291, registry, then SCRUM-292, allocation API, then SCRUM-294, release and timeouts, then SCRUM-295, metrics.
- SCRUM-298, public edge, as soon as the certificate route in §3 is unblocked. It only depends on decision D1, so it can run alongside the Allocator.
3. SCRUM-298: public player edge, TLS and Let's Encrypt¶
3.1 The blocker: getting a certificate¶
Let's Encrypt only issues a certificate after you prove you control the name, with one of:
| Challenge | Needs | Works here today? |
|---|---|---|
| HTTP-01 | Inbound TCP 80 from the internet | No: port 80 is closed |
| TLS-ALPN-01 | Inbound TCP 443 | No: port 443 is closed |
| DNS-01 | Creating a TXT record in the domain's DNS (through a DNS provider API) | Not for dp-ext8.com: the school controls its DNS |
Two ways forward. Pick one before building:
- Option A (recommended): ask the school to open TCP 80 and 443 on
51.79.241.70, inufwand in any upstream firewall. Players then usehttps://team45.dp-ext8.com, the standard port, and certificates renew automatically. Check first whether an upstream firewall exists: after opening 80 inufw, test from outside withcurl -v http://team45.dp-ext8.com/. - Option B: a domain you control, e.g. a cheap
.dev/.ggdomain whose DNS provider has an API (Cloudflare, Porkbun, etc.). Point anArecord at51.79.241.70, get the certificate with DNS-01 (no inbound port needed), and serve HTTPS on an allowed port, e.g. 5043 (https://play.<your-domain>:5043). A certificate is for a name, not a port, so this is valid TLS, just not on 443.
Either way, record the final base URL in doc 06 §12 and doc 12 §3.4, and set
AUTH_PUBLIC_SESSION_URL in deploy/compose.yaml to <base URL>/api/player/session. The
login response currently tells clients http://localhost:8080/api/player/session.
3.2 The reverse proxy (as built)¶
Decision (2026-09-28): option A. The school has no upstream firewall on 80/443, so
ufw rules for 80/443 were added. The proxy is nginx (services/edge) with the host's
certbot, not Caddy. The reason is the shared Let's Encrypt quota: 50 new certificates per
week for all of dp-ext8.com, shared by every team. Caddy requests certificates by itself
whenever its storage is missing, so a lost volume or a misconfigured restart would spend
the shared quota. With certbot, the certificate is requested once, by hand, after a
dry run. certbot's timer renews it, and a deploy hook reloads nginx.
services/edge: nginx on 80/443./auth/,/patch/and/api/player/go to the gateway,/docs/to the technical wiki, and/to an optional site. Before a certificate exists it runs in bootstrap mode (ACME only).- It sits on its own network
otomo-edge(fixed subnet, edge at172.31.250.2) with the gateway and the wiki sites. It can't reach Postgres or any internal service. - No cookies cross it in either direction. HSTS covers this host only.
- Runbook and the quota guardrails:
deploy/edge/README.mdanddeploy/edge/certbot.sh. - The gateway stays published on
127.0.0.1:8080for SSH-tunnel development.
3.3 Required gateway change: client IPs behind the proxy¶
The player gateway rate-limits per client IP (services/gateway/internal/ratelimit:
general 20 req/s, burst 40; /auth/ 5 req/s, burst 10), keyed on the TCP peer address. It
deliberately ignores X-Forwarded-For (see the TODO in ratelimit.go: "TLS termination
question… is unresolved… Do not trust X-Forwarded-For"). Behind Caddy, every player's
peer address is Caddy's, so all players would share one bucket: 20 requests per second
for the whole game.
Add to SCRUM-298 (or a sub-task):
- A gateway setting such as
GATEWAY_TRUSTED_PROXIES(CIDRs, e.g. Caddy's container address or the compose subnet). Only when the peer is in that list, take the client IP from the rightmostX-Forwarded-Forentry that Caddy appended; otherwise keep using the peer address. Never trust the header from anyone else. - Use that client IP for rate limiting and in access logs; keep the proxy's existing
"replace
X-Forwarded-Forwith the peer" behaviour for upstreams, feeding it the resolved client IP instead. - Tests: spoofed
X-Forwarded-Forfrom an untrusted peer is ignored; two clients behind the proxy get separate buckets.
3.4 Done when¶
curl https://<domain>[:port]/patch/v1/live/manifestworks from outside the host.- Plain HTTP is redirected (option A) or closed.
- A login from outside returns
services.sessionwith the public URL. - Two external clients are rate-limited independently.
- Certificates survive
up.shand container restarts without being re-issued.
4. SCRUM-290…295: the Allocator¶
4.1 Shape (SCRUM-290)¶
Build it like the other Go services (copy the layout of services/patch: internal/config,
internal/server, internal/api, internal/store, migrations, subcommands serve,
migrate, genkey), with these differences:
- No public route. It listens only on the compose network. It must not appear in either gateway's route table.
- Its own Postgres database and role (
allocator,allocator_rw), added todeploy/postgres/init/01-provision.shanddeploy/scripts/provision-upgrade.sh(the existing host's data directory never re-runs the init script). - Compose:
allocator-migrate, thenallocator, with secrets mounted like Auth's; theotomo-allocatorJenkins job already exists (ci/services/otomo-allocator/). /healthz,/readyz(startup done + database ping),/metrics, JSON logs, request IDs, the shared error body{"error":{"code","message","request_id"}}.
4.2 State (SCRUM-291)¶
Postgres is enough for M1 and survives restarts:
game_server (
server_id text primary key, -- chosen by the game server, e.g. its container name
internal_addr text not null, -- host:port on otomo-net (the proxy forwards here)
capacity smallint not null,
state text not null check (state in ('free','reserved','busy','dead')),
allocation_id uuid,
last_heartbeat timestamptz not null
)
allocation (
allocation_id uuid primary key,
party_id uuid not null,
server_id text not null references game_server(server_id),
player_ids uuid[] not null,
status text not null check (status in ('reserved','active','ended','expired')),
created_at timestamptz not null default now(),
expires_at timestamptz not null -- reserved and nobody connected by then → expired
)
-- one live allocation per party, which makes POST /internal/allocations idempotent
create unique index allocation_live_party on allocation(party_id) where status in ('reserved','active');
- Game servers heartbeat every 5 s. A reaper (every 5 s) marks servers with no heartbeat for
15 s
dead, and ends their allocations. - Reserve with
SELECT … FROM game_server WHERE state = 'free' AND last_heartbeat > now() - interval '15 seconds' ORDER BY server_id LIMIT 1 FOR UPDATE SKIP LOCKED, so concurrent allocations can never pick the same server.
4.3 Internal API (SCRUM-292, SCRUM-294)¶
Every call carries Authorization: Bearer <service key> (decision D4: one static key per
caller, mounted as a secret, compared in constant time). The key identifies the caller, so
Session can't call the game-server endpoints and vice versa.
| Caller | Call | Answer |
|---|---|---|
| Session | POST /internal/allocations {party_id, player_ids} |
201 {allocation_id, address, port, tickets: {player_id: ticket}}; the same party_id again → the same allocation (200); 503 no_capacity |
| Game server | POST /internal/servers/register {server_id, internal_addr, capacity} |
204; re-registering resets the server to free |
| Game server | POST /internal/servers/{server_id}/heartbeat {players_connected} |
204; first connection moves reserved → active/busy |
| Game server | POST /internal/servers/{server_id}/ended |
204; allocation ended, server free |
| Anyone internal | GET /.well-known/jwks.json |
the ticket-signing public key(s) (§4.4) |
address/port in the allocation answer are the Gameplay Proxy's public address, since
clients always connect through the proxy (decision D3). The proxy finds the real server from
the ticket.
Telling Session (SCRUM-294): when an allocation ends, expires, or its server dies, call
Session's internal endpoint (to be defined with Isaac in the design doc, e.g.
POST /internal/session/allocations/{allocation_id}/ended {reason}, with Session's key for
the Allocator). Session should also be able to ask GET /internal/allocations?party_id= so a
missed callback is repaired by polling.
4.4 Join tickets (SCRUM-293)¶
Signed exactly like Auth's tokens (Ed25519, golang-jwt; reuse the approach of
services/auth/internal/token), so the same verification code works in the proxy:
| Claim | Value |
|---|---|
header alg/kid |
EdDSA / the Allocator's key id (allocator genkey -kid) |
iss |
https://allocator.otomo.internal |
aud |
otomo:gameserver |
sub |
the player's ID (the Auth sub) |
alloc |
allocation ID |
srv |
game server ID (the proxy routes on this) |
jti |
random ID; single use: the proxy remembers used jtis until they expire |
exp |
issue time + 60 s |
The private key is generated out of band with genkey (never at startup), like Auth's. The
proxy and game servers fetch /.well-known/jwks.json over the internal network.
4.5 Metrics (SCRUM-295)¶
allocator_servers{state} (gauge), allocator_allocations_total{result="ok|no_capacity|error"},
allocator_allocation_seconds (histogram), allocator_reaped_total. The Dashboard picks
the service up from Prometheus once it is scraped (add it to deploy/observability).
5. SCRUM-271: what the design doc must settle with Isaac and Lew¶
- Lobby states (
forming → launching → in_game → forming) and what leave/kick/ready do in each. - Exact request and response bodies for every call in §4.3, including the Session callback.
- Timeouts: allocation call (e.g. 5 s), reservation expiry (60 s), heartbeat (5 s) and death (15 s).
- The client events Session sends (
party.launchingwith address, port and that member's ticket;party.launch_failed;party.returned), so Lew's SDK and doc 12 §8 can be final. - The proxy protocol Lew builds: how the client presents its ticket on connect (first packet or handshake), and what the proxy does on failure.
- Failure paths: no capacity, Session restarting mid-launch, a member who never connects, a server dying mid-expedition.
6. Other tickets currently assigned to you¶
Not in your stated area. Reassign or keep:
| Ticket | What | Natural owner |
|---|---|---|
| SCRUM-282, 283, 285 | Server-only Config: server manifest in releases, admin UI composer, session.rules seed |
Config/admin-UI owner, or Isaac (it feeds Session) |
| SCRUM-299 | Player gateway route review (rate-limit buckets for heartbeat/long-poll) | You (same code as §3.3) or Isaac |
| SCRUM-300 | Observability for the player services | You (orchestration) |
| SCRUM-301, 302 | End-to-end acceptance client, load tests | Shared; whoever finishes last |
| SCRUM-303 | Final docs pass | Whoever closes M1 |