Files
HexaHost-GameCloud/docs/operations/node-drain.md

2.6 KiB

Game node drain

Gracefully remove a game node from scheduling before maintenance, decommissioning, or migration.

When to drain

  • OS kernel / Docker upgrade on the node
  • Hardware maintenance or rack move
  • Shrinking cluster capacity
  • Replacing a node with new hardware (server migration)

Do not drain the last node in a region if active servers have no migration target.

Procedure

1. Set node to DRAINING

Via admin API:

PATCH /api/v1/admin/nodes/{nodeId}
Authorization: Bearer <admin-token>
Content-Type: application/json

{ "status": "DRAINING" }

The scheduler stops placing new servers on this node. Existing running servers continue until stopped by users or idle shutdown.

2. Wait for natural shutdown (optional)

If maintenance window allows, wait for idle shutdown policy to stop inactive servers. Monitor:

SELECT id, status, "nodeId" FROM "GameServer" WHERE "nodeId" = '<nodeId>' AND status = 'RUNNING';

3. Migrate or stop remaining servers

For each running server:

  • Preferred: Customer stops server via panel; reprovision on another node on next start (scheduler picks eligible node)
  • Admin: Force stop via admin API, then start — triggers reschedule if original node is DRAINING
  • Maintenance: Communicate window to affected customers

4. Confirm empty node

SELECT COUNT(*) FROM "GameServer" WHERE "nodeId" = '<nodeId>' AND status IN ('RUNNING', 'STARTING');

Expected: 0 before maintenance on Docker host.

5. Stop node-agent

sudo systemctl stop hgc-node-agent

Configured in deploy/systemd/node-agent.service.

6. Perform maintenance

Apply updates via deploy/ansible/game-node.yml or manual steps. See game node operations.

7. Return to service or decommission

Return:

PATCH /api/v1/admin/nodes/{nodeId}
{ "status": "ONLINE" }
sudo systemctl start hgc-node-agent

Decommission:

  1. Revoke enrollment token in control plane
  2. Remove GameNode row or mark OFFLINE permanently
  3. Wipe /var/lib/hgc-node data if disks are reused

Edge cases

Situation Action
Server stuck STARTING during drain Admin stop + investigate node-agent logs
Port range exhausted on other nodes Add node or expand range before drain
Node loses connectivity mid-drain Servers marked UNKNOWN; reconcile when agent returns

Automation

Future: automated drain API with --wait-empty timeout. Until then, use SQL + admin API as above.