Files
HexaHost-GameCloud/docs/operations/incident-response.md

2.8 KiB

Incident response

Operational playbook for security and availability incidents affecting HexaHost GameCloud.

Severity levels

Level Examples Response time Examples
SEV-1 API down, data breach suspected, all nodes offline 15 min Full platform outage
SEV-2 Single node down, WHMCS provision failures, DNS sync broken 1 h Partial customer impact
SEV-3 Elevatoration, non-critical bug, one customer server stuck Next business day Limited blast radius

Initial response (all severities)

  1. Acknowledge — Assign incident commander and scribe
  2. Triage — Customer-facing vs internal-only impact
  3. Communicate — Internal channel + status page if SEV-1/2
  4. Preserve evidence — Do not reboot hosts before log capture if security-related

Common incidents

API unavailable (502/503)

docker compose -f deploy/compose/compose.prod.yml ps
docker compose -f deploy/compose/compose.prod.yml logs --tail=200 api
curl -v https://api.example.net/api/v1/health/live

Check Traefik routing (traefik), PostgreSQL connectivity, disk full.

Worker queue backlog

Inspect Redis queue depth and failed jobs in logs. Scale worker container or restart after fixing root cause (DNS provider down, S3 errors).

Node heartbeat stale

Nodes not seen > 90s trigger alerts. Verify node-agent systemd status, outbound firewall to wss://api.example.net/api/v1/nodes/ws, certificate expiry.

Suspected credential leak

  1. Rotate affected secrets immediately (secrets):
    • SESSION_SECRET (invalidates all sessions)
    • WHMCS_API_SECRET per installation
    • NODE_TOKEN for compromised node
    • S3_ACCESS_KEY / database password
  2. Review audit logs and WHMCS module call logs
  3. Enable mTLS if not already active for WHMCS integration

Tenant isolation concern (SEV-1)

  1. Disable affected API endpoints if exploit is active
  2. Capture request IDs from audit log
  3. Engage security lead; reference threat model T21

WHMCS-specific

Provisioning failures often appear in WHMCS Utilities → Logs → Module Log. Cross-check GameCloud API logs with X-HGC-Integration-Id header.

Run reconciliation dry-run before live fixes: reconciliation.

Post-incident

Within 5 business days:

  • Timeline document (UTC)
  • Root cause (5 whys)
  • Action items with owners
  • Update runbooks if gaps found

Contacts and escalation

Document internally:

  • On-call rotation
  • WHMCS admin contact
  • DNS provider support
  • Hosting provider (game nodes)