2.8 KiB
2.8 KiB
Incident response
Operational playbook for security and availability incidents affecting HexaHost GameCloud.
Severity levels
| Level | Examples | Response time | Examples |
|---|---|---|---|
| SEV-1 | API down, data breach suspected, all nodes offline | 15 min | Full platform outage |
| SEV-2 | Single node down, WHMCS provision failures, DNS sync broken | 1 h | Partial customer impact |
| SEV-3 | Elevatoration, non-critical bug, one customer server stuck | Next business day | Limited blast radius |
Initial response (all severities)
- Acknowledge — Assign incident commander and scribe
- Triage — Customer-facing vs internal-only impact
- Communicate — Internal channel + status page if SEV-1/2
- Preserve evidence — Do not reboot hosts before log capture if security-related
Common incidents
API unavailable (502/503)
docker compose -f deploy/compose/compose.prod.yml ps
docker compose -f deploy/compose/compose.prod.yml logs --tail=200 api
curl -v https://api.example.net/api/v1/health/live
Check Traefik routing (traefik), PostgreSQL connectivity, disk full.
Worker queue backlog
Inspect Redis queue depth and failed jobs in logs. Scale worker container or restart after fixing root cause (DNS provider down, S3 errors).
Node heartbeat stale
Nodes not seen > 90s trigger alerts. Verify node-agent systemd status, outbound firewall to wss://api.example.net/api/v1/nodes/ws, certificate expiry.
Suspected credential leak
- Rotate affected secrets immediately (secrets):
SESSION_SECRET(invalidates all sessions)WHMCS_API_SECRETper installationNODE_TOKENfor compromised nodeS3_ACCESS_KEY/ database password
- Review audit logs and WHMCS module call logs
- Enable mTLS if not already active for WHMCS integration
Tenant isolation concern (SEV-1)
- Disable affected API endpoints if exploit is active
- Capture request IDs from audit log
- Engage security lead; reference threat model T21
WHMCS-specific
Provisioning failures often appear in WHMCS Utilities → Logs → Module Log. Cross-check GameCloud API logs with X-HGC-Integration-Id header.
Run reconciliation dry-run before live fixes: reconciliation.
Post-incident
Within 5 business days:
- Timeline document (UTC)
- Root cause (5 whys)
- Action items with owners
- Update runbooks if gaps found
Contacts and escalation
Document internally:
- On-call rotation
- WHMCS admin contact
- DNS provider support
- Hosting provider (game nodes)