85 lines
2.8 KiB
Markdown
85 lines
2.8 KiB
Markdown
# Incident response
|
|
|
|
Operational playbook for security and availability incidents affecting HexaHost GameCloud.
|
|
|
|
## Severity levels
|
|
|
|
| Level | Examples | Response time | Examples |
|
|
|-------|----------|---------------|----------|
|
|
| SEV-1 | API down, data breach suspected, all nodes offline | 15 min | Full platform outage |
|
|
| SEV-2 | Single node down, WHMCS provision failures, DNS sync broken | 1 h | Partial customer impact |
|
|
| SEV-3 | Elevatoration, non-critical bug, one customer server stuck | Next business day | Limited blast radius |
|
|
|
|
## Initial response (all severities)
|
|
|
|
1. **Acknowledge** — Assign incident commander and scribe
|
|
2. **Triage** — Customer-facing vs internal-only impact
|
|
3. **Communicate** — Internal channel + status page if SEV-1/2
|
|
4. **Preserve evidence** — Do not reboot hosts before log capture if security-related
|
|
|
|
## Common incidents
|
|
|
|
### API unavailable (502/503)
|
|
|
|
```bash
|
|
docker compose -f deploy/compose/compose.prod.yml ps
|
|
docker compose -f deploy/compose/compose.prod.yml logs --tail=200 api
|
|
curl -v https://api.example.net/api/v1/health/live
|
|
```
|
|
|
|
Check Traefik routing ([traefik](traefik.md)), PostgreSQL connectivity, disk full.
|
|
|
|
### Worker queue backlog
|
|
|
|
Inspect Redis queue depth and failed jobs in logs. Scale worker container or restart after fixing root cause (DNS provider down, S3 errors).
|
|
|
|
### Node heartbeat stale
|
|
|
|
Nodes not seen > 90s trigger alerts. Verify node-agent systemd status, outbound firewall to `wss://api.example.net/api/v1/nodes/ws`, certificate expiry.
|
|
|
|
### Suspected credential leak
|
|
|
|
1. Rotate affected secrets immediately ([secrets](../security/secrets.md)):
|
|
- `SESSION_SECRET` (invalidates all sessions)
|
|
- `WHMCS_API_SECRET` per installation
|
|
- `NODE_TOKEN` for compromised node
|
|
- `S3_ACCESS_KEY` / database password
|
|
2. Review audit logs and WHMCS module call logs
|
|
3. Enable mTLS if not already active for WHMCS integration
|
|
|
|
### Tenant isolation concern (SEV-1)
|
|
|
|
1. Disable affected API endpoints if exploit is active
|
|
2. Capture request IDs from audit log
|
|
3. Engage security lead; reference [threat model](../security/threat-model.md) T21
|
|
|
|
## WHMCS-specific
|
|
|
|
Provisioning failures often appear in WHMCS **Utilities → Logs → Module Log**. Cross-check GameCloud API logs with `X-HGC-Integration-Id` header.
|
|
|
|
Run reconciliation dry-run before live fixes: [reconciliation](../integrations/whmcs/reconciliation.md).
|
|
|
|
## Post-incident
|
|
|
|
Within 5 business days:
|
|
|
|
- [ ] Timeline document (UTC)
|
|
- [ ] Root cause (5 whys)
|
|
- [ ] Action items with owners
|
|
- [ ] Update runbooks if gaps found
|
|
|
|
## Contacts and escalation
|
|
|
|
Document internally:
|
|
|
|
- On-call rotation
|
|
- WHMCS admin contact
|
|
- DNS provider support
|
|
- Hosting provider (game nodes)
|
|
|
|
## Related
|
|
|
|
- [Disaster recovery](disaster-recovery.md)
|
|
- [Threat model](../security/threat-model.md)
|
|
- [WHMCS troubleshooting](../integrations/whmcs/troubleshooting.md)
|