# Disaster recovery Recovery objectives and procedures when a HexaHost GameCloud control plane or game node is lost. ## Recovery objectives (default targets) | Tier | RPO | RTO | Notes | |------|-----|-----|-------| | Control plane (API, DB) | 1 h | 4 h | With WAL + hourly Redis snapshot | | Object storage (backups) | 15 min | 4 h | With cross-region replication | | Single game node | 0 (running servers) | 30 min | Reschedule + DNS unchanged if edge IP stable | | Full region loss | 24 h | 24 h | Requires warm standby or restore from off-site | Adjust these in your internal SLA and document approved deviations. ## Failure scenarios ### Control plane host lost 1. Provision replacement via `deploy/ansible/control-plane.yml` 2. Restore PostgreSQL from latest dump + WAL ([backup-restore](backup-restore.md)) 3. Restore MinIO data or repoint `S3_*` to replica bucket 4. Deploy compose: `deploy/compose/compose.prod.yml` 5. Restore Traefik `acme.json` or re-issue certificates 6. Verify WHMCS integration (HMAC secret unchanged) ### Database corruption 1. Stop writes (`api`, `worker`) 2. Restore to last known-good PITR timestamp 3. Run WHMCS reconciliation dry-run from addon module 4. Review audit log gap for the corruption window ### Object storage unavailable - API continues serving metadata; backup restore and file uploads fail - Fail over to replica endpoint or restore MinIO volume from snapshot - Workers retry S3 operations with exponential backoff ### All game nodes offline - Customers cannot start servers; running containers on nodes may continue until stopped - Edge gateway returns connection errors for join-to-start - Restore nodes in priority order; scheduler assigns new provisions when nodes return ### WHMCS host lost GameCloud remains authoritative for technical state. Reinstall WHMCS module from `integrations/whmcs/`, restore WHMCS database, re-import product mappings from backup export. ## Standby strategy (recommended) | Component | Standby approach | |-----------|------------------| | PostgreSQL | Streaming replica or managed HA | | Redis | Sentinel or managed Redis | | S3 | Cross-region replication | | Control plane VMs | IaC templates in `deploy/ansible/` | | Traefik certs | Shared `acme.json` on persistent volume | ## Communication During DR: 1. Status page update (if applicable) 2. Internal war room with roles: coordinator, DB, network, WHMCS 3. Customer comms after RTO estimate is known — WHMCS can send mass email ## DR drill checklist (quarterly) - [ ] Restore PostgreSQL dump to staging environment - [ ] Restore sample backup from object storage to test server - [ ] Fail over DNS to standby edge IP (dry run) - [ ] WHMCS provisioning test order on staging - [ ] Document actual RPO/RTO vs targets ## Related - [Backup and restore](backup-restore.md) - [Server migration](server-migration.md) - [Incident response](incident-response.md)