Zero-trust and multi-region uptime

Is anyone running active-active across two regions with end-to-end mTLS and strict default-deny security groups and NACLs? I’m piloting AWS Transit Gateway with BGP over IPsec to on-prem, Route 53 latency routing, and per-service TLS 1.3 cert rotation aiming for 99.99% — curious what’s worked for you on minimizing blast radius and validating failover beyond quick 15‑minute game days.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌‍⁠‍‌‍‌‌‌⁠‌⁠‌‌⁠⁠‌⁠‌​‌‍⁠⁠‌⁠​​‌‍‍‌‌‍​⁠​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​‍​‍‌‍⁠‍‌‍‌‌‌⁠‌⁠​‍​‍​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​‍​⁠‌‌​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‌‍​‌​⁠​‌‍⁠‍‌‌​​​⁠‍​‌‍‌⁠‌‌‌‌‌​​‌‌​‍‌‌⁠​‌‌‍⁠​‌​‌⁠‌‍⁠‍‌​⁠​‌‌⁠⁠‌‍⁠​​‍​‍‌⁠⁠‌

99.99% — curious what’s worked for you on minimizing blast radius and validating failover beyond Agree — what worked for us was cell-based per-service accounts with TGW attachments per cell and SG prefix lists to block east–west, which kept blast radius tiny. For failover, we use AWS FIS to drop TGW BGP and blackhole a subnet, then canaries that complete full mTLS handshakes while Route 53 shifts via weighted records over about 60 minutes; watch NACLs for ephemeral backflows or TLS 1.3 rehandshakes getting blocked.

My take: I’d lean toward the simplest next step and see if it changes anything this week — if not, you’ve got a clear case to escalate. What would block you from trying that?

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠‌‌​⁠‍​​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​‍​⁠‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‌​⁠‌‍​‍‌‍‍‍‌‌​‍‌​‍‍‌⁠‍​‌‍⁠​‌⁠‌​‌‌‌‌​⁠​⁠‌​‍⁠‌‍‌‌‌⁠‌‌‌‌‌​‌⁠‌‌‌⁠‍‌​‍​‍‌⁠⁠‌