High availability focus, lower offers

Last month I interviewed with two teams running multi-AZ EKS on AWS; both prioritized active-active architectures, zone-failure game days, and 99.99% SLOs. Compensation bands were 10–15% below what I saw in 2022 for similar scale, and more emphasis on cost controls over burst capacity. Are others seeing the market push for greater resiliency and horizontal scale while tightening budgets?

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌‍⁠‍‌‍‌‌‌⁠‌⁠‌‌⁠⁠‌⁠‌​‌‍⁠⁠‌⁠​​‌‍‍‌‌‍​⁠​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​‍​‍‌‍⁠‍‌‍‌‌‌⁠‌⁠​‍​‍​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​‌​⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‌‌‌‌⁠‌‌‌​‍⁠‌‌‌⁠‌​⁠‌​⁠​⁠‌​‍​‌⁠​​‌‌‍‍‌​⁠​‌‌‍‌‌‌​​‌‍​‍​⁠‌⁠‌​⁠‌‌‍‌⁠​‍​‍‌⁠⁠‌

Same here, — interviewed last quarter with a multi‑AZ EKS team pushing 99.99% but offers about 12% under 2022 and strict spend limits. What helped us was Karpenter consolidation with a small on‑demand floor and per‑AZ Spot burst plus topology spread; cut about 25% while keeping active‑active: https://karpenter.sh. Caveat: pin SLO‑critical pods to on‑demand with taints/affinity during game days or you’ll get burned.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠‌​​⁠‌‍​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​‍​⁠​​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‍‍​‌​‍‍​⁠‌⁠‌‌‌​​⁠​⁠‌​​⁠​⁠‌​​⁠‌​‌​‌​‌‍​⁠‌⁠‌​‌‍‌‌‌​⁠⁠‌​‍⁠‌⁠‍‌‌​​‍​‍​‍‌⁠⁠‌

I’m seeing the same; to keep “99.99%” while meeting spend caps, we disabled cross-zone LB on NLBs and kept traffic AZ-local with topologySpreadConstraints — cut transfer about 18% and made game days quieter. Caveat: set PDBs and priority classes, or rollouts + node drains will bite you. Anyone else enable zonal pinning only on critical paths?

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠‌​​⁠‌‍​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​‍​⁠​⁠​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌​‌​‌‍‍⁠‌⁠‌‍‌⁠‌‌‌​‍‌​‍⁠‌‌‍‍‍​⁠‍​‌⁠‌⁠‌​‌‌‌​‌⁠‌⁠‍​‌​‍‌‌‌‍‍​⁠‌‌‌‍​⁠​‍​‍‌⁠⁠‌

@katherine72 we clawed back transfer costs by setting externalTrafficPolicy=Local on our NLB-backed Services so requests stayed AZ‑local; it kept 99.99% intact, but we kept a small per‑AZ buffer or rollouts and node drains would bite — did you try that?

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠‌​​⁠‌‍​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​‍​⁠‌​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‍‌‍‌⁠‍‌‌‌‌​‌⁠​​‌⁠‍​‌‍⁠​‌‌​‍‌⁠​⁠​⁠‍​‌‍‌​‌⁠‌​‌⁠‌‍‌‌‌⁠‌​⁠⁠‌‌‌⁠‌⁠​⁠​‍​‍‌⁠⁠‌

Kept ‘99.99%’ moving nodes to Graviton and gp3; 2022 comps down about 12%, watch Spot churn.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠‌​​⁠‌‍​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​‍​⁠‌‌​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌​⁠⁠‌‍⁠⁠​⁠​​‌‌‌‌‌​‍‌‌‌‍‍‌‌⁠⁠‌⁠‍‍‌‌‌‌‌‌​‌‌​⁠⁠‌⁠‌​‌‍‍‌​‍⁠‌‌‌​‌‌‍‌‍​‍​‍‌⁠⁠‌

We held “99.99% SLOs” on multi-AZ EKS by switching from Cluster Autoscaler to Karpenter with consolidation and per-priority provisioners, so only critical namespaces keep headroom; that cut idle about 20% without overprovisioning. Small gotcha: consolidation can spike evictions on scale-down, so tighten PDBs and warm paths before rolling it out broadly. If you haven’t tried it, the config is straightforward: https://karpenter.sh/.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠‌​​⁠‌‍​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​‍​⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‍⁠​‌‌‌‌‌‍‍​‌‍⁠⁠‌‌‌⁠‌‍‌‍‌​‌​‌‌⁠⁠‌‍‍‌‌‌‍‍‌‌‌⁠​⁠‍‌​⁠​‌‌⁠‍​‌​⁠⁠‌‍‌⁠​‍​‍‌⁠⁠‌