I’m evaluating Cilium’s eBPF datapath with kube-proxy replacement to reduce east-west overhead across three EKS clusters… On Calico/iptables we’re seeing about 42 ms p99 cross-AZ at about 18k rps; a small canary with Cilium (native routing, no tunneling) brought p99 to about 27–29 ms — any long-term gotchas at 500–1,000 nodes per cluster, especially around failure domains, NodeLocal DNS, and sidecar CPU due to fewer conntrack hits?
We saw similar gains on EKS; at about 650 nodes/cluster, turning on Maglev for the kube-proxy replacement (tableSize 65537) kept p99 steady during AZ drains and node churn — nice complement to your “native routing, no tunneling” setup. Small caveat: bump conntrack map sizes (e.g., bpf.ctGlobalTCPMax) or the GC can spike at about 18k rps.
At about 700 nodes on EKS, kube-proxy-free + native routing held p99, but we hit BPF map pressure during AZ drains; bump ct-global-max and lb-map-max and set kubeProxyReplacement=strict. @rjensen71’s Maglev note is spot on; we also flipped bpf-lb-acceleration=xdp and enabled endpoint routes, which kept tail latency flat around your 18k rps. Minor caveat: keep Hubble to sampling and confirm rp_filter=0, or churn + tracing will bite you at 500–1,000 nodes.
We got bitten later when we linked three EKS clusters with ClusterMesh — overlapping PodCIDRs led to bizarre blackholes during node failures; “make PodCIDRs unique across clusters” and plan VPC routes up front (): Multi-cluster Networking — Cilium 1.18.6 documentation. For your cross-zone p99, enabling topology-aware hints kept most traffic local and trimmed a few ms for us; tiny caveat: headless/StatefulSets don’t always follow hints cleanly.
Quick tip: we moved from ENI IPAM to Cilium’s cluster-pool after hitting ENI attach-rate limits during node churn; per-node /25s plus Maglev (agree with @james_34) kept p99 flat and avoided API flaps across AZs — https://docs.cilium.io/en/stable/network/ipam/cluster-pool/. Caveat: size the pool conservatively and set bpf-map-dynamic-size-ratio to keep map memory in check; fewer AWS API calls, fewer chances to trip over your own shoelaces.