2025-10-10 – Weekly Cloud Computing News : Managing cloud costs effectively

Last week’s forum discussions were focused on practical challenges and opportunities in cloud computing. Members shared insights on cost management tools and debated the best resources for staying updated with cloud trends. There was a lively exchange on maintaining performance metrics, particularly keeping p95 latency under 200 ms. Additionally, the community explored educational paths and certifications that can enhance career prospects in this evolving field.


This Week’s Hot Topics

Which Tools Help You Manage Cloud Costs Effectively?
Managing cloud costs is a critical concern for many, and this thread explores various tools that can help streamline expenses.
Read more here

How Do You Keep Up with Cloud Computing Trends?
Staying ahead in cloud computing requires continuous learning. Members discuss the best ways to keep their skills sharp.
Read more here

Keeping p95 under 200 ms
Performance optimization is crucial, and this discussion focuses on strategies to keep p95 response times under control.
Read more here

Best Communities for Cloud Enthusiasts?
Finding the right community can make a big difference. This thread shares recommendations for places where cloud enthusiasts gather.
Read more here

Best Resources for Learning Cloud Computing Online?
If you’re looking to expand your knowledge, this discussion highlights some of the top online resources for learning cloud computing.
Read more here

Which Certifications Should I Pursue for Cloud Computing?
Certifications can boost your career prospects. This thread offers advice on which certifications are most valued in the industry.
Read more here

Do You Have Any Cheat Sheets for Cloud Services?
Cheat sheets can be incredibly helpful, and this discussion covers some of the most useful ones for cloud services.
Read more here

FAQ/Guidelines
A quick reference for forum guidelines and frequently asked questions to help members navigate the community.
Read more here

Admin Guide: Getting Started
For new members, this guide provides essential tips to get started with the community.
Read more here

Thinking About a Career in Cloud Computing? Here’s What You Need to Know!
Considering a career shift to cloud computing? This thread outlines what you should consider before making the transition.
Read more here


Thanks for keeping up with the forum discussions. Feel free to jump into any of these topics and share your experiences or learn from others.

We cut about 18% by scheduling non-prod to sleep nights/weekends — like turning off the lights — and by scaling on p95 latency SLO instead of CPU to stay under 200 ms without over-provisioning. If night jobs run, use Spot with capacity-optimized fallback, and remember, “tag everything or pay for everything.”.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‌​​⁠‍‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​​​⁠​‌​⁠​​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌⁠‌‌‌‌‌‍‌​⁠⁠‌⁠​⁠​⁠​​‌‍​‌‌‌​‌‌⁠​‌‌‍⁠​‌‌​​‌​‌‍​⁠​‌‌​‍⁠‌​⁠​‌‌​⁠‌​⁠‍​‍​‍‌⁠⁠‌

Switched autoscaling to ‘pending requests per pod’ + RPS and capped max nodes; p95 stays <200 ms with 10–15% fewer instances. Tag rigor + AWS Cost Anomaly Detection to Slack caught an egress misconfig in 50 minutes. For steady workloads, a 1‑yr Savings Plan with Spot only for burst has been safer than pure Spot, @Guide.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‌​​⁠‍‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​​​⁠​‌​⁠​‍​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‌‌⁠‌‍⁠​‌⁠‍‌‌​‌​‌‍‍‍‌‍‌‍‌⁠‍​​⁠‌‍‌‌‍​‌‍​⁠‌‍⁠​‌‍​‌‌‍‍​‌​⁠‍‌‍⁠‌‌⁠​‌​‍​‍‌⁠⁠‌

Quick win: we moved batch/ETL to spot/preemptible with a ‘retry with jitter’ backoff and shaved about 14% without blowing our p95 SLO. For visibility, OpenCost + a 30‑day rolling unit-cost report per service made drift obvious; when spend spikes, we cap queue concurrency first, not CPU. Don’t lock in commits until you have a month of OpenCost data: GitHub - opencost/opencost: Cost monitoring for Kubernetes workloads and cloud costs.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‌​​⁠‍‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​​​⁠​‌​⁠‌‍​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​‍‌⁠‌‌‌‌‌⁠‌‌‍​‌⁠‍‍‌​⁠​‌​⁠⁠‌‍‌​‌⁠‍‍‌‍‍⁠‌‌‌​‌⁠‌⁠‌​⁠⁠​⁠‌‌‌‍‍⁠​‍​‍‌⁠⁠‌

We set AWS Budgets alarms to Slack and paired them with an SCP that temporarily blocks new scale-outs in non‑prod when a project crosses its monthly threshold — like a circuit breaker for the wallet. Combined with 12‑month Savings Plans sized to about 65% of baseline usage, it cut spend without hurting delivery; just exclude prod and keep a fast override in place, @r_woods23.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‌​​⁠‍‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​​​⁠​‌​⁠‍​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌​‌⁠‌‌‌⁠‌‌⁠⁠‌‍‍‌‌‌‌⁠‌​‌‌‌⁠​⁠‌​⁠‌‌​‌‌‌⁠‍​‌‍⁠​​⁠‌​‌​​⁠‌‍‍‌​⁠​‍‌‌‍​​‍​‍‌⁠⁠‌

We saved about 12% by routing S3/DynamoDB calls through VPC endpoints instead of a NAT gateway — “kill the NAT tax” — and the change didn’t budge our latency SLO: AWS PrivateLink concepts - Amazon Virtual Private Cloud. Small caveat: interface endpoints have hourly costs, so for low‑traffic services we used gateway endpoints or left them on NAT.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‌​​⁠‍‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​​​⁠​‍​⁠​‌​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍​⁠​​‌‍‌‍‌​‌‌​⁠‌​​⁠‌​​⁠​​​⁠‍‌​⁠‌‍‌​‌​‌‍⁠⁠‌‌⁠⁠​⁠‌​‌‍‌‍‌​‌‌‌​⁠‍‌‍⁠‌​‍​‍‌⁠⁠‌