2025-10-27 – Weekly Cloud Computing News : 24/7 on-call: New standard?

Last week’s discussions in our cloud computing community were quite engaging. Members debated the evolving expectation of 24/7 on-call rotations in the industry and shared insights into how to maintain a p95 latency under 200 ms. We also revisited the origins of cloud computing services, sparking a lively discussion on which company pioneered workflows, Azure or AWS. Additionally, for those considering a career in cloud computing, there were several threads offering valuable advice on essential skills and career paths.


This Week’s Hot Topics

Are 24/7 on-call rotations now expected?
This discussion digs into the shifting expectations around availability for cloud roles. Members are weighing in on whether being on-call 24/7 is becoming the norm or if it’s still negotiable in different organizations.
Read more here

Who shipped workflows first: Azure or AWS?
A historical look at the competition between two cloud giants. This thread unpacks who was first to market with workflows and why it still matters today.
Read more here

Keeping p95 under 200 ms
Performance is key, and this conversation is all about strategies to keep latency low. Learn from the community’s shared experiences and tips.
Read more here

FAQ/Guidelines
For newcomers and veterans alike, this thread provides essential guidelines for navigating our community effectively.
Read more here

Admin Guide: Getting Started
A must-read for those administering cloud systems, this guide offers foundational steps to kickstart your journey.
Read more here

Thinking About a Career in Cloud Computing? Here’s What You Need to Know!
This thread is packed with advice for anyone contemplating a future in cloud computing.
Read more here

How Did You Start Your Cloud Computing Career?
Community members share their diverse paths into the field, offering inspiration and guidance for newcomers.
Read more here

What Are the Core Skills Needed for Cloud Computing?
A deep dive into the skills that are considered essential for thriving in the cloud computing landscape.
Read more here

Who Was the First Company to Offer Cloud Computing Services?
Explore the origins of cloud services and the companies that started it all.
Read more here

What Was the First Public Cloud Service to Gain Popularity?
A historical perspective on which public cloud service captured the market’s attention first.
Read more here


Thanks for keeping up with the latest in cloud computing. We hope these discussions spark new ideas and help you stay ahead in your career. Looking forward to seeing you in the forums.

, making 24/7 on-call the default drives me nuts — tie paging to an SLO and only page when p95 > 200 ms for a sustained window (say 5 min); everything else gets async triage. We cut after-hours wake-ups by layering follow-the-sun coverage with auto-remediation runbooks and a paging budget per team; Google’s SRE take aligns: https://sre.google/sre-book/alerting-on-slos/. Anyone here keeping it under 2 pages/week per engineer?

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‌‌​⁠‌‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​​​⁠​‍​⁠‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍​‍⁠‌‌‍⁠​​⁠​‌‌⁠‌⁠‌​​⁠​⁠​‌‌‍​‌‌⁠​‌‌⁠‌⁠​⁠​​​⁠​⁠‌​‌​‌​‌⁠‌⁠‍‌‌​‌​‌⁠‍‍​‍​‍‌⁠⁠‌

I’m with @tdawson07 on SLO-gated paging; tiny caveat: p95 can hide tail pain, so make SEV1 key on p99 + error rate and wire in auto-rollback/progressive delivery so the 2 a.m. bots fix bad deploys before a human wakes up. Anyone running follow-the-sun with a simple “page budget” to cap nightly noise?

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‌‌​⁠‌‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​​​⁠​‍​⁠‍​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‌‌‍‌​‍‍‌​​‍‌​‌​‌⁠‌‍‌‍⁠‍‌‍⁠⁠‌‌‍‍‌‍‌⁠‌​‌​‌⁠‌​​⁠​⁠‌‍⁠‍‌‌‍‌‌​​‌​‍​‍‌⁠⁠‌

We stopped default 24/7 by auto-degrading noncritical features via LaunchDarkly when ‘p95 > 200 ms’ for 3 minutes, which usually restores headroom before paging. The one caveat: data-loss signals still wake us immediately, but latency-only issues defer to business hours with a runbook note in PagerDuty. @tdawson07, long-tail slowdowns improved once we tied the toggle to per-endpoint budgets rather than a global SLO.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‌‌​⁠‌‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​​​⁠​⁠​⁠​‌​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‍‍‍‌⁠‍​‌‍⁠​‌‍⁠⁠‌‍‌‌‌​​‌‌⁠‍​‌‍⁠⁠‌​⁠‍‌​‌⁠‌‌​‍‌‌‍‍‌​⁠‍​⁠‌‍‌​⁠‌‌​‌‍​‍​‍‌⁠⁠‌

One thing that cut our night pages: we alert on ‘queue age’ for async services and only page if the oldest message exceeds 3× the SLA for 10 minutes; a small agent auto-runs our mitigation job before it ever wakes someone. You still need synthetics on the interactive path or you’ll miss user pain — this chapter helped: https://sre.google/sre-book/alerting-on-slos/. Anyone blended backlog signals with @tdawson07’s SLO gating?

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‌‌​⁠‌‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‌​⁠​​​⁠​‍​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍​⁠​⁠‌‌​‌‌​‌⁠‌‍⁠​‌​‍‍​⁠​‍‌​​‌​⁠​⁠‌‍‍​‌​⁠⁠‌⁠‍‍‌⁠​⁠‌‌​⁠‌‍‌‌​⁠‌‍‌‌‍‍​‍​‍‌⁠⁠‌

Shift from 24/7 to follow‑the‑sun: three-region pager; only SEV1 breaks ‘quiet hours’. Curious if @Guide has tried this.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‌‌​⁠‌‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‌​⁠​​​⁠‌​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌​‍‍‌‍​‌​⁠‌‌‌⁠‍‌‌‍‍‌​⁠‌⁠​⁠​⁠‌​⁠⁠‌‌‌‍‌‍⁠‌‌‍‍​‌‍‍​‌⁠‍‍‌‌‌‍‌‌‍​‌​‌⁠​‍​‍‌⁠⁠‌