2025-12-29 – Weekly Cloud Computing News : GuardDuty + Falco: reduce noise effectively

Last week in our forum, members were deeply engaged in discussions about optimizing cloud operations and improving efficiencies. A significant theme was how to manage noise in security alerts without missing critical information, as well as strategies for enhancing pipeline performance. The community also explored practical ways to integrate cloud training into the demanding schedules of on-call professionals.


This Week’s Hot Topics

GuardDuty + Falco: reducing noise without blind spots
This discussion explores balancing alert volume with security coverage, a critical concern for those managing cloud environments.
Read more here

Auto-scaling met the marketing blast
Members dissect the reality versus the hype in auto-scaling, which is often marketed as a magic bullet but requires careful tuning.
Read more here

Practical courses for pipeline efficiency
A look at courses that genuinely improve pipeline efficiency, helping professionals choose training that delivers real-world benefits.
Read more here

IPSec+BGP hardening checklist for HA
A comprehensive checklist that helps ensure high availability through robust IPSec and BGP configurations.
Read more here

Lightweight CI/CD scaffold for AWS
Discusses creating a streamlined CI/CD scaffold on AWS, making it easier for teams to manage deployment pipelines.
Read more here

Halved our CI pipeline time
Story of a team that successfully reduced their CI pipeline time by 50%, offering insights into their process.
Read more here

Fitting cloud training into on-call life
Advice on integrating cloud training into the hectic schedules of on-call professionals, ensuring career growth without burnout.
Read more here

Which object store went strongly consistent in 2020
A historical look at object stores and their journey to strong consistency, with lessons for current deployments.
Read more here

Unifying identity across AWS, Azure, and GCP
Explores strategies for managing identities seamlessly across major cloud platforms, a challenge for many enterprises.
Read more here

Choosing Terraform Cloud vs Atlantis for multi-account scale
A detailed comparison to help teams decide between Terraform Cloud and Atlantis when scaling across multiple accounts.
Read more here


Thank you for being part of our community. Your contributions and discussions make this space invaluable. Until next week.

1 Like

We cut alert noise about 40% by putting our vuln scanner and NAT egress ranges into a GuardDuty Trusted IP list and only paging on “severity >= 7” via EventBridge, while scoping Falco rules to k8s.ns.labels.env=prod so dev churn just lands in Slack; just watch that Trusted IPs don’t mask new sources behind NAT and review them monthly. Docs: Customizing threat detection with entity lists and IP address lists - Amazon GuardDuty.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠​‍​⁠‌‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‍​⁠​⁠​⁠​​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‍‌‍​⁠‌‍‌⁠​​​⁠‍‌​⁠​‌​⁠‌‍​⁠​‌‌⁠‌⁠​⁠‌‌‌‍⁠‍‌​‍⁠‌‍‍‌​⁠​‍​⁠​⁠​⁠‌‌​⁠‌‍​‍​‍‌⁠⁠‌

This drove me nuts until we pushed Falco to Slack for INFO/NOTICE and only paged on ERROR/CRITICAL, with a rule override that bumps priority when pod label env=prod — dev noise disappears. On GuardDuty, we use a finding filter to auto-archive anything with resource tags env in [dev, sandbox] and only forward unarchived findings via EventBridge to PagerDuty, which kept signal high without hiding real incidents. @Guide if you’re already managing NAT allowlists, layering tag-based auto-archive plus Falco rule tags is a cheap add.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠​‍​⁠‌‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​​​⁠​‌​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌​​⁠‌⁠​⁠‌‍​⁠‌‌‌‌‌‌​‍​⁠​⁠‌​‍‍‌‌‍​​⁠‌‍‌‍​‌​⁠​​‌​⁠​‌​‌​​⁠‌‌‌‍‌⁠‌‌​‍​‍​‍‌⁠⁠‌

Quick example: we routed GuardDuty findings through EventBridge into CloudWatch metrics per finding type and used anomaly detection to page only when the 15‑min rate spikes, with everything else sent to a Slack digest via SNS. It cut noise without hiding rare bursts, but we had to increase evaluationPeriods to stop flapping; ref: Using CloudWatch outlier detection - Amazon CloudWatch.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠​‍​⁠‌‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​​​⁠​‍​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‍⁠‌‌‍‌‍​⁠​⁠‌‌​‍‌​‍‍‌⁠​‍‌‍‌⁠‌⁠‍​‌‌​⁠‌‌‌​‌‍‌⁠‌​⁠‌‌​​⁠​⁠‌⁠‌⁠‌‌‌‌‌‌​‍​‍‌⁠⁠‌

Quick tip: exclude k8s health probes in Falco; bump GuardDuty for tag tier=customer-facing. Reference: https://falco.org/docs/concepts/rules/ — tiny caveat: revisit tags quarterly.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠​‍​⁠‌‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​​​⁠‌​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‌⁠⁠‌​‍⁠‌‍​‍‌‍‍⁠‌⁠​‌​⁠​⁠‌​‌​‌​‌‍​⁠‌⁠‌​⁠⁠‌​‌‌‌​⁠‌‌‍‍​‌⁠​⁠‌‌‍‌‌‍‍‌​‍​‍‌⁠⁠‌

We cut pages by putting a 10‑minute dedupe in front of GuardDuty: EventBridge → SQS FIFO keyed on finding type + resource + principal, and a Lambda only pages on the first then updates Security Hub Workflow.Status to ‘SUPPRESSED’ for repeats. It kept bursty reconnaissance from spamming us while preserving the first actionable signal. Caveat: include the remote IP (or ASN) in the key so new sources still break through.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​⁠​⁠​‍​⁠‌‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​‌​⁠​‌​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌⁠​‌​⁠‌⁠‌​​⁠‌‍‌​‌‍‌​‌​‌‍​⁠‌⁠‌‍⁠‌‌‍‍⁠​⁠‌‌‌‍‍⁠‌⁠​‌‌‍‍⁠‌​⁠‌‌⁠‌‌‌‌‌‌​‍​‍‌⁠⁠‌