I’m mapping my Q1 study plan and want courses that improve throughput and data integrity, not just theory. If you’ve taken a hands-on track (Databricks Lakehouse, AWS Data Engineer, Airflow/dbt labs, etc.) that helped you ship idempotent, backfill-safe jobs and cut compute spend by about 20% or improved SLA adherence, which one delivered and why?
I got the most ROI from Databricks Optimizing Apache Spark + DLT labs (https://academy.databricks.com/) — we cut about 22% compute with AQE, autoscaling job clusters, and OPTIMIZE/ZORDER on hot Delta tables. DLT expectations plus MERGE-based upserts made runs idempotent and safe to backfill, like tightening the bolts on a race bike. If you’re mostly Airflow/dbt, Astronomer Academy + dbt Advanced Testing is a close second; are you on AWS or Azure?
Quick example: Astronomer Academy’s Airflow DAG Authoring + Data-Aware Scheduling labs (https://academy.astronomer.io) taught us to switch long S3/ExternalTask sensors to “deferrable” and schedule on datasets with catchup, making backfills idempotent and trimming worker hours about 21%.