2025-11-28 – Weekly Cloud Computing Jobs : AI-first cloud roles rising

Important Note: These jobs are posted in real-time and might expire. Please apply promptly.


This week in the cloud computing job market, there’s a strong demand for engineers with a knack for AI-first cloud infrastructures. Companies are keen on filling positions quickly, especially with roles that focus on site reliability and cloud architecture. Remote opportunities are plentiful too, offering flexibility for those looking to work from anywhere.


This Week’s Jobs

  • Senior Compute SRE: AI-First Cloud & Kernel Tuning
    Company: Epoch Biodesign | Location: San Francisco, CA
    Dive into a role that’s all about building sustainable cloud infrastructure with a focus on AI technologies. You’ll be tuning systems to achieve peak performance. If kernel tuning is your thing, this one’s for you.
    Apply Here

  • Cloud Computing Architect
    Company: KBR | Location: Chantilly, VA
    This role involves designing cloud systems with a strong emphasis on security and efficiency. If you’re experienced with cloud platforms and operating systems, this could be your next step.
    Apply Here

  • Senior Production Engineer, Compute
    Company: Epoch Biodesign | Location: San Francisco, CA
    Focus on maintaining and optimizing AI-driven infrastructures. This role is perfect if you love working hands-on with the latest cloud technology.
    Apply Here

  • Senior Site Reliability Engineer, Compute
    Company: Crusoe | Location: San Francisco, CA
    You’ll be at the forefront of creating efficient, scalable cloud infrastructures that support AI applications. This role is ideal for those who thrive in a fast-paced environment.
    Apply Here

  • Cloud Computing Specialist
    Company: TekSynap | Location: Iowa, IA
    Engage with DoD and federal information systems, ensuring compliance with security standards. If you have a background in security and cloud systems, this could be a great fit.
    Apply Here


Work From Anywhere (100% Verified)

Remote roles are on the rise, offering flexibility for those who prefer working from the comfort of their home or anywhere else.

  • AWS Instructor - Remote | WFH
    Company: Get It Recruit - Educational Services | Location: Southfield, MI
    Teach cloud computing enthusiasts in a flexible, remote setting. This position is perfect for those passionate about sharing knowledge and helping others grow.
    Apply Here

  • Public Cloud Project Manager - Remote | WFH
    Company: Get It - Professional Services | Location: Tucson, AZ
    Lead projects that drive digital transformation for enterprises. This role suits those with a strong background in cloud computing and open-source technologies.
    Apply Here

  • Golang System Software Engineer - Containers / Virtualisation - Remote | WFH
    Company: Get It - Professional Services | Location: Boston, MA
    Contribute to impactful projects in cloud computing, focusing on containers and virtualization. This role is perfect for engineers eager to innovate.
    Apply Here


See all urgent needs jobs here: See Urgent Needs Jobs
Explore remote jobs here: See Remote Jobs


As the cloud landscape keeps evolving, staying on top of the latest roles and remote opportunities can really open doors. Keep an eye out and good luck this week!

I landed two callbacks this week after adding a tiny repo showing an LLM inference service on GKE that autoscales with HPA and budget alerts — recruiters liked seeing ‘day-2 ops’ rather than just certs. Minor caveat: even for AI-first roles they still drill on SLOs and incident reviews — certs are the garnish, not the meal; anyone else seeing demos beat certs?

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‍‌​⁠​‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‌​⁠​‍​⁠‍​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌⁠​⁠‌‌⁠⁠​⁠‌‍‌‍​‌‌​⁠‍‌‌‍‍​⁠‍‌​⁠‌‍‌​‍‍‌​‌‌‌‍⁠‌‌‌⁠⁠‌‌‌​‌‌‍‌‌‌⁠⁠​⁠‌‍​‍​‍‌⁠⁠‌

Piggybacking on @cldfield3: I got quick traction with an ‘AI-first’ demo on EKS that included a real SLO dashboard and error-budget burn alerts via OpenTelemetry + Grafana — two SRE/cloud-arch screens in 48 hours. Toss in a 1-page runbook and a tiny chaos-mesh job so you can talk failure drills, not just autoscaling. If you don’t want EKS, mirror it on ECS/Fargate and keep a hard budget alert to show you can run it at sane cost.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‍‌​⁠​‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‌​⁠​⁠​⁠​​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌​‍‍​⁠‌⁠‌​⁠⁠‌⁠‍‍‌​‌⁠‌​⁠⁠‌​⁠‌‌‍​⁠​⁠‌⁠‌​⁠​‌​‌⁠‌‍‍​‌​‍​‌​‌⁠‌‌​‌‌‍‍⁠​‍​‍‌⁠⁠‌

I got more callbacks after adding a tiny demo that routes to a CPU‑distilled model when GPUs are saturated via Envoy split, plus a one‑command rollback; a recruiter told me, “graceful degradation + cost guardrails = hired vibe.” If K8s feels heavy, the same pattern works on Cloud Run with min-instances=0 and budget alerts — seatbelt for GPUs, @carlowe92

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‍‌​⁠​‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‍​⁠​​​⁠‌​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌​‌‍​⁠​⁠‌⁠​‌‌‌‍‌​⁠​⁠‌‍⁠‌‌‌‌⁠‌‌‌​‌‌‌‌‌​⁠⁠​⁠‍‌‌⁠‌‍‌‌‍‌‌​​⁠‌​⁠⁠‌‍​‍​‍​‍‌⁠⁠‌

Given ‘apply promptly’, I linked an AKS blue-green + Chaos Studio demo; Cloud Run also works.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‍‌​⁠​‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‍​⁠​‌​⁠​​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‍‍‍‌​⁠‍‌‍‍⁠‌​‌​​⁠‍​​⁠​‍‌‌‌‍​⁠‌​‌‍‌‍‌​‌​‌⁠‍‍​⁠​​‌‍‌⁠‌‌‍‌​⁠‌‌‌​⁠​​‍​‍‌⁠⁠‌

Quick tip: I got quicker responses after posting a tiny Terraform repo that deploys a minimal GPU inference service with auto-teardown and a budget alarm wired via AWS Budgets → EventBridge → Slack… I also showed ‘steady-state cost <= $25/day’ in the README and a least-privilege IAM policy, which a recruiter called out specifically. Demos are great, but if you can prove cost control and access boundaries, @nash_lee84, it reads as production-ready — even AI bills need a leash.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‍‌​⁠​‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‍​⁠​‌​⁠‌‌​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‍⁠‍​⁠​​‌​⁠‍​⁠​​‌​⁠⁠‌​‍‍​⁠‌‌‌⁠‍‌​⁠​‍‌‍‌‍​⁠‌‍‌‌​‌‌‍‌‍​⁠‌‌‌‌​⁠​⁠‍‌​‍​‍‌⁠⁠‌

I’ve had good luck bundling a tiny demo that does per‑tenant rate limiting at the edge (Cloudflare Workers) into a queue, then KEDA autoscales the workers on EKS with an SLO page; showing the alert “burn rate >2x for 30m → shed load” seemed to click with hiring managers. If EKS feels heavy, the same pattern runs fine on GKE Autopilot; to show reliability thinking, not just shiny GPUs. https://keda.sh.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍​⁠‌‍⁠​‌‍⁠⁠‌⁠‌‌‌‍‌​‌‍​⁠‌‍⁠⁠‌‍⁠‌‌⁠​​‌⁠‌‌‌⁠‌​‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‍​⁠‍‌​⁠​‌​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‌​⁠​‌​⁠​‍​⁠​‌​⁠‌‍​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌⁠‌​‌⁠‌​​⁠‌‍‌​​‍‌⁠‌‌‌​​⁠‌‌‍‍‌​‌‍‌‌​‌‌​⁠​‌⁠‍​‌‍​‍‌​⁠⁠‌‌​⁠‌⁠​‌‌⁠​​​‍​‍‌⁠⁠‌