Senior HPC Engineer - Fleet Engineering at Lambda
- Company: Lambda
- Location: Remote / San Francisco Office (Fremont St)
- Employment type: Full-time
- Salary: $227K – $356K • Multiple Ranges
- Posted: 2026-09-01
- Technologies: AWS, Terraform, Ansible
About the role
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU. If you'd like to build the world's best AI cloud, join us. What You’ll Do - Build and operate monitoring and alerting for cluster health — fabric, GPU, power/thermal, and job-level signals — to detect and respond to issues proactively - Remotely deploy and configure large-scale HPC clusters for AI workloads using automation wherever possible - Automate cluster lifecycle: operating systems, firmware, drivers, and networking, managed as code (Ansible, Terraform) rather than by hand - Create runbooks and automated remediations for common cluster failure modes, designed so…
Apply on Lambda's official careers page: https://jobs.ashbyhq.com/lambda/d4e6f207-73c3-43d1-aa4b-d840ad2254d4