Senior AI Inference Engineer - Model Optimization & Deployment at Zoox
- Company: Zoox
- Location: Foster City, CA
- Employment type: Full-time
- Salary: USD 225,000 – 305,000
- Posted: 2026-04-11
- Technologies: CUDA
About the role
The Perception team is pioneering the development of a multi-modality foundation model to drive the next generation of autonomous system intelligence. As a Model Optimization & Deployment Engineer, you will focus on bringing highly efficient, production-ready large-scale models to our on-vehicle stack. We are looking for experts with hands-on experience in compressing, accelerating, and deploying complex models (LLMs, VLMs, or FMs) for power- and thermal-constrained vehicle SOCs. You will optimize the ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution on edge devices.
Apply on Zoox's official careers page: https://jobs.lever.co/zoox/c88c8b02-71b6-492c-a666-584458ac8c6e