Inference Engineering, Co-op at Inferact
- Company: Inferact
- Location: San Francisco
- Employment type: Internship
- Posted: 2026-09-15
About the role
Overview Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build. About the Role We're looking for exceptional University of Waterloo co-op students who want to work on the systems that determine how fast, efficiently, and reliably frontier AI models run on frontier workloads at scale. This is not a sandboxed internship project. We will match your strengths and interests to a real engineering problem across the vLLM stack, from model execution and low-level accelerator code to distributed serving and the cloud platform that makes it all usable.…
Apply on Inferact's official careers page: https://jobs.ashbyhq.com/inferact/2b6032f9-12a5-4083-b5e6-4bec5376cbaa