[Junior / Senior / Staff] Software Engineer, Inference / Compute Infrastructure Engineering at Together AI
- Company: Together AI
- Location: San Francisco
- Employment type: Full-time
- Posted: 2026-07-16
- Technologies: Kubernetes
About the role
About the Role . We're looking for a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You'll design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving stack, scheduler, or hardware pool is doing the work. You'll also build the systems that keep the fleet efficient, not just running, including defragmentation and rebalancing logic that consolidates scattered workloads back into contiguous capacity, and scheduling/bin-packing improvements that push GPU utilization up without hurting latency.…
Apply on Together AI's official careers page: https://job-boards.greenhouse.io/togetherai/jobs/5186628007