Site Reliability Engineer, Post Training at Thinking Machines Lab
- Company: Thinking Machines Lab
- Location: San Francisco
- Employment type: Full-time
- Posted: 2026-08-31
About the role
ABOUT THINKING MACHINES The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it. ABOUT THE ROLE We're hiring a Site Reliability Engineer (SRE) to keep our post-training and reinforcement learning (RL) systems fast, reliable, and easy for researchers to iterate on. Think of this as a production engineering or site reliability role built around model training: you'll own the health of the training runs, clusters, and pipelines that power post-training and RL at Thinking Machines.…
Apply on Thinking Machines Lab's official careers page: https://jobs.ashbyhq.com/thinkingmachines/a9469410-04c7-4e6a-b8b4-64c15933a2bf