Software Engineer, Tinker Platform
Thinking Machines Lab
- Location
- San Francisco
- Employment
- Full Time
- Work model
- Remote
- Level
- Senior
- Salary
- $350k – $475k/yr
- Posted
- 2h ago
Skills
About this role
The mission of Thinking Machines is to build AI that extends human will and judgment.
About the Role
We're hiring a senior software engineer to help build and scale the distributed systems that power Tinker, our post-training platform. You'll work on the infrastructure that schedules training and sampling jobs across large GPU clusters, keeps state consistent under failure, and gives every Tinkerer a fast, reliable path from idea to a customized model. This role sits at the center of Tinker's reliability and scale. You'll design and operate the systems that decide how work is placed, checkpointed, recovered, and observed across thousands of accelerators running concurrently — with real users depending on the platform staying up. You'll partner closely with research and infrastructure teams to translate emerging training and sampling workloads into systems that are fast, correct, and resilient by default. This is a hands-on engineering role with significant ownership. You'll be trusted to make architectural calls, drive incidents to resolution, and set technical direction for parts of the platform with minimal oversight. About Tinker Tinker is our fine-tuning API that empowers researchers and developers to customize frontier AI to their needs — opening access to capabilities that have previously been concentrated in a handful of labs. We manage the infrastructure while allowing Tinkerers full flexibility in training open weights models with their own data, algorithms, and for their own needs. Tinker is rapidly adding new customers, features, and novel use-cases, and the systems underneath it run as a live, always-on production platform.
What You'll Do
Design, build, and operate the distributed systems underlying Tinker, including job scheduling, resource orchestration, checkpointing, and fault recovery across large multi-node GPU clusters. Own the reliability, latency, and throughput of live production systems serving concurrent training and sampling workloads for external users. Build for graceful degradation and fast recovery: handle node failures, preemptions, and network partitions without losing user progress or data. Improve observability across the platform — metrics, tracing, and alerting that make it possible to detect and diagnose issues before or as they affect users. Lead incident response for the systems you own, drive root-cause analysis, and turn findings into lasting fixes. Partner with research and ML infrastructure teams to understand new training and sampling workloads and to evolve the platform's architecture to support them at scale. Make build-vs-buy and architectural tradeoff calls for core infrastructure components, and mentor other engineers on distributed systems practices.
Skills & Qualifications
Minimum Qualifications 7+ years of experience building, running, and scaling distributed systems in production, with direct on-call ownership of live services. Deep understanding of distributed systems fundamentals: consensus, consistency models, failure detection, scheduling, and state management under partial failure. Track record of designing systems that operate reliably at scale, including diagnosing and resolving production incidents under time pressure. Strong systems-level coding ability in a language such as Python, Go, Rust, or C++, and comfort working close to the infrastructure layer.
Preferred Qualifications
Experience operating large-scale GPU or accelerator clusters, including job orchestration, resource scheduling, or ML training infrastructure. Familiarity with the training or inference stack for large language models, and an understanding of the systems demands unique to fine-tuning or post-training workloads. Experience with checkpointing, distributed storage, or high-throughput networking (e.g., RDMA, NCCL) in performance-critical settings. Experience building developer-facing infrastructure or APIs where reliability and usability are both first-class concerns. A track record of