yoinka

Research, Mid Training

Thinking Machines Lab

RemoteSan Francisco, CAFull TimeMid$350k – $475k/yr
Sign in to applyVerified 1h ago
Location
San Francisco, CA
Employment
Full Time
Work model
Remote
Level
Mid
Salary
$350k – $475k/yr
Posted
1h ago

Skills

Deep LearningJAXLLMMachine LearningPyTorchPythonTensorFlow

About this role

About Thinking Machines The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

Mid-training is a step between pre-training and post-training, where we take a base model and train it into the foundation for reasoning. This role owns the late-stage training responsibility that shape what our models are fundamentally capable of, including things like synthetic data strategies, the data mix, quality uplift, context extension and capabilities across coding, math, reasoning, and so on. This role blends fundamental research and practical engineering, as we do not distinguish between the two internally. It's an excellent fit for someone comfortable working across the boundary of pre-training and post-training, and who wants to shape what our models can do at their core. What You’ll Do Own the data. Decide what the model needs to see for each capability and each area of knowledge, then source, curate, and synthesize it. Build the pipelines that filter, deduplicate, verify, and rewrite raw material into training-grade datasets. You'll be responsible for producing data that actually moves the model, and for making sure it reflects how people really use these models, not just clean benchmark-style tasks. Improve what the model knows. Design and measure interventions that increase knowledge: targeted corpora, synthetic rephrasings and QA over source documents, knowledge-dense mixes. Characterize how knowledge scales with data and when it is retained through post-training versus forgotten. Instill behaviors and set the prior. Introduce new behaviors during mid-training and measure them the right way: not just whether an eval goes up, but whether post-training becomes easier, more sample-efficient, and more stable as a result. Work closely with the post-training team to decide which behaviors belong in mid-training and which belong in RL. Build the quality pipeline. Own the automatic filtering and scoring stack: train and calibrate quality classifiers and LLM-based judges, build verifiers for synthetic data, and use them to raise the quality bar of the mix by a lot, not a little. Give the team fine-grained control over data attributes (difficulty, domain, format, style, correctness) and measure the effect of each. Develop and tune the recipe. Iterate on mid-training recipes: the collection of datasets, training stages, annealing schedules, and hyperparameters. Measure how recipe choices affect metrics, including downstream of post-training. Iterate on evals. Mid-training involves a never-ending loop of defining a set of evaluations, optimizing them, and then realizing your existing evals don't capture what matters. You'll be responsible for both making numbers go up and making sure the numbers are meaningful. Debug and understand. While tuning the details of a training configuration, we often observe results that don't quite make sense. You'll be responsible for both getting things to work and developing a deeper understanding we can bring to the next problem. Scale and explore. Mid-training will involve a combination of scaling existing methodologies and developing new ones. We'll want to both measure how performance scales with dataset size and explore using completely different kinds of training data.

Skills and Qualifications

Required qualifications: Proficiency in Python and familiarity with deep learning frameworks (e.g., PyTorch, TensorFlow, or JAX). Comfort debugging distributed training and writing code that scales. Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding. Clarity in communication, an ability to

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Research, Mid Training at Thinking Machines Lab, San Francisco, CA | Yoinka