yoinka

Principal Research Engineer

Microsoft

United States, Washington, RedmondPrincipalH-1B sponsor company
Sign in to applyVerified 1h ago
Location
United States, Washington, Redmond
Work model
On-Site
Level
Principal
H-1B history
2,066 approvals (FY2023)
Posted
2h ago

Skills

JavaMachine LearningPyTorchPython

About this role

Overview

Microsoft Research Americas is seeking a Principal Research Engineer to help advance how foundation models are trained, adapted, evaluated, and improved. You will work with researchers and engineers to turn promising model ideas into reproducible experiments and measurable improvements that can support multiple research initiatives. In this role, you will lead hands-on work across model pretraining, continued pretraining, fine-tuning, and post-training. You will deepen your experience in foundation model development, model evaluation, and large-scale experimentation while partnering with machine learning systems and infrastructure specialists to scale successful approaches across GPU clusters. This role is based in Redmond, Washington, with an expectation of three days per week in the office. Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Responsibilities

Lead model-training programs spanning pretraining from scratch when appropriate, continued pretraining of existing base models, supervised fine-tuning, and post-training. Improve model quality through changes to model architecture, training data, learning objectives, optimizers, hyperparameters, and end-to-end training recipes. Design and implement controlled experiments, ablation studies, and evaluation methods for accuracy, reasoning, robustness, generalization, safety, and domain performance. Diagnose and resolve training instability, convergence failures, numerical issues, data-quality defects, overfitting, catastrophic forgetting, and capability regressions. Own technical and design decisions and develop reusable training, evaluation, experimentation, and model-release practices that support multiple research initiatives. Partner with machine learning systems and infrastructure specialists to scale successful approaches across GPU clusters, improve training efficiency, and ensure reliable and reproducible execution. Provide technical direction across research and engineering teams, mentor engineers, communicate trade-offs, and translate ambiguous goals into measurable milestones.

Qualifications

Required Qualifications: Bachelor's Degree in Computer Science, Machine Learning, Applied Mathematics or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, Python, C, C++, C#, or Java OR equivalent experience.

Preferred Qualifications

Master's Degree or Doctorate in Computer Science, Machine Learning, Applied Mathematics, or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, Python, C, C++, C#, or Java OR Bachelor's Degree in a related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, Python, C, C++, C#, or Java OR equivalent experience. Experience training language or foundation models, including developing models through pretraining or continuing the training of existing base models. Experience improving model quality through changes to training data, model architecture, learning objectives, optimization methods, hyperparameters, or post-training techniques. Experience developing and debugging machine learning systems using Python and a modern machine learning framework such as PyTorch or JAX. Experience designing model evaluations or ablation studies used to assess training changes and guide technical decisions. Experience with one or more model-adaptation methods, such as continued pretraining, instruction tuning, parameter-efficient fine-tuning, preference optimization, reinforcement-learning-based post-training, or

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Principal Research Engineer at Microsoft, United States, Washington, Redmond | Yoinka