yoinka

Lead Machine Learning Engineering, (Hybrid)

Cisco

RemoteSeattle Washington, USSenior
Sign in to applyVerified 2h ago
Location
Seattle Washington, US
Work model
Hybrid
Level
Senior
Posted
Aug 15, 2026

Skills

GenAILLMMachine LearningPyTorchPythonTensorFlow

About this role

The application window is expected to close on: 09/28/2026 Job posting may be removed earlier if the position is filled or if a sufficient number of applications are received . This is a hybrid role based out of Cisco's Seattle or San Jose office. Meet the Team The Cisco AI Research team brings together AI researchers, machine learning engineers, data engineers, and networking domain experts to build the next generation of AI-powered networking. We work at the intersection of generative AI, large-scale data systems, and networking, developing Large Language Models (LLMs), agents, and domain-specific AI systems. Our work spans research and engineering, with a strong focus on translating advances in AI into scalable systems and real-world impact.

Your Impact

As a Lead Machine Learning Engineer, you will build and improve the data and ML systems that power our LLMs and AI models . A major focus of this role is solving one of the most important challenges in modern AI: creating high-quality training and evaluation data at scale . You will design and build scalable data pipelines, improve human data labeling workflows, create synthetic datasets, and develop automated approaches for continuously measuring and improving dataset quality. This is a hands-on technical role at the intersection of m achine learning engineering and data engineering . You will work closely with researchers, engineers, and domain experts to determine what data our models need, how to create it efficiently, and how to measure its impact on model performance. Design, build, and maintain robust, scalable data pipelines that support the full lifecycle of ML and LLM development, from initial data ingestion to production-ready model deployment. Architect and manage human-in-the-loop labeling workflows, including task generation, quality control, and feedback integration to ensure high-fidelity training data. Develop scalable strategies for synthetic data generation, filtering, and validation to enhance dataset diversity, coverage, and overall quality. Leverage LLMs and advanced ML techniques to automate data generation, labeling, scoring, and evaluation processes, increasing efficiency and consistency. Establish rigorous systems to measure and mitigate dataset failure modes—such as bias, contamination, and distribution shifts—while designing experiments that directly link dataset composition to model performance. Collaborate closely with researchers and ML engineers to define dataset requirements for fine-tuning, preference learning, and agent development, ensuring alignment with project goals. Provide technical direction on infrastructure, compute, and storage decisions while fostering engineering excellence through design reviews, best practices, and team mentorship.

Minimum Qualifications

Bachelor’s degree in a STEM field with 8+ years of relevant experience, OR Master’s degree in a STEM field with 6+ years of relevant experience, OR PhD in STEM or a relevant technical field with 3+ years of industry or academic research experience. 3+ years of hands-on experience building, curating, and scaling datasets for machine learning training and evaluation. 5+ years of professional programming experience using Python, C++, or Go within a production or research environment. 5+ years of experience using machine learning frameworks such as PyTorch, TensorFlow, or equivalent technologies to develop, train, evaluate, and deploy machine learning models.

Preferred Qualifications

Expertise in curating, scaling, and managing datasets for the entire LLM lifecycle—including synthetic data generation, augmentation, and post-training workflows like SFT and RLHF. Proficiency in designing human-in-the-loop labeling systems and proactively mitigating complex dataset failure modes such as label noise, bias, contamination, and distribution shift. Demonstrated success using LLMs for data generation, model-assisted labeling, and evaluation, with

Listing verified 2h ago. Applications go through the company's official careers site.

← Back to Yoinka

Lead Machine Learning Engineering, (Hybrid) at Cisco, Seattle Washington, US | Yoinka