Sr. Software Development Engineer, Inference Team - AWS Neuron
Amazon
- Location
- US, WA, Seattle
- Employment
- Full Time
- Work model
- On-Site
- Level
- Senior
- Posted
- Aug 17, 2026
Skills
About this role
AWS Neuron is the complete software stack for AWS Inferentia and Trainium, AWS purpose-built accelerators for cloud-scale machine learning. This senior software engineering role is part of the Machine Learning Inference Applications team and focuses on delivering high-performance model inference solutions for customer workloads running on Inferentia- and Trainium-powered instances. The engineer will lead the development of core serving technologies within open-source frameworks (such as vLLM and SGLang) to enable customers to serve large-scale inference efficiently on AWS Neuron. The engineer will also drive excellence in inference framework optimization and engineering best practices. The engineer will work closely with model development engineers, performance engineers, compiler engineers, and runtime engineers to ensure end-to-end model performance and deliver production-ready accuracy, scalability, and efficiency across a broad range of models and customer use cases. Key job responsibilities The engineer will lead the customization and optimization of open-source inference frameworks (vLLM, SGLang) on AWS Neuron—including core framework logic such as scheduling and model execution—applying state-of-the-art techniques in kernel development, parallel computation, distributed KV cache, speculative decoding, and systems engineering to deliver best-in-class LLM serving performance. The engineer will also influence the team's technical roadmap by evaluating emerging inference research and translating it into production-ready features, and raise the bar through design leadership and mentorship. This role requires close collaboration across model development, compiler, runtime, and performance engineering teams.
About the team
Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we’re building an environment that celebrates knowledge-sharing and mentorship. Our senior members enjoy one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects that help our team members develop your engineering expertise so you feel empowered to take on more complex tasks in the future.