Principal Technical Account Manager, ES - NAMER - US-Frontier AI
Amazon
- Location
- US, CA, San Francisco
- Employment
- Full Time
- Work model
- On-Site
- Level
- Principal
- Posted
- Sep 18, 2026
Skills
About this role
Are you ready to transform how businesses leverage artificial intelligence and machine learning at scale? Join the Amazon Web Services (AWS) Support team and become a strategic partner in delivering Amazon AI/ML solutions that empower our Frontier AI customers to innovate, optimize, and achieve unprecedented operational excellence. Amazon Web Services (AWS) is seeking an experienced Sr. TAM with expertise in AI/ML, HPC, and/or other technologies to join our Frontier AI Technical Account Management (TAM) team. You'll be at the forefront of solving complex AI/ML model training and inference implementation challenges, guiding Frontier research customers through their most ambitious machine learning transformation journeys. By combining deep technical expertise with collaborative problem-solving, you'll help organizations unlock the full potential of artificial intelligence and machine learning technologies — from distributed model training on GPU clusters to production-grade inference at scale. The TAM role is not directly hands on keyboard within the customer’s environment for troubleshooting customer support issues, rather you will work with appropriate engineers and service teams to see issues through to resolution. You will help our customers design, build, operate, and secure their cloud environments. More importantly, you will work proactively to help craft and execute strategies to drive our customers' adoption and use of AWS services, including EC2, S3, DDB, RDS, and many more. Your technical acumen and customer-facing skills will enable you to effectively represent AWS within a customer environment, and drive discussions with senior leadership regarding incidents, trade-offs, support and risk management. You will provide advocacy and strategic technical guidance to help plan and build solutions using best practices and proactively keep your customers’ AWS environments operationally healthy and resilient. The close relationships developed with your customers will allow you to understand their business/operational needs and technical challenges to help them achieve the greatest value from AWS. This position will require the ability to travel 10% or more as needed. The TAM is the centerpiece of value to our Enterprise Support customers. If you wish to be at the forefront of innovation, come join us! Key job responsibilities Deliver Strategic Technical Engagements — Lead comprehensive technical deep-dives and performance optimization for enterprise AI/ML workloads, including distributed training cluster architecture using AWS Parallel Computing Service (PCS) and AWS ParallelCluster, the latest GPU-accelerated computing (i.e. P6/P6e , G7/G7e instances), AWS Trainium-based training (Trn3 UltraServers), and multi-node NCCL communication tuning over EFA’s SRD protocol. Architect and Validate Innovative Solutions — Supporting customers who design and implement production-grade AI/ML training and inference solutions leveraging Slurm-based job scheduling, distributed training frameworks (PyTorch FSDP, DDP, DeepSpeed, Megatron-LM), SageMaker HyperPod for managed GPU clusters with automated health checks and node replacement, high-performance parallel storage (Amazon FSx for Lustre), and container runtimes on Deep Learning AMIs (DLAMIs) against reference architectures and HPC lens to ensure performance, reliability, and cost governance at scale. Architect solutions using P6e UltraServers for multi-trillion parameter frontier models and Trn3 with the AWS Neuron SDK for cost-optimized training and inference. Enable Customer Success — Support customers in implementing business-critical HPC capabilities, including the development of large language model (LLM) (Llama, GPT-class models), physics-informed neural networks (PINNs) and surrogate models, MLOps pipelines, simulation-ML hybrid architectures orchestrated by AWS Step Functions and AWS Batch, distributed data processing, cluster observability, and governance controls