Principal Cloud & AI Workload Performance Analysis Engineer
AMD
- Location
- Texas, United States
- Employment
- Full Time
- Work model
- On-Site
- Level
- Principal
- Posted
- 1h ago
Skills
About this role
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.
THE ROLE
AMD is looking for a systems-minded performance engineer to own the workloads most exposed to hyperscaler custom CPUs, Arm ecosystem momentum and AI-system architecture. You will measure and explain cloud-native performance, price-performance, performance-per-watt, software maturity and host-CPU effects on accelerator utilization. The work requires fair methods for comparing AMD EPYC with Intel, Arm Neoverse-based platforms, cloud-custom CPUs, NVIDIA Grace-class systems and other emerging designs. THE PERSON: You are curious at every layer of the system. You can move from containers, orchestration and application throughput down to NUMA placement, memory bandwidth, interconnect behavior and CPU-to-accelerator data movement - then explain the business consequence in a concise, evidence-based way. You care as much about deployability and software maturity as about a benchmark score.
KEY RESPONSIBILITIES
Define and maintain the cloud-native and AI workload taxonomy used in competitive analysis and forecasting. Design reproducible methods for web services, microservices, containers, scale-out data processing, caching, cloud infrastructure and CPU-support functions in AI systems. Compare AMD, Intel, Arm and cloud-custom environments using aligned software versions, tuning policies, instance shapes and service-level objectives. Analyze host-CPU impact on accelerator utilization, input pipelines, communication overheads, memory movement, NUMA behavior and end-to-end AI system throughput. Measure and model performance-per-watt, density, utilization and price-performance where the inputs can be normalized defensibly. Track compiler, kernel, library, orchestration and migration factors that affect Arm and cross-ISA deployability. Work with architecture experts to explain memory, interconnect and platform bottlenecks behind observed results. Translate current scaling behavior and software trends into 24-36 month forecast inputs and early risk or advantage assessments. Develop cross-ISA methodology documentation that can withstand partner, customer and internal technical scrutiny. Support validation with approved cloud, OEM, ISV and ecosystem partners and produce decision-ready technical and executive reports.
KEY RESPONSIBILITIES
Deep hands-on experience benchmarking and profiling cloud, distributed or accelerated-system workloads. Strong Linux systems skills and proficiency with automation or scripting for deployment and analysis. Experience with containers, orchestration and cloud-native workloads at meaningful scale. Ability to design fair experiments across dissimilar CPU architectures and platform configurations. Experience with application profiling, hardware performance counters and system telemetry. Understanding of AI host-side bottlenecks, CPU-to-accelerator data movement, NUMA, memory bandwidth and interconnect effects. Strong technical writing and presentation skills, plus the ability to collaborate with external partners and reproduce results outside one lab. PREFERRED RESPONSIBILITIES: Hands-on experience with Arm Neoverse or cloud-custom Arm platforms. Experience benchmarking major cloud-service-provider instances and services. Experience measuring AI system throughput or accelerator utilization as a function of host-CPU and platform behavior. Familiarity with LPDDR or HBM-attached CPU systems,