Senior Technical Marketing Engineer - DSX AI Infrastructure Software
NVIDIA
- Location
- US CA Santa Clara
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 394 approvals (FY2023)
- Posted
- Aug 26, 2026
Skills
About this role
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. NVIDIA DSX brings together facilities infrastructure, hardware, software, simulation, and partner technologies to build and run efficient AI factories. We are looking for a Senior Technical Marketing Engineer to show and educate our AI factory ecosystem how to bring up and operate the entire stack, ranging from facilities and multi-node GPU infrastructure to provisioning, networking, storage, cluster orchestration, security, observability, and workload enablement.
What you'll be doing
Stand up and validate complete DSX-aligned software stacks on multi-node GPU systems. Capture the dependencies, configuration order, validation steps, and operational handoffs as you go. Turn working deployments into useful technical content: reference architectures, quick-starts, installation and upgrade guides, troubleshooting runbooks, code examples, blogs, whitepapers, and demo videos. Build reusable examples and automation with APIs, Python or shell scripting, infrastructure-as-code, containers, Kubernetes, Slurm, Helm, GitOps or equivalent experience, and CI/CD where they fit. Build demos, labs, and training that address the practical aspects of operating an AI factory, from initial deployment and tenant setup to upgrades, monitoring, scheduling, fault isolation, remediation, capacity management, and security. Show how the layers of the stack fit together. Work with TME, Product, Engineering, and Marketing to demonstrate how data center hardware, infrastructure and cluster management software, orchestration, AI platforms, and the workloads on top operate as one system. Test pre-release software using representative training and inference workloads. Identify rough edges, assess interoperability and resiliency, and provide Product and Engineering with clear feedback before customers face similar issues. Help solution architects, field teams, cloud and OEM partners, ISVs, and system integrators use the stack successfully through repeatable assets, train-the-trainer sessions, live demos, and direct support on important engagements. Collaborate with open-source and cloud-native communities to demonstrate practical integration approaches, address documentation and usability shortcomings, and assist partners in expanding and developing the DSX software stack. Listen for recurring problems from customers, partners, the field, and developers. Use those signals to set content priorities and recommend product improvements, then track whether the work reduces deployment time and improves operational success. Present your work in customer briefings, partner workshops, industry events, webinars, and internal training. Some travel will be required. What we need to see: BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or another technical field, or equivalent experience. 8+ years of experience in infrastructure engineering, systems engineering, solutions architecture, software engineering, technical marketing engineering, site reliability engineering, or a related role. Hands-on experience deploying and operating Linux-based data center, cloud, HPC, or AI infrastructure, including multi-node GPU systems and production operational practices. Strong working knowledge of Kubernetes and/or Slurm, including containers, operators,