Senior Storage Platform Engineer
NVIDIA
- Location
- US, CA, Santa Clara
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 394 approvals (FY2023)
- Posted
- Sep 10, 2026
Skills
About this role
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables unique creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence. Join our team at NVIDIA as a Senior Storage Platform Engineer responsible for designing, deploying, and operating the storage platforms that power NVIDIA's EDA FARM, software engineering, AI/ML teams, and engineering workflows at scale. You will own the full lifecycle of our multi-vendor storage infrastructure while simultaneously building the automation, pipelines, and integrations that turn storage into a scalable, self-service platform.You will drive Infrastructure as Code adoption across various Storage platforms, and ensure our storage estate is tightly integrated with CMDB, observability, configuration management, and self-service tooling.
What You'll Be Doing
Lead end-to-end deployment of storage systems across NetApp, Pure Storage, Cloudian, and DDN, owning timelines, configuration quality, and delivery against project milestones. Define and enforce configuration standards, baselines, and operational runbooks across all platforms; ensure every deployment is consistent, documented, and auditable. Design and implement Infrastructure as Code frameworks (Ansible, Terraform, or equivalent) to automate provisioning, configuration management, and lifecycle operations across all storage platforms. Build and maintain CI/CD pipelines for storage configuration deployments — ensuring every change is reviewed, tested, and rolled out in a repeatable, validated manner. Develop and own integrations between storage platforms and the broader infrastructure ecosystem: CMDB for asset discovery and inventory sync, observability stacks for metrics and alerting, configuration management tools for drift detection, and self-service portals that let engineering teams provision storage on demand. Manage day-2 operations including capacity planning, firmware and software lifecycle management, performance tuning, and root cause analysis. Partner with stakeholders like chip design teams, software engineering, research and application teams to deliver integrated, fit-for-purpose storage solutions for engineering workloads Champion GitOps practices for storage — every configuration tracked, every change reviewed, no manual snowflakes. What we need see: 8+ years of hands-on experience with enterprise storage systems in large-scale production environments. Deep working knowledge of at least three platforms from our stack: NetApp ONTAP (NFS/NAS), Pure Storage FlashArray or FlashBlade, Cloudian HyperStore (S3 object), DDN (Lustre/EXAScaler or high-performance NAS). Strong command of storage protocols — NFS, SMB, iSCSI, NVMe-oF, S3, and Lustre. Proven experience building infrastructure automation with Ansible, Terraform, or equivalent IaC tools. Proficiency in Python, Go, or similar, with a track record of building reusable tooling rather than one-off scripts. Experience integrating storage systems with CMDB platforms (ServiceNow, Nautobot, Cerebro or equivalent) and observability stacks (Prometheus/Grafana, Splunk, Datadog, or equivalent). Solid experience with CI/CD platforms (GitHub Actions, Jenkins, GitLab CI, or equivalent) and Git-based workflows. Ability to work across teams and translate infrastructure needs into platform capabilities. MS Degree in Computer Science or equivalent experience Way to stand out from the crowd: You've worked in HPC, AI/ML, or large-scale research computing environments where storage is mission-critical and throughput matters. You've built self-service storage