Manager, Infrastructure Engineering and DevOps
NVIDIA (Eightfold)
- Location
- Israel, Yokneam
- Work model
- On-Site
- Level
- Mid
- Posted
- Aug 17, 2026
Skills
About this role
NVIDIA is looking for an outstanding Manager, Infrastructure Engineering - Networking to lead a high-impact infrastructure engineering team. The team develops and maintains engineering infrastructure solutions that enable R&D teams to provision, validate, test, debug, and recover complex server and networking environments at scale. In this role, you will lead a team that provides infrastructure platforms and hands-on engineering support for internal customers across firmware, driver, hardware, software, and verification organizations. The position combines deep technical leadership with people management, execution ownership, customer support, and operational excellence in a fast-paced R&D environment. We are looking for a strong technical manager who can grow and mentor engineers, set technical direction, drive recovery of complex systems, optimize customer flows, and deliver reliable infrastructure capabilities at NVIDIA scale. What you’ll be doing : Lead and grow an infrastructure engineering team responsible for bare-metal provisioning, VM infrastructure, server fleet automation, CI/CD infrastructure, customer-facing debug support, and high-performance networking environments. Own the team’s technical roadmap, priorities, execution plans, and delivery commitments across multiple infrastructure initiatives, while balancing long-term platform improvements with day-to-day customer needs. Drive ownership of VM box inventory and lifecycle management across many Linux distributions, including image readiness, OS compatibility, package baselines, kernel configurations, provisioning flows, and production availability. Build infrastructure capabilities that enable engineering and verification teams to run provisioning, testing, validation, and debug workflows efficiently and reliably, without positioning the team as the owner of verification itself. Lead customer support, debug, and optimization of internal customer flows, including triage, root-cause analysis, bottleneck removal, workflow improvements, and clear communication with engineering stakeholders. Guide complex system debug and recovery in a firmware R&D environment, including server bring-up issues, driver and firmware interactions, boot failures, networking problems, lab instability, automation failures, and environment recovery. Provide technical leadership for Linux-based automation platforms, including server lifecycle management, OS installation, kernel configuration, driver setup, inventory management, resource allocation, observability, and production readiness. Partner closely with firmware, driver, hardware, software, cloud, and verification teams to define requirements, improve reliability, and deliver infrastructure solutions that accelerate engineering productivity. What we need to see: B.Sc. in Computer Engineering, Computer Science, Electrical Engineering, or a related technical field, or equivalent experience. 8+ overall years of experience in Linux systems administration, infrastructure automation, DevOps, system software, firmware infrastructure, lab infrastructure, or related engineering domains. 3+ years of experience leading or managing engineering teams, technical projects, or cross-functional infrastructure initiatives. Strong technical background in Linux environments, including systemd, package management, kernel parameters, GRUB, sysctl tuning, NFS, networking, boot flows, and service management. Hands-on experience designing, implementing, and debugging automation software using Python, scripting, CI/CD workflows, and modern software development practices. Experience managing infrastructure across multiple Linux distributions, including OS image management, compatibility issues, provisioning flows, package dependencies, and environment consistency. Proven ability to support internal customers in complex technical environments, including issue triage, root-cause analysis, flow optimization, incident handling, and communication with