Senior Escalation Engineer, Hypershield (remote)
Cisco
- Location
- San Jose California, US
- Work model
- Remote
- Level
- Senior
- Posted
- Sep 17, 2026
Skills
About this role
The application window is expected to close on: 10/04/2026 Meet the Team Isovalent , now part of Cisco, was founded by the creators of Cilium and eBPF , and builds open-source software and enterprise solutions for the networking, security, and observability needs of modern infrastructure. The Customer Reliability Engineering team is the deep technical escalation tier for Cisco Hypershield on the Cisco Nexus N9300 Series Smart Switches. The team owns the hardest break/fix and reliability cases raised by Cisco TAC, applying Site Reliability Engineering practices across the full stack: the data-center fabric and the on-premises Kubernetes controller that manages the security policy enforced on it. The work demands methodical diagnosis, composure under incident pressure, and the ability to operate at the seam between customer environments and engineering.
Your Impact
The ideal candidate combines deep networking expertise with strong troubleshooting skills, customer-facing experience, and a passion for improving reliability across complex product environments. Own Hypershield cases raised from Cisco TAC through to resolution, engaging customers directly as the incident requires Diagnose complex production failures through the Hypershield surface: the N9300 Smart Switch fabric and the on-premises Kubernetes controller that manages its security policy Localize faults across the layered architecture: switching and forwarding , security services and enforcement, and the control plane Develop a deep understanding of each customer's architecture and configuration, and diagnose failures in unfamiliar production environments Reproduce customer failures, partner with engineering to drive fixes, and own the fix back to the customer Convert individual cases into systemic improvements: runbooks, diagnostics, knowledge-base content, and product feedback to engineering Help build the team's comprehensive view of customer health, developing new monitoring, tooling, and reliability practices as the installed base grows Minimum Qualifications Bachelor's + 8 years of experience, Master's + 6 years, or equivalent industry experience Experience supporting enterprise customers in a commercial capacity Experience operating and trou bleshooti ng Cisco Nexus / NX-OS in production; equivalent depth on another major vendor accepted Prior experience to localize failures across a layered data-center architecture spanning switching/forwarding, services/enforcement, or control-plane domains Linux operations or operations experience including working exposure to containers, Kubernetes, or Virtual Machines (VMs) Preferred Qualifications Experience supporting enterprise customers by diagnosing and resolving complex production incidents, with a demonstrated ability to clearly communicate status, root cause, and remediation to both technical and executive audiences, verbally and in writing Direct experience using network troubleshooting tooling as a primary diagnostic method, including packet capture and flow-telemetry analysis (NetFlow/IPFIX) Knowledge of enterprise virtualization, sufficient to troubleshoot a VM-based appliance deployment; vSphere is the current deployment target Operational Kubernetes and Helm proficiency , including diagnosing failures beyond the workload level. Working knowledge of VXLAN EVPN fabrics, including the Smart Switch's placement within them and the ability to isolate faults across the fabric Experience with automation and APIs; familiarity with Cisco Nexus Dashboard a plus CCNP Data Center, CCNP Enterprise, DevNet Professional, CCIE Data Center, CCIE Enterprise, or DevNet Expert (a plus) Why Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and