Staff Site Reliability Engineer
Palo Alto Networks
- Location
- Office - India - Bangalore Bagmane Tech Park
- Work model
- On-Site
- Level
- Staff
- H-1B history
- 168 approvals (FY2023)
- Posted
- Sep 17, 2026
Skills
About this role
Our Mission
At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.
Who We Are
In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us! We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes.
Job Summary
Palo Alto Networks is looking for a cloud infrastructure and observability professional to help build and operate reliable, scalable, and secure technology platforms. You will combine expertise in cloud operations, automation, and observability to improve infrastructure performance and reliability, while exploring opportunities to apply AI and machine learning to IT operations. Your Career Join a team of senior engineers operating in a large-scale, multi-cloud production environment supporting tens of thousands of enterprise customers worldwide. This is not a typical SRE role — you'll work at the intersection of Site Reliability Engineering and AI-driven automation, pioneering the next generation of intelligent infrastructure operations. As a Staff AI-SRE, you will be managing cloud infrastructure comprising Kubernetes (GKE) and API Gateway (Kong), while leading the adoption of AIOps, Large Language Models (LLMs), and AI-assisted workflows to transform how we build, monitor, and operate systems. You'll work alongside experienced DevOps professionals in a fast-paced, cybersecurity-focused organization committed to AI-first operations. Own and operate large-scale, global production environments with an AI focus — leveraging AIOps and machine learning to drive autonomous operations Pioneer AI-driven automation: Implement LLM-powered runbooks, AI-assisted incident diagnostics, and intelligent alerting systems that predict issues before they impact customers Design self-healing infrastructure: Build systems that leverage AI for automated incident remediation, anomaly detection, and proactive service monitoring Lead architecture, deployment, and operations of Kong API Gateway infrastructure — including rate limiting, authentication, traffic management, and plugin customization Design, deploy, and manage production-grade GKE clusters — cluster upgrades, node pool management, workload optimization, and multi-tenancy Implement predictive scalability: Use AI modeling to forecast resource needs, prevent bottlenecks, and optimize capacity planning across GKE and cloud infrastructure Actively monitor, investigate, and resolve P1/P2 incidents using AI-assisted diagnostics and automated playbooks Drive end-to-end troubleshooting across complex, distributed systems — augmented by AI tools for faster root cause analysis Implement and maintain Infrastructure as Code using Terraform and Helm — integrating with AI APIs for intent-based infrastructure management Champion "vibe coding" workflows: Leverage AI coding assistants (GitHub Copilot, Claude) to accelerate development — focusing on high-level architecture and intent while AI handles boilerplate Develop and maintain automation and tooling (Python, Bash, Go)