Sr Staff Site Reliability Engineer (Wildfire) - NetSec - Bangalore
Palo Alto Networks
- Location
- Office - India - Bangalore Bagmane Tech Park
- Work model
- On-Site
- Level
- Staff
- H-1B history
- 168 approvals (FY2023)
- Posted
- Aug 12, 2026
Skills
About this role
Our Mission
At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.
Who We Are
In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us! We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes.
Job Summary
The Team Palo Alto Networks’ Cloud-Delivered Security Services (CDSS) is the intelligence engine of our Next-Generation Security platform. We provide a suite of AI-driven, subscription-based services—including Advanced Wildfire, Advanced Threat Prevention, DNS Security, URL Filtering, and IoT Security—integrated natively into our firewalls. Our infrastructure processes trillions of events daily, delivering real-time protection to over 85,000 global enterprises. Working in CDSS means building the backbone of global cybersecurity at a scale few companies in the world ever reach.
Job Summary
We are seeking an ambitious, technically sharp Senior Staff Site Reliability Engineer to drive the reliability, operational scalability, and infrastructure engineering for Palo Alto Networks’ WildFire malware analysis platform. In this high-impact role, you will take technical ownership of system resilience across WildFire services and appliance platforms. You will bridge the gap between threat analysis pipelines, low-level system execution, and high-throughput malware sandboxing infrastructure - ensuring sub-second detection telemetry and high availability for enterprise deployments worldwide.
Key Responsibilities
Hybrid & Cloud Infrastructure Resilience: Architect, scale, and maintain operational reliability across WildFire’s multi-tenant public clouds (AWS/GCP/Azure/OCI), private cloud appliances, and hybrid inspection pipelines. Infrastructure Resilience: Architect, scale, and maintain the overarching operational reliability for WildFire’s cloud analysis engines, virtualized sandboxes, and distributed appliance infrastructure. Autonomous Operations & IaC: Spearhead the transition to fully automated operational workflows using Terraform, Ansible, and GitOps (ArgoCD) to eliminate operational toil across multi-tenant and edge environments. SLO & Error Budget Governance: Define and enforce SLIs, SLOs, and SLAs across malware inspection pipelines. Partner with security engineering leads on Error Budget strategies to balance rapid threat signature deployment with platform stability. Observability & Threat Telemetry: Architect end-to-end observability stacks (Prometheus, Grafana, OpenTelemetry, Datadog/ELK) optimized for low-latency tracing, high-concurrency sample processing monitoring, and rapid MTTR. Cloud & Platform Release Engineering: Build and scale enterprise CI/CD automation (GitHub Actions / GitLab CI) empowering engineering teams to safely deploy cloud microservices, threat analysis engines, and platform firmware updates. On-Call & Incident Response: Lead production on-call rotations, establishing escalation paths, automated alerting, and incident response procedures to ensure 24/7 reliability