Senior Software Engineer - SRE & AIOps
ServiceNow
- Location
- Santa Clara, CALIFORNIA, United States
- Employment
- Full Time
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 185 approvals (FY2023)
- Posted
- 1h ago
Skills
About this role
Company Description
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started. Join us to put AI to work for people.
Job Description
About the role: ServiceNow is seeking a Senior Software Engineer - SRE & AIOps to contribute to infrastructure automation, operational resilience, and toil elimination across our hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, you will implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable our global engineering teams to operate reliably at scale.
This role combines solid hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with growing technical leadership capabilities. You will contribute to SRE tooling design, develop auto-remediation capabilities, and help establish patterns that maintain ServiceNow's cloud platform reliability while minimizing operational toil across follow-the-sun global teams. What you get to do in this role: Deploy, operate, and troubleshoot production Kubernetes clusters across hybrid and multi-cloud environments, maintaining operational standards and supporting high-velocity application deployments. Implement and maintain closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures, leveraging automation frameworks and machine learning insights to reduce MTTR and on-call burden. Contribute to the design and evolution of SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations that support global on-call operations. Develop and maintain SLO frameworks, alerting policies, and automated runbooks that empower on-call engineers to resolve issues autonomously while managing alert fatigue. Build and maintain Infrastructure-as-Code frameworks and GitOps pipelines that enable reproducible infrastructure deployments across hybrid and multi-cloud environments with security and compliance guardrails. Support hybrid cloud and data center operations, including on-premises infrastructure, public cloud environments, and workload optimization across multi-region deployments. Contribute to adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices and network security controls. Support on-call rotation operations and incident response processes across different time zones, helping develop runbooks and contributing to post-incident reviews that drive continuous improvement. Share knowledge and mentor junior SRE engineers on reliability patterns, incident investigation techniques, and automation best practices. Champion a culture of blameless incident analysis, data-driven decision-making, and continuous improvement through knowledge sharing and documentation. Identify and systematically automate repetitive operational tasks, from infrastructure provisioning to incident response, improving team efficiency and capacity.
Qualifications
To be successful in this role you have: Kubernetes Proficiency: Solid hands-on experience operating production Kubernetes clusters, including deployment models, pod orchestration, resource management, network policies, and troubleshooting runtime