yoinka

Staff Software Engineer – SRE & AIOps

ServiceNow

Vancouver, British Columbia, CanadaFull TimeStaffH-1B sponsor company
Sign in to applyVerified 1h ago
Location
Vancouver, British Columbia, Canada
Employment
Full Time
Work model
On-Site
Level
Staff
H-1B history
185 approvals (FY2023)
Posted
1h ago

Skills

CI/CDKubernetesMachine LearningServiceNow

About this role

Company Description

It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started. Join us to put AI to work for people.

Job Description

About the Role ServiceNow is seeking a Staff Software Engineer – SRE & AIOps to drive infrastructure automation, operational resilience, and toil elimination across our hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, you will design and implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable our global engineering teams to operate reliably at scale.

This role combines strong hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with technical leadership influence across infrastructure teams. You will architect SRE tooling, develop auto-remediation capabilities, and establish patterns that allow ServiceNow's cloud platform to maintain high reliability while minimizing operational toil across follow-the-sun global teams. What you get to do in this role: Design, deploy, and operate enterprise-scale Kubernetes clusters across hybrid and multi-cloud environments, establishing governance, scaling policies, and operational practices that support high-velocity application deployments at 99.99%+ availability targets. Architect and implement closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures without human intervention, leveraging agentic AI and machine learning frameworks to predict failures, trigger preventive actions, and continuously reduce MTTR and on-call burden. Design and evolve the SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations, that support global follow-the-sun on-call operations and enable data-driven incident response. Establish SLO frameworks, error budgets, and alerting policies that balance rapid incident response with alert fatigue management, while developing automated runbooks and playbooks that empower on-call engineers to resolve issues autonomously. Design and maintain Infrastructure-as-Code frameworks and GitOps pipelines that enable reproducible, auditable infrastructure deployments across hybrid and multi-cloud environments with consistent security and compliance guardrails. Architect hybrid cloud and data center operations, spanning on-premises infrastructure, public cloud environments, and edge computing, including workload migration strategies, disaster recovery patterns, and cost optimization practices across multi-region deployments. Drive adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices, service mesh architectures, and network security controls that enable rapid, safe release cycles. Design on-call rotation schedules, escalation policies, and incident command systems that span across different time zones, ensuring 24/7 incident response while driving post-incident review processes that capture learning and drive systemic improvements. Mentor and guide junior SRE engineers and infrastructure teams on reliability patterns, incident investigation techniques, automation best practices, and agentic AI applications for infrastructure

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Staff Software Engineer – SRE & AIOps at ServiceNow, Vancouver, British Columbia, Canada | Yoinka