yoinka

Staff Site Reliability Engineer

Johnson Controls

Richmond Hill-Ontario-CanadaStaff
Sign in to applyVerified 1h ago
Location
Richmond Hill-Ontario-Canada
Work model
On-Site
Level
Staff
Posted
Aug 27, 2026

Skills

AWSAzureDatadogGrafanaKubernetesTerraform

About this role

About Johnson Controls   Johnson Controls, a global leader in thermal management, mission-critical building systems, energy efficiency, and decarbonization, helps customers use energy more productively, reduce carbon emissions, and operate with the precision and resilience required in rapidly expanding industries such as data centers, healthcare, pharmaceuticals, advanced manufacturing, and higher education.   For more than 140 years, Johnson Controls has delivered performance where it really matters. Backed by advanced technology, lifecycle services and an industry-leading field organization, we elevate customer performance, turn goals into real-world results and help move society forward.   Visit  johnsoncontrols.com  for more information and follow @Johnsoncontrols on social platforms.

What you will do

OpenBlue from Johnson Controls is a cyber-secured smart building ecosystem that unifies data, AI, and automation to transform how buildings perform. By connecting systems that have historically stood apart and applying award-winning analytics, we give customers real-time visibility, predictive insight, and automated action across the entire building lifecycle. None of that reaches a customer without the platform underneath it. Our data platform, our enterprise SaaS portfolio, and OpenBlue Airwall run continuously for enterprise, public sector, and government customers, in environments where an outage or a data integrity problem carries real operational consequence for the buildings and the people inside them.   Johnson Controls is seeking a Staff Site Reliability Engineer. You will be the senior technical owner of escalated production problems, capable of debugging a failure across application code, data pipelines, cloud services, and network paths, and driving it through to permanent corrective action rather than a restart and a hopeful note in the ticket. You will also plan and execute the infrastructure work that keeps those platforms healthy, with Terraform as your native language and change safety as your standing constraint. You will be a core team member of our engineering department, and participate in the on-call rotation, and occasionally join customer conversations when the severity of an issue warrants an engineer in the room.   This role is based in Canada. We support Canadian government customers and commercial customers with Canadian data residency requirements, and this position works directly with those systems and their data.   How you will do it Application reliability and L3 escalation   Serve as the senior escalation owner for production issues across the   OpenBlue   Data Platform, our enterprise SaaS products, and Airwall   Debug complex, cross layer failures spanning application code, data pipelines, cloud infrastructure, and network paths   Lead root cause analysis and drive both interim and permanent corrective action to closure with the owning engineering teams, including the code or configuration change that prevents recurrence   Turn recurring escalations into engineering work by feeding defect patterns, reliability gaps, and supportability problems back into the product backlog   Improve detection ahead of the customer by strengthening instrumentation, monitors, dashboards, and runbooks in Datadog and Grafana   Participate in the on-call rotation and act as a senior technical lead during major incidents   Represent Engineering directly with customers during high severity incidents and post incident reviews when the situation calls for it   Infrastructure and platform engineering   Plan and execute infrastructure upgrades, migrations, and platform changes across Azure and AWS with minimal customer disruption   Own infrastructure as code in Terraform, including module design, state management, and drift remediation   Operate and improve Kubernetes workloads across capacity, autoscaling, resource limits, and deployment reliability   Raise release and

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Staff Site Reliability Engineer at Johnson Controls, Richmond Hill-Ontario-Canada | Yoinka