Staff Engineer, Site Reliability Engineering
General Motors
- Location
- Markham, Ontario, Canada
- Work model
- On-Site
- Level
- Staff
- H-1B history
- 267 approvals (FY2023)
- Posted
- Sep 1, 2026
Skills
About this role
Job Description
Vacancy Status: Yes - This posting is for an existing vacancy within the organization and is open to new applications. (Backfill) AI Disclosure: As part of the application process, Artificial Intelligence will be used in the hiring process for this role Work Arrangement: Hybrid: This role is categorized as hybrid. This means the successful candidate is expected to report to Markham office three times per week, at minimum.
About the role
General Motors is transforming the automotive landscape through its next-generation Software-Defined Vehicle platform. Data is central to that transformation , powering safety, personalization, energy optimization, operational decision-making, and connected customer experiences. We are seeking a Staff Engineer to help make GM’s data platforms reliable, observable, operable, and scalable. This is a senior technical leadership role for someone who can move comfortably between system-level design, production operations, incident response, automation, and customer partnership. You will help define and spread the engineering patterns that make services easier to operate . You will work with SRE, data engineering, infrastructure, developer experience, application, and product teams to improve reliability from design through production and continuously improve how the organization operates . What you’ll do Lead the design and implementation of scalable, fault-tolerant, and observable infrastructure supporting vehicle telemetry, data ingestion, and platform operations. Lead production readiness efforts across multiple teams—engaging directly in code, shaping reliability standards, guiding architectural improvements, and ensuring applications launch with resilient deployments, strong observability, and predictable operations. Design, implement, and improve CI/CD delivery pipelines that make releases repeatable, safe, observable, and fast. Establish appropriate quality gates, artifact promotion, deployment verification, progressive delivery, and rollback practices . Partner across SRE, product, and application teams to design and implement meaningful SLOs, SLIs, observability, monitoring and alerting, runbooks, and operational best practices. Build and improve reusable AI workflows, skills, and evaluations. Apply appropriate validation techniques, including regression testing, structured evaluations, and LLM-as-a-judge approaches where useful. Automate operational work, including incident intake, triage, diagnostics, remediation, evidence collection, service requests, and customer-facing status workflows. Participate in a weekly on-call rotation with 12-hour shifts; the rotation cycles every eight weeks. Lead incident response, communicate clearly under pressure, and coordinate effective mitigation and recovery. Participate in post-incident reviews and drive durable, system-level fixes that prevent recurrence rather than relying on short-term patches or repeated manual workarounds. Partner directly with internal customers to understand their needs, explain technical trade-offs, and improve service outcomes with tact, empathy, and clear communication. Influence technical direction across teams, mentor engineers, and raise engineering standards through design reviews, code reviews, documentation, and hands-on leadership with cross-functional engineering projects. Balance reliability, performance, security, delivery speed, and cost when making technical decisions—especially under pressure. What you bring 8+ years in SRE, DevOps, or systems engineering, including experience managing or mentoring high-impact teams. Track record of designing and building and maintaining high-scale, cloud-native systems in production (preferably Azure, AWS, or GCP). Hands-on experience architecting observability patterns, including standardized instrumentation, OTEL c ollector