yoinka

Senior Lead Engineer, Platform Operations & Observability

McKesson

CAN, QC, Montreal, Ville Saint-LaurentSeniorH-1B sponsor company
Sign in to applyVerified 1h ago
Location
CAN, QC, Montreal, Ville Saint-Laurent
Work model
On-Site
Level
Senior
H-1B history
49 approvals (FY2023)
Posted
Aug 14, 2026

Skills

AzureCI/CDKubernetesPrometheus

About this role

McKesson is an impact-driven, Fortune 10 company that touches virtually every aspect of healthcare. We are known for delivering insights, products, and services that make quality care more accessible and affordable. Here, we focus on the health, happiness, and well-being of you and those we serve – we care. What you do at McKesson matters. We foster a culture where you can grow, make an impact, and are empowered to bring new ideas. Together, we thrive as we shape the future of health for patients, our communities, and our people. If you want to be part of tomorrow’s health today, we want to hear from you.

About the Role

Team/Project: Canada B2C Digital Solution. Main application in the portfolio is a B2C Platform for pharmacy patients. Team of around 20 people. McKesson is seeking a Senior Lead Engineer, Platform Operations & Observability to lead the reliability, observability, and operational excellence of enterprise healthcare technology platforms. In this role, in accordance with Application Monitoring & Observability Lead, you will drive monitoring strategies, incident management practices, change governance, and root cause analysis initiatives while helping build scalable, secure, and resilient systems. You will collaborate with software engineering, platform, security, and operations teams to improve service reliability, automate operational processes, and establish best practices for production readiness. This position also provides technical leadership and mentorship to engineering teams while influencing reliability standards and long-term platform strategy.

What You'll Do

Lead monitoring and observability strategies across enterprise applications, platforms, and services. Design, implement, and optimize dashboards, alerts, telemetry, logging, and performance monitoring solutions. Drive incident management processes, major incident response, escalation coordination, service restoration activities and postmortem incident. Conduct root cause analysis (RCA) investigations and lead corrective and preventive action planning. Partner with engineering teams to improve platform reliability, resiliency, scalability, and operational readiness. Lead change management reviews and promote safe deployment and release practices. Ability to execute regression and validation test plans following each production deployment. Provide technical leadership, coaching, and mentoring to engineers while establishing engineering best practices. Influence architecture, automation, CI/CD, and operational excellence initiatives supporting enterprise platforms.

Basic Requirements

7+ years of professional experience in Software Engineering, Site Reliability Engineering, Platform Engineering, DevOps, or related technical roles. Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent experience. Experience supporting large-scale production environments and enterprise applications. Hands-on experience with monitoring, observability, logging, alerting, and application performance monitoring tools. Proven experience leading incident management and production support activities. Experience performing root cause analysis and implementing preventive solutions. Experience with CI/CD, automation, DevOps practices, and software delivery pipelines. Experience with microservices, APIs, distributed systems, and cloud-based architectures. Lead production readiness reviews and operational acceptance activities prior to major releases. Preferred Skills / Experience Experience with tools such as Dynatrace, Prometheus, Dotcom Monitor or similar observability platforms. Experience with Kubernetes, containers, and cloud platforms such as Azure. Knowledge of ITIL-aligned incidents, problems, and change management practices. Experience defining SLAs, MTTR and service reliability metrics. Demonstrated technical leadership and mentorship of engineering teams. Experience operating in regulated or highly compliant environments.

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Senior Lead Engineer, Platform Operations & Observability at McKesson, CAN, QC, Montreal, Ville Saint-Laurent | Yoinka