Associate Director, Observability and Service Reliability
Kyndryl
- Location
- Toronto (KCA51701) HQ
- Work model
- On-Site
- Level
- Senior
- Posted
- Aug 31, 2026
Skills
About this role
Who We Are
At Kyndryl, we run and reimagine the mission-critical technology systems that drive advantage for the world’s leading businesses. We are at the heart of progress; with proven expertise and a continuous flow of AI-powered insight, enabling smarter decisions, faster innovation, and a lasting competitive edge. For our people—Kyndryls—that means doing purposeful work that powers human progress. Join us and experience a flexible, supportive environment where your well-being is prioritized and your potential can thrive.
The Role
Enterprise Observability Strategy Own and mature the enterprise observability and service reliability strategy. Define standards for monitoring applications, infrastructure, cloud platforms, networks, endpoints, APIs, databases, middleware, and other critical technology services. Establish expectations for metrics, logs, traces, events, synthetic monitoring, real user monitoring, digital experience, service health, and business transaction visibility. Create a consistent enterprise approach while allowing teams to use monitoring technologies suited to their platforms and services. Identify monitoring gaps, redundant capabilities, excessive alerting, and opportunities to improve visibility. Move the organization from traditional monitoring toward proactive, predictive, and automated operations. Service Reliability Engineering Establish and mature the organization’s service reliability framework. Partner with technical service owners to define monitoring requirements for critical applications and services. Ensure monitoring reflects the complete service, including application performance, infrastructure, dependencies, integrations, user experience, business transactions, capacity, and failure conditions. Define minimum observability requirements based on service criticality and business impact. Help teams establish meaningful Service Level Indicators, Service Level Objectives, availability targets, performance thresholds, and health measures. Use reliability data, incidents, problem records, capacity trends, and telemetry to identify systemic weaknesses and prioritize improvements. Technical Thought Leadership Serve as the enterprise technical authority for observability, monitoring, and service reliability. Provide architectural guidance to application, infrastructure, cloud, engineering, DevOps, SRE, cybersecurity, and operations teams. Influence solution design so services are observable, measurable, supportable, and resilient by design. Develop enterprise monitoring patterns, reference architectures, standards, and reusable capabilities. Guide technical teams in selecting appropriate monitoring methods and technologies for specific platforms and use cases. Evaluate emerging observability, AIOps, automation, analytics, and service reliability capabilities for measurable operational value. Observability Platform Leadership Provide strategic oversight for the enterprise observability and monitoring tool ecosystem. Lead the strategy, architecture, governance, adoption, and optimization of major platforms, including Dynatrace, Nexthink, and related enterprise monitoring technologies. Ensure monitoring tools operate as an integrated ecosystem rather than isolated platforms. Establish standards for instrumentation, tagging, alerting, dashboards, integrations, service mapping, ownership, and data quality. Partner with technical teams to maximize platform value while reducing tooling duplication and complexity. Manage strategic technology and vendor relationships to ensure observability investments deliver measurable operational value. Dynatrace Platform Strategy Provide strategic leadership for enterprise use of Dynatrace across applications, infrastructure, cloud, and digital services. Drive adoption of application performance monitoring, distributed tracing, real user monitoring, synthetic monitoring, infrastructure monitoring, logs, topology, service health, and intelligent problem