Associate Director, Incident, Problem, Change & Operational Resilience
Kyndryl
- Location
- Buenos Aires, Argentina
- Work model
- On-Site
- Level
- Senior
- Posted
- Aug 24, 2026
Skills
About this role
Who We Are
At Kyndryl, we run and reimagine the mission-critical technology systems that drive advantage for the world’s leading businesses. We are at the heart of progress; with proven expertise and a continuous flow of AI-powered insight, enabling smarter decisions, faster innovation, and a lasting competitive edge. For our people—Kyndryls—that means doing purposeful work that powers human progress. Join us and experience a flexible, supportive environment where your well-being is prioritized and your potential can thrive.
The Role
Major Incident Management & Executive Communications Lead the Major Incident Management function and ensure consistent execution across the organization. Direct the response to Severity 1, Severity 2, and other business critical incidents, ensuring rapid engagement, clear accountability, effective escalation, and timely service restoration. Provide senior operational leadership during major incidents and ensure technical teams remain focused on recovery priorities. Own and coordinate executive communications throughout major incidents. Translate complex technical information into concise, business focused updates covering impact, risk, recovery progress, dependencies, decisions required, and next steps. Establish appropriate communication cadences and ensure messaging remains consistent across executives, business stakeholders, technical teams, vendors, and customer facing organizations. Lead post incident reviews and ensure corrective and preventive actions have clear ownership and are completed. Continuously improve major incident processes, playbooks, tooling, automation, training, and response readiness. Incident & Problem Management Own the enterprise Incident and Problem Management practices and associated governance. Ensure incidents receive appropriate prioritization, ownership, escalation, investigation, resolution, and documentation. Drive improvements that reduce service disruption and improve Mean Time to Acknowledge, Engage, and Restore. Identify recurring issues, operational trends, and systemic risks through proactive and reactive Problem Management. Ensure significant and recurring incidents receive appropriate root cause analysis. Maintain governance over problem records, known errors, corrective actions, and remediation commitments. Hold accountable teams and service owners for resolving underlying causes rather than only restoring service. Change Management Own the enterprise Change Management practice and ensure technology changes appropriately balance speed, business value, stability, and risk. Define and maintain change governance, policies, risk classifications, approval models, and control requirements. Provide oversight for Change Advisory Board activities and ensure Normal, Standard, and Emergency changes follow appropriate governance. Use historical incident data, service criticality, testing evidence, dependencies, implementation plans, and rollback strategies to improve change risk assessment. Monitor change success rates, failed changes, emergency changes, and change related incidents. Partner with engineering, development, infrastructure, cybersecurity, and operations teams to improve change quality while reducing unnecessary administrative friction. Operational Resilience & Service Reliability Establish and mature an enterprise approach to technology operational resilience and service reliability. Identify services, technologies, and dependencies that create significant business or operational risk. Partner with service owners, engineering teams, cybersecurity, infrastructure, business continuity, and disaster recovery teams to improve the resilience of critical services. Use incident, problem, change, availability, and reliability data to identify systemic weaknesses and prioritize improvement initiatives. Drive programs that reduce repeat failures, eliminate single points of failure, strengthen recovery capabilities, and improve service reliability.