yoinka

Site Reliability Engineer

TJX Companies

Mississauga, ON L5R 0G1SeniorH-1B sponsor company
Sign in to applyVerified 1h ago
Location
Mississauga, ON L5R 0G1
Work model
On-Site
Level
Senior
H-1B history
8 approvals (FY2023)
Posted
Sep 1, 2026

Skills

AzureCI/CDDatadogGenAINew RelicSplunk

About this role

TJX Companies At TJX Canada, every day brings new opportunities for growth, exploration, and achievement. You’ll be part of our vibrant team that embraces diversity, fosters collaboration, and prioritizes your development. Whether you’re working in our Distribution Centers, Corporate Offices, or Retail Stores—WINNERS, HomeSense, and Marshalls, you’ll find abundant opportunities to learn, thrive, and make an impact. Come join our TJX family—a Fortune 100 company and the world’s leading off-price retailer. Here at TJX Canada, we are an equal opportunity employer committed to the inclusion and accommodation of all individuals.

Job Description

We’re looking for a Site Reliability Engineer to help shape the future of intelligent operations by leveraging AIOps, GenAI, automation, and self-healing technologies across enterprise platforms. In this role, you’ll drive initiatives that enhance system reliability, improve observability, and accelerate incident response while reducing manual effort through innovative automation solutions. You'll have the opportunity to influence modern engineering practices, work with cutting-edge AI technologies, and make a measurable impact on the scalability and performance of critical business applications. Join a collaborative team where innovation, continuous improvement, and operational excellence are at the heart of everything we do. Why Work With Us? We value integrity, respect, and teamwork, encouraging a unique and inclusive culture. Enjoy Associate discounts at our stores, available to you and eligible family members. Immediate access to our Group Benefits package, including a Health Care Spending Account, Retirement Savings Program, Associate & Family Assistance Program, and various well-being resources A competitive vacation package, paired with a Vacation Trade Program that allows you to opt in for an extra week.  Comprehensive training and development resources designed to help you learn, grow, and succeed. Exciting career paths with growth opportunities and tuition reimbursement to support your career progression. What You’ll Do: Implement SRE best practices across production support, incident management, problem management, change management, release, and deployment processes. Build and maintain automation scripts, runbooks, self-healing workflows, and operational tools to reduce manual effort and improve MTTR. Support AI-driven operational capabilities such as incident summarization, alert enrichment, anomaly detection, log analysis, event correlation, and root cause recommendations. Configure and enhance observability across logs, metrics, traces, dashboards, alerts, and synthetic monitoring. Work with monitoring/APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor , or similar tools. Support implementation of SLIs, SLOs, SLAs, service health metrics, reliability KPIs, and operational dashboards . Integrate monitoring, ITSM, CI/CD, cloud, and automation platforms using APIs, scripts, and workflows. Participate in incident response, troubleshooting, RCA, postmortems, and corrective/preventive action planning. Drive shift-left reliability by embedding monitoring, alerting, automation, and operational readiness into SDLC and CI/CD pipelines. Support reliable application and platform operations in Azure Cloud . Create and maintain documentation, runbooks, knowledge articles, and automation playbooks. Ensure AI-enabled operations follow enterprise security, compliance, data privacy, and responsible AI guidelines.

About You

6+ years of experience in SRE, DevOps, Production Support Engineering, Cloud Operations, Automation Engineering, or related roles. Hands-on experience in application support, incident management, problem management, change/release management, deployment support, monitoring, and documentation. Hands-on experience with observability/APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Site Reliability Engineer at TJX Companies, Mississauga, ON L5R 0G1 | Yoinka