yoinka

Staff Site Reliability Engineer

Levi Strauss

RemoteSpain - RemoteStaffH-1B sponsor company
Sign in to applyVerified 1h ago
Location
Spain - Remote
Work model
Remote
Level
Staff
H-1B history
17 approvals (FY2023)
Posted
Aug 25, 2026

Skills

AzureBigQueryGCPLLMRESTTerraform

About this role

Job Location:  Spain   Calling all originals: At Levi Strauss & Co., you can be yourself — and be part of something bigger.   We’re   a company of people who like to forge our own path and leave the world better than we found it. Who   believe   that what makes us different makes us   stronger.   So   add your voice. Make an impact. Find your fit — and your future.   We're seeking an exceptional   Staff Site Reliability Engineer  to join our Data & AI Platform Engineering team. In this role, you'll own and elevate the reliability, scalability, and operability of our enterprise data and AI platforms — the platforms that power everything from the design of our iconic jeans to the optimization of our global retail and supply chain. As a hands-on technical leader, you'll embody the principles of Google's SRE discipline: eliminating toil, engineering for reliability, and building a culture of shared ownership between development and operations. This is a unique opportunity to shape how a legendary brand runs production at scale on Google Cloud Platform, with a growing multi-cloud footprint across GCP and Azure.

About the Job

Reliability & Incident Management Define, instrument, and enforce   SLOs, SLIs, and error budgets  across all platform services, ensuring alignment with business and product commitments Drive continuous reduction in   MTTD and MTTR  through improved observability, automated alerting, and runbook-driven incident response Lead   blameless post-mortems  and translate findings into durable reliability improvements, ensuring systemic issues are eliminated rather than patched   Toil Reduction & Automation Systematically identify, measure, and eliminate operational toil; track toil percentage per sprint and enforce guardrails to keep it below 50% of engineering capacity Build and maintain   self-serve infrastructure capabilities  — enabling product and data engineering teams to provision, scale, and operate their own resources safely and consistently Automate deployment pipelines, configuration management, and operational workflows using   Infrastructure-as-Code  principles (Terraform, Helm, GitOps) Platform Engineering & Architecture Serve as the primary   GCP subject matter expert  — architecting and optimizing workloads across GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI Lead   multi-cloud architecture  decisions across GCP and Azure, ensuring consistent security posture, cost efficiency, and operational practices across environments Design and implement   self-healing infrastructure patterns , auto-scaling strategies, and capacity planning models to support high-availability data and AI platforms Champion   data security and governance  best practices — including encryption at rest and in transit, IAM least-privilege, secrets management, and audit logging AI, Agentic Systems & Modern Observability Apply SRE principles to   agentic AI workloads  — defining reliability expectations for LLM-based and multi-agent systems, including latency SLOs, fallback patterns, and model observability Partner with AI Platform teams to productionize agentic pipelines with robust monitoring, drift detection, and rollback capabilities Drive adoption of   AI-assisted operations tooling  to enhance observability, anomaly detection, and predictive incident management Leadership & Culture Guide and mentor junior and mid-level SREs — conducting code reviews, running reliability reviews, and elevating the team's engineering craft Collaborate cross-functionally with Data Engineering, Software Engineering, Security, and Product teams to embed reliability as a shared value from design through deployment Champion a culture of   psychological safety, continuous learning, and reliability excellence Communicate platform health, risk posture, and reliability roadmaps clearly to both technical and executive audiences About You   Required Qualifications Master's

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Staff Site Reliability Engineer at Levi Strauss, Spain - Remote | Yoinka