Senior Cloud Platform Engineer (SMTS)
Salesforce
- Location
- Washington Bellevue
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 498 approvals (FY2023)
- Posted
- Sep 18, 2026
Skills
About this role
To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword — it’s a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all. Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce.
Job Description
Senior Member of Technical Staff (SMTS) – Monitoring Cloud Infrastructure Location: Bellevue / Seattle / San Francisco / Palo Alto/ Hybrid / On-Site Role Level: Software Engineering Senior MTS Team: Infrastructure Engineering / Monitoring Cloud Position Overview As a Senior Member of Technical Staff (SMTS) within our Monitoring Cloud team, you will be a key owner and operator of the systems that keep Salesforce reliable. You won't just be "using" tools; you will be productizing infrastructure to ensure our monitoring capabilities evolve at the scale of our multi-cloud footprint. Your mission is to bridge the gap between high-level feature design and deep-system stability. From automating the "paved path" across AWS and GCP to securing air-gap environments for our most sensitive customers, you will ensure our monitoring stack is invisible, resilient, and intelligent. This is an AI-first engineering role. You will use AI-assisted development tools (e.g., Claude Code) as the default for every inner-loop activity, code authoring, Terraform and Kubernetes scaffolding, test generation, refactoring, log/trace analysis, runbook drafting, and documentation. We expect AI to compound your throughput on routine implementation so you can focus your human judgment on architecture, security, on-call response, and customer outcomes.
Core Responsibilities
1. Infrastructure as Code (IaC) & Automation Design and implement automation frameworks using Terraform and Kubernetes to manage monitoring infrastructure. Standardize "paved path" deployments across AWS and GCP, eliminating manual configuration errors and ensuring global consistency. Use AI-assisted tooling as the default for authoring, refactoring, and reviewing IaC modules, Helm charts, and automation scripts while directing intent, validating output, and owning the final result. 2. Infrastructure Upkeep & Productization Own the lifecycle of the Monitoring Cloud stack, including version upgrades and performance tuning. Productize core components (e.g., Grafana, custom Terraform providers) to make them consumable as reliable services by internal engineering teams. Leverage AI for upgrade planning, release-note analysis, migration scaffolding, and boilerplate-heavy productization work (API wiring, schema plumbing, SDK generation), while retaining accountability for design and rollout. 3. Secure & Air-Gapped Operations Deploy and manage the full monitoring stack within highly isolated, air-gapped environments. Ensure that our most secure customer segments receive the same level of observability and reliability as our public cloud offerings. Apply AI assistance during development of the artifacts that ship into these environments; operate them in-network with the disciplined, human-driven workflows these environments require. 4. Operational Excellence & Health Participate in the team’s on-call rotation, providing the deep technical expertise required to maintain strict SLAs and availability targets. Conduct root-cause analysis (RCA) for complex system failures and implement long-term preventative fixes. Address support requests