yoinka

Site Reliability Engineer II, GovCloud

Medallia

McLean, VirginiaFull TimeMid$103k – $155k/yr
Sign in to applyVerified 2h ago
Location
McLean, Virginia
Employment
Full Time
Work model
On-Site
Level
Mid
Salary
$103k – $155k/yr

Skills

AWSArgoCDAzureCI/CDGCPGitJenkinsKafkaKubernetesLinuxPostgreSQLPythonRedisTerraform

About this role

Overview

Medallia is the pioneer and market leader in Experience Management. Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the management of experiences, insights, and actions for candidates, customers, employees, patients, and residents alike.   We believe that every experience is a memory that can last a lifetime. Experiences shape the way people feel about a company. And they greatly influence how likely people are to advocate, contribute, and stay. At Medallia, we are committed to creating a world where organizations are loved by their customers and their employees. We empower exceptional people to create extraordinary experiences together.  Bring your whole self. The Role and Team We are growing our GovCloud team and looking for a Site Reliability Engineer II to help operate and improve Medallia’s US public-sector cloud platform. You will support federal agencies and other regulated customers in a highly available, secure, and compliant environment built on AWS GovCloud and Kubernetes.   This is a hybrid role based near Tysons, Virginia, with regular in-office collaboration and remote flexibility. You will be a hands-on engineer who helps operate and improve production systems, works closely with engineering and security teams, and grows toward owning larger systems while helping keep the platform reliable as we scale.

Responsibilities

Build, operate, and improve highly available, secure cloud infrastructure on AWS, including networking, access management, Kubernetes clusters, DNS, certificates, and shared platform services. Implement, operate and optimize  AWS cloud networking, specifically managing VPCs, subnets and routing, security groups/NACLs, VPC endpoints/PrivateLink and load balancing. Monitor, maintain and support production PostgreSQL — replication, backups and recovery, routine performance tuning, and upgrades — as part of the platform's data tier. Monitor production systems, respond to incidents, and drive fixes that improve reliability and reduce repeat issues. Develop and maintain Infrastructure-as-Code (primarily Terraform) and Kubernetes deployment workflows using Git, CI/CD, and GitOps practices. Improve observability across metrics, logs, and uptime monitoring; help tune alerts and operational runbooks. Partner with software engineering, security, and release management teams to deploy changes safely and resolve production issues. Contribute to platform upgrades, security patching, and compliance-driven maintenance in a regulated cloud environment. Participate in an on-call rotation for production support. Use AI-assisted tooling responsibly, with attention to security, privacy, and customer data boundaries. Learn the platform and grow your scope with mentorship from senior engineers. Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite.

Qualifications

Minimum Qualifications Eligibility Requirement: Must reside in the U.S. and hold U.S. Citizenship or a Green Card (Lawful Permanent Resident) to meet AWS GovCloud compliance requirements. Bachelor’s degree or equivalent experience in Computer Science or a related field. 2+ years of experience in Site Reliability Engineering, platform engineering, DevOps, or related production infrastructure roles (or equivalent hands-on production experience). Production experience with: Core services (IAM, compute, object storage, encryption/key management) and cloud networking on AWS (strongly preferred), Google Cloud (CGP), Azure, or a similar public cloud platform Terraform or comparable infrastructure-as-code tools Git and CI/CD pipelines Linux and foundational systems concepts (networking, DNS, TLS/certificates) PostgreSQL or another relational database — basic operations (queries, backups) Familiarity with Kubernetes concepts, container orchestration, and microservices management Programming and Automation: Proficiency in Python and/or Go experience to build

Listing verified 2h ago. Applications go through the company's official careers site.

← Back to Yoinka

Site Reliability Engineer II, GovCloud at Medallia, McLean, Virginia | Yoinka