yoinka

Senior Software / Site Reliability Lead Engineer

General Dynamics

RemoteUS--TeleworkSenior$142.7k – $158.3k/yrClearance required
Sign in to applyVerified 1h ago
Location
US--Telework
Work model
Remote
Level
Senior
Salary
$142.7k – $158.3k/yr

Skills

CI/CDCloudFormationDatadogDockerGrafanaKubernetesPrometheusPythonShellTerraform

About this role

Basic Qualifications Bachelor's degree in Software Engineering, or related Science, Technology, Engineering or Mathematics field, plus a minimum of 8 years of relevant experience; or Master's degree, plus 6 years relevant experience.CLEARANCE REQUIREMENTS: Ability to obtain a Department of Defense Secret security clearance is required at time of hire. Applicants selected will be subject to a U.S. Government security investigation and must meet eligibility requirements for access to classified information. Due to the nature of work performed within our facilities, U.S. citizenship is required. Responsibilities for this Position What You Will Own

Cross-pod reliability standards . Set the reliability bar and ensure it is met consistently across applications. Collaborate with Functional SREs to connect technical reliability metrics to business-side outcomes. You own the engineering signal; together you tell the full reliability story. SLOs and reliability metrics . Own definitions of service level objectives for every AI service that goes to production. Establish error budgets and use them to drive engineering decisions — not just measure uptime. Monitoring and observability . Implement and maintain the full observability stack — logging, metrics, tracing, and dashboards. You will know when something is degrading before users do. Design and manage alerting infrastructure that tells you what's wrong, not just that something is wrong. Alerts you build catch real problems; they don't cry wolf. Incident response. Own on-call procedures, escalation paths, and incident management end-to-end. Lead post-incident reviews and maintain the reliability improvement backlog. When something breaks, you coordinate the response and ensure it doesn't break the same way again. Production Readiness. Define and enforce the criteria that determine whether an AI service is ready for production. You are the gate between "it works in dev" and "it's ready to ship." Toil elimination. Identify and automate repetitive operational tasks. If a human is doing something a script could do, you fix that.

What You Won't Own

Infrastructure provisioning — IT provides the infrastructure; you define what's needed and validate it works Business process decisions or backlog prioritization Business-side reliability metrics - you partner with the Functional SRE on those, but they own that domain

What Makes This Role Different

AI services have failure modes that traditional applications don't — model drift, token budget exhaustion, prompt injection, upstream data quality degradation. You will build monitoring for problems that most SRE teams have never encountered. You are applying SRE principles from scratch. There is no existing SRE practice to inherit — you will define it for the platform. Your production readiness criteria directly determine whether AI services go live. You have real authority to say "not ready." You operate across projects simultaneously — embedded deeply enough to understand large-scale systems, while maintaining consistent standards across all projects. Your software engineering background means you can engage directly with development teams at the design level — catching reliability problems before they become operational ones.

Required Qualifications

Bachelor’s degree in Computer Science, Software Engineering, or a related field, plus 8 years of experience; or Master’s degree plus 6 years of experience Production SRE or DevOps experience — you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines Hands-on experience with monitoring and observability tools — Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that caught real problems. Strong scripting and automation skills — Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar) Experience with containerized environments

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Senior Software / Site Reliability Lead Engineer at General Dynamics, US--Telework | Yoinka