yoinka

Site Reliability Engineer

Vannevar Labs

RemoteRemote; San Diego, CaliforniaSeniorClearance required
Sign in to applyVerified 1h ago
Location
Remote; San Diego, California
Work model
Remote
Level
Senior
Posted
2h ago

Skills

AWSAgileCI/CDDatadogDockerElasticsearchPulumiPythonSQLShellTerraform

About this role

Vannevar is a defense technology company building AI to deter our adversaries. In the 21st century, conflict moves at algorithmic speed and foresight equals firepower. Our agentic AI is purpose-built to compete with China—from cross-Strait conflict to gray zone coercion. Trained on the most mission-relevant datasets in defense, our technology models adversary behavior, simulates campaigns, and recommends the best course of action to decision makers. Our AI systems are some of the most trusted in the industry and actively used on the front lines of the Indo-Pacific to keep the peace and save lives.

Exceptional technology starts with exceptional people. Vannevar is a small agile team combining world-class engineers with veteran strategists who bring deep expertise in defense and tradecraft. We’re building a company defined by mission impact, user empathy, and disciplined growth. In just three years, we grew from $3M to $80M in ARR, achieved early profitability, and reached unicorn status—proving that disruption doesn’t require an ego, and staying power doesn’t mean standing still.

About the role

We are looking for an Site Reliability Engineer to own the reliability, health, and deployment automation of the platform at Vannevar Labs. In this role you'll be the person watching the system's pulse — monitoring dashboards, catching health issues before they become incidents, and owning the debugging process from first alert to resolution. Your decisions today will have a large impact on the company's future. We believe that simple systems are easier to understand, maintain, and scale. You will be making trade-offs as you work to ensure that our systems are prepared to operate reliably in high-side environments at scale. A strong sense of judgment matters here: knowing when to dig deeper into a problem yourself and when to pull in the right people to escalate. Clear, calm communication — during an incident and in day-to-day work — is a must.

What you'll do

• Monitor dashboards and system telemetry to detect health issues, performance degradation, and reliability risks — often before anyone else notices them.

• Own the debugging and incident response process end to end, exercising good judgment about when to investigate more deeply and when to escalate.

• Build logging, monitoring, and observability tooling to visualize the state of the platform and continuously mature our SRE practices.

• Develop, maintain, and be responsible for overall platform health, scaling, and capacity planning.

• Understand and help improve the deployment process, and automate build & deployment pipelines.

• Identify bottlenecks in engineering workflows and drive improvements that make the whole team faster and more reliable.

• Develop self-service tools and automation to improve engineering efficiency.

• Play a critical part in implementing a secure, robust, high-availability delivery pipeline.

• Communicate system status, trade-offs, and post-incident learnings clearly with teammates and stakeholders.

Qualifications

• 5+ years of experience in SRE, DevOps, or software engineering.

• Hands-on experience monitoring production systems and responding to incidents — comfortable owning a debugging process and making the call on when to dig in versus escalate.

• Excellent communication skills, especially the ability to stay clear and organized while troubleshooting live issues.

• Experience with the PLG stack, Datadog, or other enterprise monitoring/observability tools.

• Experience participating in an on-call rotation and running or contributing to post-mortems.

• Knowledge of AWS cloud technologies.

• Familiarity with

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Site Reliability Engineer at Vannevar Labs, Remote; San Diego, California | Yoinka