yoinka

Cloud Senior Monitoring and Observability Engineer

Leidos

Remote6314 Remote/Teleworker USSeniorH-1B sponsor company
Sign in to applyVerified 2h ago
Location
6314 Remote/Teleworker US
Work model
Remote
Level
Senior
H-1B history
26 approvals (FY2023)
Posted
Aug 24, 2026

Skills

AnsibleCI/CDDatadogGrafanaKubernetesLinuxNew RelicOpenShiftPrometheusServiceNowSplunk

About this role

The Senior Monitoring and Observability Engineer will support the Leidos SEC ISS2 contract by engineering, operating, and continuously improving enterprise monitoring and observability capabilities across hybrid infrastructure, cloud, and container platforms. This hands-on role is responsible for monitoring coverage, platform integration, agent deployment, tagging and normalization, dashboards, alerting, logs, APM, synthetic monitoring, automation, and operational integrations. Datadog is the primary enterprise observability platform used in the environment. Strong Datadog experience is preferred, however, candidates with substantial experience engineering and operating other enterprise monitoring or observability platforms will be considered where they demonstrate strong transferable monitoring expertise and the ability to rapidly develop proficiency with new technologies. The engineer partners with Operations and engineering teams to improve visibility, alert quality, incident detection, troubleshooting, performance analysis, and operational reliability across the enterprise.

Primary Responsibilities

In this Role you will: Engineer, operate, maintain, and continuously improve the enterprise monitoring and observability platform, including dashboards, monitors, metrics, logs, APM, synthetic monitoring, tagging, integrations, and related capabilities. Assess monitoring coverage across enterprise systems and applications, identify visibility gaps, and coordinate onboarding or remediation with the appropriate technical teams. Maintain monitoring coverage across Windows, Linux, cloud, OpenShift/Kubernetes, virtualized, database, network, storage, middleware, and application environments. Support monitoring and observability for Red Hat OpenShift, Kubernetes, OpenShift Virtualization, and virtual machine workloads running on OpenShift. Configure and troubleshoot monitoring agents, integrations, collectors, APIs, and related platform components. Build and maintain consistent tagging, metadata, dashboards, alerts, service health views, and operational reporting. Automate monitoring deployment, configuration, tagging, onboarding, upgrades, and integrations using Ansible, APIs, scripting, CI/CD, infrastructure-as-code, or similar technologies. Develop and maintain integrations between observability platforms, ServiceNow, notification systems, on-call workflows, and other enterprise operational systems. Use monitoring and observability data to troubleshoot complex performance and availability issues, support incident response and root-cause analysis, and recommend technical remediation. Correlate infrastructure, application, platform, and dependency telemetry to identify service degradation and recurring technical issues. Partner with Operations and engineering teams to improve monitoring coverage, alert quality, service visibility, incident detection, escalation, and operational response. Analyze telemetry and historical trends to identify capacity risks, recurring issues, monitoring gaps, and opportunities for improvement. Develop actionable performance, availability, capacity, and monitoring coverage reporting for technical and leadership stakeholders. Maintain monitoring standards, technical documentation, configuration guidance, and operational procedures.

Basic Qualifications

BS degree and 8-12 years of prior relevant experience, or Master's degree with 6-10 years of prior relevant experience. Additional relevant experience may be considered in lieu of degree requirements where permitted by contract. Strong hands-on experience engineering and operating enterprise monitoring or observability platforms. Strong Datadog experience is preferred; however, substantial experience with ScienceLogic SL1, SolarWinds, Dynatrace, New Relic, Splunk Observability, LogicMonitor, Prometheus/Grafana, or comparable enterprise platforms will be considered based on demonstrated monitoring and observability engineering expertise. Demonstrated ability

Listing verified 2h ago. Applications go through the company's official careers site.

← Back to Yoinka

Cloud Senior Monitoring and Observability Engineer at Leidos, 6314 Remote/Teleworker US | Yoinka