Senior DevOps Engineer
Qualys
- Location
- Pune
- Work model
- On-Site
- Level
- Senior
- Posted
- Aug 28, 2026
Skills
About this role
Come work at a place where innovation and teamwork come together to support the most exciting missions in the world!
Role
Overview We are seeking a Senior DevOps Engineer, Observability to design, build, operate, and continuously improve a high-scale observability platform built around ClickHouse, HyperDX, OpenTelemetry, Kubernetes, and modern DevOps automation practices. This role is intended for a senior hands-on engineer who can take end-to-end ownership of observability infrastructure for production environments. The engineer will be responsible for building reliable telemetry pipelines for logs, metrics, and distributed traces, optimizing ClickHouse for large-scale observability workloads, operating HyperDX for troubleshooting and application performance analysis, and partnering with application, platform, and SRE teams to improve production visibility, reliability, and incident response. The ideal candidate is deeply technical, operationally disciplined, automation-oriented, and comfortable working in high-volume, production-critical environments. Senior-level expectation: This role requires ownership beyond task execution, including technical judgment, production accountability, automation-first delivery, mentoring, and clear communication during incidents and escalations. Key Responsibilities • Design, build, and operate scalable observability platforms using ClickHouse, HyperDX, OpenTelemetry, Kubernetes, Prometheus, Grafana, Alertmanager, Fluent Bit, and Filebeat. • Architect and optimize ClickHouse for high-volume observability workloads, including logs, traces, metrics, and telemetry analytics. • Design and manage ClickHouse schemas, partitioning strategies, ordering keys, TTLs, materialized views, retention policies, storage efficiency, and query optimization. • Build, operate, and improve HyperDX for log search, distributed tracing, service analysis, dashboards, telemetry correlation, troubleshooting, and root-cause analysis. • Build and maintain scalable OpenTelemetry Collector pipelines for collecting, processing, enriching, filtering, sampling, and routing telemetry data. • Implement reliable correlation across logs, metrics, and traces to support faster application troubleshooting, service dependency analysis, and incident resolution. • Design resilient telemetry pipelines with batching, queuing, retries, backpressure handling, sampling, rate limiting, and cardinality controls. • Deploy and operate observability infrastructure on Kubernetes, with focus on scalability, high availability, capacity planning, resiliency, and operational safety. • Automate infrastructure deployment, configuration management, platform upgrades, application onboarding, and recurring operational tasks using DevOps best practices. • Build and maintain CI/CD workflows using Jenkins, infrastructure automation using Terraform and Ansible, and service discovery or secrets management integrations using HashiCorp Consul and Vault. • Partner with application engineering, platform engineering, SRE, and operations teams to troubleshoot production issues, improve observability coverage, and reduce mean time to detect and resolve incidents. • Analyze production performance issues across infrastructure, applications, telemetry pipelines, and ClickHouse queries, and drive corrective actions to closure. • Define and improve standards for telemetry instrumentation, log quality, metric hygiene, trace propagation, dashboard design, alert quality, and production readiness. • Participate in incident response, post-incident reviews, capacity planning, operational reviews, and remediation tracking for observability services. • Mentor junior engineers, review designs and automation changes, and raise the overall technical and operational maturity of the team. Required Qualifications • 6+ years of experience in DevOps, SRE, Platform Engineering, Infrastructure Engineering, or Observability Engineering roles. • Strong