yoinka

Senior DevOps Engineer – Observability Platform

Qualcomm

Chennai, Tamil Nādu, IndiaSeniorH-1B sponsor company
Sign in to applyVerified 6h ago
Location
Chennai, Tamil Nādu, India
Work model
On-Site
Level
Senior
H-1B history
22 approvals (FY2023)
Posted
Jul 9, 2026

Skills

KubernetesAWSGrafanaPrometheusPythonDockerJavaC++ShellTerraformCloudFormation

About this role

Company: Qualcomm India Private Limited Job Area: Engineering Group, Engineering Group > Software Engineering General Summary: As a Senior DevOps Engineer – Observability Platform , you will be responsible for building and maintaining scalable, reliable infrastructure and deployment pipelines with a strong emphasis on observability — metrics, logs, and traces — across systems running on Kubernetes and AWS. You will work closely with development teams to improve development velocity while ensuring system reliability, security, and performance. This role is critical in providing a standardized, observability platform that gives both internal engineering teams and external, customer-facing services deep, reliable visibility into system health, performance, and reliability Minimum Qualifications: •    Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience. OR Master's degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Software Engineering or related work experience. OR PhD in Engineering, Information Systems, Computer Science, or related field. • 2+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc. Infrastructure Management : Design, implement, and maintain cloud-based infrastructure using Infrastructure as Code principles Automation : Develop automation scripts and tools to streamline operations and eliminate manual processes Containerization : Manage containerization strategies and orchestration using Docker and Kubernetes Observability Platform: Design, build, and operate a standardized, self-service metrics, logs, and tracing platform (Prometheus, Grafana, Loki, OpenTelemetry) serving both internal teams and external, customer-facing services running on Kubernetes and AWS Instrumentation & Telemetry: Partner with engineering teams to instrument applications and infrastructure, standardizing telemetry collection with OpenTelemetry SLOs & Alerting: Define and maintain SLIs/SLOs and error budgets, build actionable dashboards, and tune alerting to maximize signal and reduce noise Performance Optimization : Use observability data to analyze and optimize system performance, scalability, and cost-efficiency Documentation : Create and maintain thorough documentation for infrastructure, deployment processes, and operational procedures Incident Response & Escalation: Provide second-tier engineering escalation during business hours and own the telemetry, SLO, and alerting tooling that powers incident detection and reduces MTTD/MTTR; front-line 24/7 on-call is owned by the dedicated SRE team, not observability engineers. Lead post-mortem analysis for observability-platform incidents Requirements Qualifications 5+ years of experience in DevOps, Observability, or similar roles, including hands-on production experience operating Kubernetes based stack. Advantage - Strong background in software development with security focus Technical Skills Cloud Platforms : Extensive hands-on experience with AWS, including its observability services (CloudWatch, X-Ray, Amazon Managed Service for Prometheus, Amazon Managed Grafana) Infrastructure as Code : Proficiency with Terraform, AWS CloudFormation, or similar IaC tools Containerization : Advanced knowledge of Docker and Kubernetes ecosystem Observability Stack: Hands-on experience with Prometheus, Grafana, Loki, Tempo or Jaeger, OpenTelemetry, and Alertmanager; experience scaling metrics storage with Thanos, Mimir, or Cortex Programming/Scripting : Strong coding skills in Python, Bash, or Go Soft Skills Problem-Solving:  Excellent analytical and troubleshooting skills. Communication:  Strong verbal and written communication skills. Collaboration:  Ability to work effectively in a team environment and collaborate with cross-functional teams. Leadership:  Proven leadership skills and the ability

Listing verified 6h ago. Applications go through the company's official careers site.

← Back to Yoinka