Machine Learning Ops Engineer
ResMed
- Location
- Bangalore, India
- Work model
- On-Site
- Level
- Mid
- Posted
- Sep 4, 2026
Skills
About this role
Global Technology Solutions (GTS) at ResMed is a division dedicated to creating innovative, scalable, and secure platforms and services for patients, providers, and people across ResMed. The primary goal of GTS is to accelerate well-being and growth by transforming the core, enabling patient, people, and partner outcomes, and building future-ready operations. The strategy of GTS focuses on aligning goals and promoting collaboration across all organizational areas. This includes fostering shared ownership, developing flexible platforms that can easily scale to meet global demands, and implementing global standards for key processes to ensure efficiency and consistency.
About the role
ResMed’s AI platform powers dozens of data scientists and a growing set of GenAI / Agentic AI products that touch patients, clinicians, and providers worldwide. We run on AWS and Kubernetes , provisioned with Terraform , and shipped through modern CI/CD. We are looking for AI/ML Platform Engineer whose core is Kubernetes, AWS, Terraform, AI and platform observability — someone who can design, build, and operate the platform end-to-end and instrument it so nothing is a mystery in production. You should also bring an AI working mindset : curious about how ML and agentic workloads run on the platform, comfortable partnering with data scientists and GenAI teams, and eager to grow the platform toward LLMOps and Agentic AI as those workloads scale. What you’ll do Design, build, and operate the AI/ML platform on AWS + Kubernetes — clusters, networking, IAM, storage, cost, and reliability. Provision and evolve infrastructure with Terraform ; treat infra as code with real review and rollback. Own CI/CD for data pipelines, ML models, and AI applications — from repo to production with confidence. Stand up and evolve the platform observability stack — Prometheus, Loki, Grafana / Datadog — for metrics, logs, traces, dashboards, alerting, and SLOs. Automate what shouldn’t be manual: environment provisioning, golden-path pipelines, self-serve tooling for data scientists. Partner with product, data science, and GenAI teams to make their workloads first-class on the platform — model serving, evaluation, cost/latency controls, and safe rollout. Run POCs to pull promising tech into the platform without accumulating debt. Participate in code review, mentoring, and process improvement; raise the engineering bar. What we’re looking for Must-have 3 + years of engineering experience in a complex, technical environment. Deep, hands-on Kubernetes in production. Hands-on AWS — comfortable with 3+ of: EKS, Lambda, EC2, S3, IAM, Networking (VPC, ALB/NLB), RDS, EMR, Glue, Athena, Batch, SageMaker, MWAA/Airflow. Working command of Terraform — modules, state, reviews, drift. Platform observability experience: Prometheus, Loki, Grafana and/or Datadog — metrics, logs, dashboards, alerting, SLOs. Strong production Python (and SQL for data work). Experience building CI/CD pipelines and APIs end-to-end — GitHub / GitHub Actions, CodePipeline or Jenkins. Hands-on working experience with an AI/ML platform in production — data science tooling, model lifecycle, feature / inference infrastructure, and self-serve enablement for DS and GenAI teams. Deploying AI agents / LLM workloads on Kubernetes — containerizing agent workloads, autoscaling (HPA/KEDA), GPU scheduling where needed, secure egress for tool calls, secrets and rate-limit management, and running long-lived / stateful sessions safely. Exposure to the modern AI / Agentic AI stack is required — working familiarity with at least a few of: an agent framework ( LangChain / LangGraph / CrewAI / AutoGen / Strands / Semantic Kernel / PydanticAI ), LLM serving ( vLLM , KServe , Ray Serve, TGI), a RAG / vector-store setup (OpenSearch, pgvector , Pinecone, Weaviate ), LLM observability ( Langfuse , LangSmith , Arize