yoinka

Senior Machine Learning Engineer

HP Inc.

Bengaluru Karnataka IndiaSeniorH-1B sponsor company
Sign in to applyVerified 2h ago
Location
Bengaluru Karnataka India
Work model
On-Site
Level
Senior
H-1B history
37 approvals (FY2023)
Posted
Aug 25, 2026

Skills

AWSCI/CDDatabricksDeep LearningLLMMLOpsMachine LearningServerless

About this role

Senior Machine Learning Engineer Description - We are looking for a Senior MLOps Engineer to design, build, and operate the infrastructure that enables machine learning models and large language models to be deployed safely, reliably, and at scale. In this role, you will create the end-to-end capabilities required to move models from experimentation into production, expose them through secure and highly available endpoints, and enable users and applications to interact with AI-powered services. You will work across AWS and Databricks to establish robust CI/CD pipelines, model-serving infrastructure, observability, governance, rollback mechanisms, and operational standards. You will partner closely with data scientists, machine learning engineers, software engineers, security teams, and platform engineers. The ideal candidate combines strong cloud and DevOps engineering skills with a practical understanding of machine learning systems, LLM deployment patterns, and production reliability.

Key Responsibilities

MLOps Platform and Architecture Design and implement a scalable MLOps platform using AWS and Databricks. Define reference architectures and reusable deployment patterns for traditional machine learning models, deep learning models, and large language models. Build standardized workflows that move models from development and validation into staging and production. Develop self-service capabilities that allow data scientists and ML engineers to deploy models without manually managing infrastructure. Establish clear separation between development, testing, staging, and production environments. Design multi-region or multi-availability-zone architectures where required by business continuity and availability objectives. CI/CD and Model Deployment Build automated CI/CD pipelines for model code, inference services, infrastructure, configuration, and model artifacts. Implement automated testing across the deployment lifecycle, including: Unit testing Integration testing Model validation Data contract validation API and endpoint testing Security testing Performance and load testing Regression testing Automate model packaging, containerization, versioning, approval, promotion, and deployment. Support deployment strategies such as blue-green deployments, canary releases, shadow deployments, and controlled traffic shifting. Implement reliable rollback and roll-forward mechanisms for application code, infrastructure, model versions, prompts, and configuration. Ensure deployments are reproducible, auditable, and recoverable. Model and LLM Serving Design and operate secure, scalable, low-latency inference endpoints. Deploy models using appropriate services and patterns across AWS and Databricks, such as: Databricks Model Serving MLflow Model Registry Amazon SageMaker Amazon ECS or EKS AWS Lambda, where appropriate API Gateway Application Load Balancers Build synchronous, asynchronous, batch, and streaming inference capabilities. Design serving architectures for LLM-powered applications, including: Hosted foundation models Open-source models Fine-tuned models Retrieval-augmented generation Embedding services Vector search Prompt and response orchestration Tool-calling and agentic workflows Optimize inference performance, scalability, GPU utilization, concurrency, throughput, latency, and cost. Implement autoscaling, request throttling, queuing, caching, timeout handling, and graceful degradation. Reliability, Recovery, and Business Continuity Build recoverable model-serving endpoints with clearly defined recovery time and recovery point objectives. Implement automated health checks, failover mechanisms, retry policies, circuit breakers, and service recovery procedures. Design backup and recovery processes for: Model artifacts Model registry metadata Feature definitions Deployment configurations Infrastructure state Prompts and application configuration Vector indexes and knowledge-base assets Create disaster recovery procedures and

Listing verified 2h ago. Applications go through the company's official careers site.

← Back to Yoinka

Senior Machine Learning Engineer at HP Inc., Bengaluru Karnataka India | Yoinka