Staff ML Engineer
Cohesity
- Location
- Pune Panchshil India Office
- Work model
- On-Site
- Level
- Staff
- Posted
- Aug 28, 2026
Skills
About this role
Cohesity is the leader in AI-powered data security. Over 13,600 enterprise customers, including over 85 of the Fortune 100 and nearly 70% of the Global 500, rely on Cohesity to strengthen their resilience while providing Gen AI insights into their vast amounts of data. Formed from the combination of Cohesity with Veritas’ enterprise data protection business, the company’s solutions secure and protect data on-premises, in the cloud, and at the edge. Backed by NVIDIA, IBM, HPE, Cisco, AWS, Google Cloud, and others, Cohesity is headquartered in Santa Clara, CA, with offices around the globe. We’ve been named a Leader by multiple analyst firms and have been globally recognized for Innovation, Product Strength, and Simplicity in Design , and our culture . Want to join the leader in AI-powered data security? Staff ML Engineer Level: Staff Software Engineer (Level 5) Team: Gaia Emblem Engine Location: India | Type: Full-Time About the Role Cohesity is looking for a Staff ML Engineer (Level 5) to help design and scale the Gaia Emblem engine, the intelligence layer powering our next-generation data platform. In this role, you will architect and build large-scale, production-grade ML/LLM systems — spanning model serving, retrieval-augmented generation (RAG), GPU infrastructure, and distributed backend services. As a Staff-level engineer, you will set technical direction, mentor senior engineers, and partner closely with Product and Data Engineering to bring Gaia Emblem’s roadmap to life. Key Responsibilities • Architect and drive the technical roadmap for the Gaia Emblem engine, including LLM integration, RAG pipelines, and prompt-engineering frameworks. • Design and operate scalable, GPU-aware infrastructure on Kubernetes (vanilla K8s, OpenShift, EKS, GKE, or equivalent), including GPU scheduling, autoscaling, and costefficient resource utilization. • Build resilient, high-throughput distributed systems and microservices exposed via gRPC and REST APIs. • Own data engineering pipelines feeding ML workflows, including ingestion, transformation, and storage across SQL, NoSQL, and search systems (PostgreSQL, Redis, Elasticsearch). • Define and enforce engineering best practices: CI/CD, test automation, observability, and infrastructure-as-code (Terraform, Helm). • Lead design reviews, set coding and architectural standards, and mentor engineers across the organization. • Partner with Product, Applied ML, and SRE teams to translate business requirements into scalable technical solutions. • Drive root-cause debugging and performance optimization across distributed, GPUbacked services. • Evaluate and introduce new tools, frameworks, and infrastructure patterns to keep Gaia Emblem at the forefront of applied ML engineering.
Minimum Qualifications
Experience • 10+ years of professional software engineering experience, including 8+ years building distributed, production-grade backend systems. • 3+ years of hands-on experience with LLM-based systems, RAG architectures, or applied ML infrastructure. • Demonstrated experience operating services on Kubernetes (vanilla K8s, OpenShift, EKS, GKE, or similar) at production scale, including GPU-backed workloads. • Track record of technical leadership: driving architecture decisions, leading crossteam initiatives, and mentoring engineers at Staff/Senior level. Education • Bachelor’s degree in Computer Science, Engineering, or a related technical field required. • Master’s degree or PhD in Computer Science, Machine Learning, or a related field preferred. Required Skills • Languages: Python, Java, Go, JavaScript • ML/AI: LLMs, Prompt Engineering, Retrieval-Augmented Generation (RAG) • Infrastructure & Cloud: Kubernetes (platform-agnostic — vanilla K8s, OpenShift, EKS, GKE, or equivalent), GPU Scheduling, Docker, Helm Charts, Terraform, AWS, Cloud Computing • Distributed Systems & APIs: Distributed Systems, Microservices, gRPC, REST