yoinka

PCAI And AI Factory

Juniper Networks

Bengaluru, Karnātaka, IndiaMidH-1B sponsor company
Sign in to applyVerified 1h ago
Location
Bengaluru, Karnātaka, India
Work model
On-Site
Level
Mid
H-1B history
140 approvals (FY2023)
Posted
Aug 18, 2026

Skills

AirflowAnsibleComputer VisionGrafanaKubernetesLLMNLPSpark

About this role

PCAI And AI Factory This role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office.

Who We Are

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.

Job Description

HPE Operations is our innovative IT services organization. It provides the expertise to advise, integrate, and accelerate our customers’ outcomes from their digital transformation. Our teams collaborate to transform insight into innovation. In today’s fast paced, hybrid IT world, being at business speed means overcoming IT complexity to match the speed of actions to the speed of opportunities. Deploy the right technology to respond quickly to market possibilities. Join us and redefine what’s next for you. What you’ll do: We are seeking a Subject Matter Expert (SME) – Admin, Operate & Manage (HPE PCAI & AI Factory Solutions) to manage and optimize HPE’s next-generation AI infrastructure platforms. The ideal candidate will have deep hands-on expertise in AI, HPC, and GPU-accelerated environments, with strong knowledge of HPE Ezmeral, NVIDIA AI Enterprise, Containerized workloads, and Automation frameworks. This role focuses on the operational stability, lifecycle management, and continuous improvement of large-scale Private Cloud for AI (PCAI) and AI Factory deployments.

Key Responsibilities

1. Platform Administration • Administer and maintain HPE PCAI and AI Factory environments, ensuring optimal uptime and performance. • Manage compute nodes (HPE DL380a, DL325, Cray XD670), GPU clusters (NVIDIA L40S/H100/H200), and InfiniBand NDR networks. • Administer virtualization and container platforms such as vSphere, RHEL/RHOS, Ezmeral Runtime Enterprise, Kubernetes, and Rancher Harvester. • Perform configuration, patching, version upgrades, and firmware updates across hardware and software layers. 2. Operational Monitoring & Incident Management • Proactively monitor system health using DCGM, NetQ, Grafana, and Exivity dashboards. • Handle alerts, performance anomalies, and incidents across GPU, network, and storage layers. • Lead root cause analysis (RCA) and corrective action plans to prevent recurring issues. • Maintain operational documentation, runbooks, and incident logs. 3 . Lifecycle & Configuration Management • Manage cluster lifecycle through Ansible, AWX, HPE Performance Cluster Manager (HPCM), and SLURM. • Oversee automation for provisioning, scaling, and patch management of Compute and Containerized workloads. • Manage configuration changes, infrastructure templates, and version baselines in production and staging environments. 4. AI Platform & Software Operations • Operate HPE Ezmeral Unified Analytics, Data Fabric, and AI Essentials platforms. • Support NVIDIA AI Enterprise (NVAIE) components including NIMs, NeMO frameworks, and RAPIDS runtime. • Manage and monitor AI/ML workloads (LLM, NLP, Computer Vision, Chatbots) on containerized clusters. • Ensure smooth operation of development tools like Jupyter, Spark, Airflow, MLflow, Kubeflow, and Ray. 5. Storage & Data Operations • Administer VAST, WEKA, and Alletra MP storage solutions for file, object, and distributed storage. • Monitor storage performance, replication, and capacity utilization. • Coordinate with storage engineering teams for performance optimization

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

PCAI And AI Factory at Juniper Networks, Bengaluru, Karnātaka, India | Yoinka