Data Center Operations Engineer, Senior – Cloud AI - Riyadh, KSA
Qualcomm
- Location
- Riyadh, Saudi Arabia
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 22 approvals (FY2023)
- Posted
- 16h ago
Skills
About this role
Company: Qualcomm Middle East Information Technology Company LLC Job Area: Information Technology Group, Information Technology Group > IT Engineering General Summary:
About the Role
Qualcomm is leveraging its expertise in hardware‑accelerated AI to deliver powerful, energy‑efficient generative AI and computer vision workloads at rack‑scale in modern data centers. The Qualcomm Cloud AI team develops and operates a vertically integrated hardware and software platform for large‑scale AI inference acceleration across global deployments. We are seeking a Senior Data Center Operations Engineer to support Qualcomm’s Cloud AI data center environments through hands-on administration of MAAS, virtual machines, server operating systems, and operational infrastructure. This role is responsible for reliable deployment, imaging, provisioning, troubleshooting, and maintenance of systems that support Cloud AI deployments. The role requires strong senior-level technical execution, practical troubleshooting judgment, coordination across internal and external teams, and disciplined ownership of documentation, inventory, and operational procedures. YOU MUST BE A SAUDI NATIONAL TO BE CONSIDERED FOR THIS ROLE Key Responsibilities will include Operational Excellence & Reliability Administer MAAS servers, including deployment configuration, bare-metal provisioning, image management, imaging workflows, and troubleshooting of provisioning issues. Provision, configure, and support virtual machines, including compute, memory, storage, and network resource allocation based on operational requirements. Install, patch, maintain, and troubleshoot server operating systems across bare-metal and virtualized environments. Hardware & Infrastructure Leadership Support day-to-day server resources, storage coordination, and infrastructure incidents impacting Cloud AI data center operations. Troubleshoot server hardware, operating system, provisioning, VM, storage, and infrastructure issues, escalating when needed and supporting resolution through closure. Coordinate with HUMAIN, internal Qualcomm teams, and vendors to support deployments, maintenance activities, incident response, and operational follow-up. Maintain accurate system inventory, deployment records, configuration information, operational documentation, and standard operating procedures. Automation, Tooling & Continuous Improvement Contribute to improvements in monitoring, alerting, ticketing, and operational workflows to reduce repeat issues and improve support efficiency. Develop and maintain runbooks, checklists, and troubleshooting procedures for MAAS, VM provisioning, OS maintenance, and infrastructure support activities. Participate in incident reviews and help translate lessons learned into practical operational process improvements. Cross‑Functional Collaboration Serve as a trusted technical partner to Cloud AI software, systems, security, supply chain, and facilities teams. Communicate complex operational issues, tradeoffs, and recommendations clearly to engineers, leadership, and external partners. Influence operational roadmaps and priorities through data‑driven insights and recommendations. Required Qualifications & Experience You demonstrate senior-level engineering ownership through hands-on technical depth, sound troubleshooting judgment, and the ability to execute independently across operational infrastructure responsibilities: Hands-on experience administering MAAS or similar bare-metal provisioning platforms, including deployment, image management, and troubleshooting. Experience provisioning, configuring, and supporting virtual machines, including resource allocation and operational support. Strong Linux or server operating system administration experience, including installation, patching, maintenance, and troubleshooting. Ability to diagnose infrastructure incidents involving compute, storage, operating systems, provisioning workflows, and virtualization layers.