HPC Systems Administrator
Eli Lilly
- Location
- San Francisco, California, United States of America
- Employment
- Full Time
- Work model
- On-Site
- Level
- Mid
- Posted
- Sep 16, 2026
Skills
About this role
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. Where AI Meets Medicine: Build the Future of Drug Discovery in the Heart of Silicon Valley! Making medicine that’s never been made means doing what’s never been done. If you’re an engineer, scientist, or builder who thrives on problems no one has solved before, this is your invitation; we want you on the team. We are ready to challenge the status quo and push medicine forward, all in the name of health. Are you up for the challenge? If so, join us! About the Lilly and NVIDIA Partnership Lilly and NVIDIA are launching a new AI co-innovation lab in the heart of Silicon Valley — an up-to-$1 billion, multi-year commitment to solve drug discovery’s toughest challenges. The lab brings Lilly scientists, technologists, chemists and biologists together with NVIDIA engineers under one roof. Together, we are building purpose-built foundation and frontier AI models trained on Lilly data at scale, tightening the feedback loop between automated wet labs and computational dry labs, designing the next generation of medicines for millions of patients across the globe. What You’ll Be Doing The HPC Systems Administrator will build and operate scalable AI and high-performance computing platforms that power advanced machine learning and scientific workloads. You will partner with AI scientists, engineers, and domain experts to enable efficient model training, inference, and experimentation across GPU, cloud, and on-premises environments. Through platform engineering and automation, you will improve productivity, performance, and access to advanced computing resources while ensuring the reliability, availability, and efficiency of Lilly's AI and HPC infrastructure. How You’ll Succeed Deliver highly available, secure, and performant AI and HPC platforms that meet the needs of research and engineering teams. Drive operational excellence through automation, standardization, monitoring, and continuous improvement. Balance infrastructure reliability, scalability, and cost efficiency across environments. Enable efficient ML workflows through automation for orchestration, resource scheduling, data access, and reproducibility Collaborate effectively across scientific, engineering, and infrastructure teams to solve complex technical challenges. Adapt quickly to evolving AI, GPU, and HPC technologies and translate new capabilities into business value. What You Should Bring Deep expertise in Linux systems administration, automation, and infrastructure management, with strong scripting skills in Python, Bash, and/or Ansible. Experience building, administering, and optimizing large-scale HPC, GPU, or AI/ML computing environments, including job scheduling and resource management platforms such as Slurm or Grid Engine. Proficiency with automation, configuration management, and container technologies, including Ansible, Kubernetes, Docker, and related tooling. Solid understanding of distributed computing, high-performance networking, storage architectures, and cluster infrastructure. Experience supporting large-scale distributed training and inference workloads across multi-GPU and multi-node environments. Knowledge of GPU infrastructure, hardware lifecycle management, monitoring, and observability practices. Demonstrated ability to solve complex infrastructure challenges, identify root causes, and