yoinka

HPC Operations Engineer

Tower Research Capital

New YorkMid$175k – $225k/yrH-1B sponsor company
Sign in to applyVerified 2h ago
Location
New York
Work model
On-Site
Level
Mid
Salary
$175k – $225k/yr
H-1B history
9 approvals (FY2023)
Posted
2h ago

Skills

AnsibleGrafanaLinuxMachine LearningPrometheusPythonShell

About this role

Tower Research Capital is a leading quantitative trading firm founded in 1998. Tower has built its business on a high-performance platform and independent trading teams. We have a 25+ year track record of innovation and a reputation for discovering unique market opportunities.

Tower is home to some of the world’s best systematic trading and engineering talent. We empower portfolio managers to build their teams and strategies independently while providing the economies of scale that come from a large, global organization.

Engineers thrive at Tower while developing electronic trading infrastructure at a world class level. Our engineers solve challenging problems in the realms of low-latency programming, FPGA technology, hardware acceleration and machine learning. Our ongoing investment in top engineering talent and technology ensures our platform remains unmatched in terms of functionality, scalability and performance.

At Tower, every employee plays a role in our success. Our Business Support teams are essential to building and maintaining the platform that powers everything we do — combining market access, data, compute, and research infrastructure with risk management, compliance, and a full suite of business services. Our Business Support teams enable our trading and engineering teams to perform at their best.

At Tower, employees will find a stimulating, results-oriented environment where highly intelligent and motivated colleagues inspire each other to reach their greatest potential.

Summary

This is an operations role, not a platform engineering one. It is about the daily health of the Research compute fleet: you will be the first line of support for HPC users and the primary owner of day-to-day operations across scheduling, compute, storage, and access. The work is transactional by nature, with tickets, triage, provisioning, and maintenance done well, every day.

The fleet has grown fast, with close to a thousand machines added recently, and this role exists so that infrastructure gets dedicated, high-standard operational care. You will keep an eye on system health, queues, node status, and service availability; work job failures, scheduler errors, and resource constraints as they come in; and drive every issue to resolution or a clean, well-documented escalation.

You will sit inside the HPC team, next to the engineers who build and run the platform. That proximity matters: your diagnostics feed their root-cause work, your runbooks capture what the team learns, and the recurring issues you surface become candidates for automation and permanent fixes. For someone who wants to grow into HPC engineering, this is a strong place to start.

Responsibilities

• Provide first-line support for HPC users across scheduling, compute, storage, and access issues.

• Troubleshoot job failures, scheduler errors, and resource constraints, driving each issue to resolution or a clean handoff.

• Triage infrastructure incidents: gather diagnostics, apply known fixes, and escalate to subject-matter experts when a problem extends beyond defined ownership.

• Monitor fleet health (queues, node status, storage, and service availability) and act on what you see before users have to report it.

• Carry out established operational procedures for maintenance, patching, and configuration updates across the Research fleet.

• Provision new machines into the Research fleet (OS installation, configuration, validation, and handoff into service), and handle reinstalls and decommissions as routine work.

• Write and maintain runbooks, knowledge-base articles, and user guides so the next occurrence of a problem is faster to fix than the first.

• Spot recurring issues and propose practical refinements, such as better workflows or automation candidates the HPC team can pick up, so the same ticket stops coming

HPC Operations Engineer at Tower Research Capital, New York | Yoinka