yoinka

GPU Systems Engineer

Tower Research Capital

New YorkMid$200k – $300k/yrH-1B sponsor company
Sign in to applyVerified 2h ago
Location
New York
Work model
On-Site
Level
Mid
Salary
$200k – $300k/yr
H-1B history
9 approvals (FY2023)
Posted
2h ago

Skills

AnsibleLinuxMachine LearningPythonREST

About this role

Tower Research Capital is a leading quantitative trading firm founded in 1998. Tower has built its business on a high-performance platform and independent trading teams. We have a 25+ year track record of innovation and a reputation for discovering unique market opportunities.

Tower is home to some of the world’s best systematic trading and engineering talent. We empower portfolio managers to build their teams and strategies independently while providing the economies of scale that come from a large, global organization.

Engineers thrive at Tower while developing electronic trading infrastructure at a world class level. Our engineers solve challenging problems in the realms of low-latency programming, FPGA technology, hardware acceleration and machine learning. Our ongoing investment in top engineering talent and technology ensures our platform remains unmatched in terms of functionality, scalability and performance.

At Tower, every employee plays a role in our success. Our Business Support teams are essential to building and maintaining the platform that powers everything we do — combining market access, data, compute, and research infrastructure with risk management, compliance, and a full suite of business services. Our Business Support teams enable our trading and engineering teams to perform at their best.

At Tower, employees will find a stimulating, results-oriented environment where highly intelligent and motivated colleagues inspire each other to reach their greatest potential.

Summary

Trading and research at the firm run around the clock and across the globe, and they run on infrastructure this team designs, builds, and operates.

As part of R&D, you will join the engineers responsible for the compute, storage, operating systems, and automation behind that work at serious scale: hundreds of petabytes of storage and large CPU and GPU clusters spanning thousands of nodes.

The role is broad by design. One week you might be shaping the architecture of a new AI cluster, the next profiling a training job that will not scale, the next writing automation that keeps the whole fleet healthy with minimal human intervention.

Responsibilities

• Design, deploy, and scale distributed GPU clusters, from hardware selection and network topology through to production operation.

• Track down performance bottlenecks across the full stack: compute, storage, network, and the seams between them.

• Partner with researchers to profile and benchmark GPU workloads, then turn the findings into measurable speedups.

• Build the automation that lets a small team operate thousands of nodes: provisioning, monitoring, diagnostics, and self-healing.

• Own infrastructure projects end to end, from scope and design through implementation and long-term support.

• Qualify new generations of hardware and software, and work directly with vendors to root-cause complex issues.

Qualifications

• 5+ years engineering large-scale Linux systems in HPC, AI, or distributed-infrastructure environments.

• Deep Linux fundamentals: installation, performance tuning, and debugging, down to the kernel when the problem calls for it.

• Hands-on troubleshooting of distributed GPU workloads, with a strong mental model of GPU performance.

• Working experience with GPUDirect RDMA. You understand how data moves between GPUs and the network, and what to check when it does not.

• Solid Python for automation and tooling, plus CUDA or C/C++ experience. You can read, profile, and debug GPU code, not just operate the clusters it runs on.

• Familiarity with configuration management tools such as Salt, Ansible, Puppet, or Chef.

• Comfort diagnosing problems that cross hardware, OS, and network boundaries rather than stopping at one layer.

• Clear communication. You will work daily with

GPU Systems Engineer at Tower Research Capital, New York | Yoinka