Infrastructure Engineer
Fractile
- Location
- London/Bristol
- Work model
- On-Site
- Level
- Mid
- Posted
- 1h ago
Skills
About this role
Infrastructure Engineer
About Fractile
Fractile was founded in 2022 on the bet that, eventually, the world’s most capable AI systems would be limited in their impact by the time taken to produce useful outputs. We bet everything on the logical conclusion: that the only way to truly unlock this latent value, to make speed viable at scale, was to radically re-invent the hardware that we run our frontier AI models on. Ever since, we have been building chips and systems that tackle this problem: how to efficiently generate output at thousands of tokens per second, while handling the complexity and capacity challenges of operating large models at very long contexts.
The workloads that push to the limits of the current frontier are already transformational; it is the technical and economic limits on inference speed that are constraining progress. The defining work of the 21st century will be marked by the engine of inference delivering immense and diffuse chains of intellectual inquiry, in drug discovery, in software engineering, in materials discovery, in any field where progress is driven by deep reasoning and intelligence to resolve complex problems.
We are seeking an Infrastructure Engineer to help build and maintain the compute, storage, and networking foundations that support our silicon development workloads, supporting an on-premise compute environment used for computationally-intensive engineering work, while growing your skills across Linux systems administration, HPC-style cluster management, networking, storage, and performance tuning. This is a great opportunity for someone with 2-3+ years of solid Linux experience who picks up new technical areas quickly, works methodically, and is comfortable operating independently, and who wants to develop deeper expertise in infrastructure, high-performance computing, and large-scale systems engineering.
Key Responsibilities
• Help maintain and troubleshoot a fleet of on-premise Linux (Rocky/RHEL-family) servers.
• Support the deployment and upkeep of on-premise compute infrastructure using infrastructure-as-code tooling (e.g. Ansible).
• Assist with monitoring and observability — setting up and maintaining tooling for resource utilisation, service health, and machine failures (e.g. Prometheus, node_exporter, Grafana, Zabbix).
• Help diagnose and resolve networking issues (DNS, VLANs, bonding, routing) and storage issues (network filesystems, capacity management, performance troubleshooting).
• Support day-to-day operation of a cluster compute/job scheduling environment (e.g. Slurm).
• Help with user and identity management tasks (e.g. FreeIPA/LDAP), including onboarding and access provisioning.
• Document infrastructure changes and contribute to runbooks and playbooks.
• Work with engineers across the organisation to understand and help resolve infrastructure-related bottlenecks.
Requirements
• 2-3+ years of solid, hands-on Linux system administration experience.
• Good working knowledge of monitoring solutions (e.g. Prometheus, Grafana, Zabbix, or similar).
• Solid networking fundamentals — TCP/IP, DNS, VLANs, routing, basic troubleshooting.
• Good understanding of storage concepts — network filesystems, parallel/distributed filesystems, disk/volume management, basic performance troubleshooting.
• Comfortable working from the command line and with scripting (Bash, Python, or similar) to automate routine tasks.
• A methodical, curious approach to troubleshooting, with the ability to pick up new tools and technical areas quickly.
• Comfortable working independently and taking ownership of problems through to