Senior Production Engineer
Clear Street
- Location
- London, UK
- Work model
- On-Site
- Level
- Senior
- Posted
- 1h ago
Skills
About this role
About Clear Street
Clear Street’s mission is to give every sophisticated investor access to every asset, in every market, through a unified platform built for speed, transparency and scale.
We give our clients the technology, tools, and service once reserved for the largest institutions, rebuilt with modern infrastructure. Our single, cloud-native, end-to-end capital markets platform powers investor growth today and is transforming how they can interact with markets tomorrow.
For more information, visit https://clearstreet.io.
The Role
As a Production Engineer, you sit at the intersection of software reliability and operational excellence. You own the health, resilience, and recovery of our production systems—while spending equal energy innovating solutions that eliminate human toil, reduce incident blast radius, and raise the reliability bar across the entire platform. You will partner closely with engineering, operations, and business teams to understand daily pain points and translate them into lasting automated solutions. Half your time is spent in the trenches—supporting production, responding to incidents, and deeply understanding how our systems behave under real conditions. The other half is yours to build: automation, tooling, and observability platforms that make tomorrow's on-call shift meaningfully easier than today's.
You will work on challenges like
● Design and build comprehensive monitoring and observability platforms that surface the right signal at the right time—eliminating alert fatigue and accelerating root-cause analysis. ● Develop intelligent automation and self-healing capabilities that diagnose issues, trigger recovery workflows, and reduce mean time to recovery (MTTR) without manual intervention. ● Analyze incidents, identify systemic trends, and engineer solutions that prevent entire classes of failures from recurring. ● Build reusable runbooks, diagnostic tooling, and recovery playbooks that turn tribal knowledge into scalable platform capabilities. ● Create golden-path operational workflows—making the safest, most reliable path also the easiest one for engineering teams to follow. ● Partner with Platform Engineering to influence CI/CD pipelines, deployment safety, and infrastructure resilience from a production reliability perspective. ● Champion Infrastructure as Code, GitOps, and SRE best practices while helping teams <span style="font-size: