Senior Site Reliability Engineer, IaaS
Algolia
- Location
- Paris, France
- Work model
- On-Site
- Level
- Senior
- Posted
- 607d ago
Skills
About this role
At Algolia, we’re proud to be a pioneer and market leader in AI Search, empowering 18,000+ businesses to deliver blazing-fast, predictive search and browse experiences at internet scale. Every week, we power over 30 billion search requests — four times more than Microsoft Bing, Yahoo, Baidu, Yandex, and DuckDuckGo combined.
In 2021, we raised $150 million in Series D funding, quadrupling our valuation to $2.25 billion. This strong foundation enables us to keep investing in our market-leading platform and serving incredible customers like Under Armour, PetSmart, Stripe, Gymshark, and Walgreens.
The team
The Infrastructure as a Service team is at the center of one of Algolia’s most consequential engineering transformations.
For years, Algolia has operated a production fleet of approximately 4,000 bare-metal servers to deliver the reliability, low latency, and scalability that our customers expect. We are now building the foundations of a unified cloud and Kubernetes platform designed to support Algolia’s growth for years to come.
This is not a lift-and-shift project. It is an opportunity to rethink how Algolia provisions, secures, operates, observes, upgrades, and scales production infrastructure and to build it as a platform that engineers can safely consume, rather than a queue of manual requests.
The opportunity
As a Senior Site Reliability Engineer in IaaS, you will help shape the next generation of Algolia’s production infrastructure.
You will lead major parts of the Cloud Baseline and the reliable lifecycle capabilities that enable teams to operate and migrate workloads safely on a cloud-native platform. You will work across cloud foundations, Kubernetes, platform engineering, automation, reliability, and large-scale production operations.
This role is for an engineer who enjoys solving infrastructure problems where the answer must work not once, but hundreds or thousands of times: creating repeatable cloud environments, enabling a growing fleet of production clusters, reducing manual operations, and maintaining the reliability and cost efficiency our customers expect throughout the transition.
YOU WILL
• Lead the design and evolution of Cloud Baseline capabilities across cloud providers, including identity and access, networking, account structure, security, auditability, tagging, inventory, and cost visibility.
• Design and automate cloud infrastructure foundations that enable a growing fleet of production Kubernetes clusters.
• Lead complex infrastructure initiatives, such as cloud-environment standardisation, cluster lifecycle automation, upgrade strategies, or infrastructure-drift reduction.
• Ensure cloud and Kubernetes foundations, lifecycle operations, and operational guardrails are reliable and scalable enough to support large-scale workload migration without compromising customer experience.
• Treat the platform as a product: define clear interfaces, reusable modules, self-service workflows, documentation, and reliable operational standards for the engineers who consume it.
• Build automated guardrails for security, compliance, reliability, and safe change management, allowing teams to move faster without weakening production protections.
• Improve platform efficiency through capacity planning, rightsizing, autoscaling, resource governance, and clear cost visibility.
• Use automation and AI-assisted engineering tools where appropriate to improve fleet-scale analysis, infrastructure documentation, and safe, repeatable operational changes.
• Mentor engineers, share knowledge, and raise the quality of infrastructure design and operations across the team.
• Collaborate with Infrastructure, Security, FinOps, and engineering teams to align technical decisions and deliver