AI First DevOps Engineer
Tomorrow.io
- Location
- 21 Ha'Arba'a St., 20th floor, Tel Aviv
- Work model
- On-Site
- Level
- Mid
- Posted
- 55m ago
Skills
About this role
Tomorrow.io's Engineering department is focused on building life-changing software and products at scale, from infrastructure that handles massive amounts of data to outstanding customer-centric user experiences in B2B, B2C and B2D products that change billions of lives worldwide.
We're looking for a DevOps Engineer to power the reliability, security, and efficiency of the world's most impactful weather platform. You'll build self-service platforms that give developers and weather scientists true independence, weave AI into how we operate, and work side-by-side with R&D to push performance and scale further. You'll evolve our cloud infrastructure to match the pace of the business, hold the line on cost, and stay close to production through on-call. The people who thrive here bring a product mindset, take ownership without waiting to be asked, and leave the people and systems around them better than they found them.
As a DevOps Engineer at Tomorrow.io, you'll
Evolve and maintain our cloud infrastructure, delivery and observability - and the self-service platforms above them - so they serve our business strategy and let scientists and developers work on their own as we grow at scaleDevelop and adopt tools, with AI at the center, that make development and operations processes measurably more efficientPartner with the engineering teams behind the Tomorrow.io platform and its API to improve service performance, reliability, scale, security and costKeep production available to its SLOs - taking your turn on the on-call rotation and following each incident through to whatever stops it from happening againWork as part of the team: share what you know, learn from the people around you, and help us keep improving how we work together
What you bring
• 4+ years as a DevOps / Site Reliability Engineer, including 2+ years hands-on with a major cloud (GCP, AWS or Azure) and with IaC such as Terraform (Crossplane or another Kubernetes-native approach is a plus)
• Production Kubernetes depth - containerized deployments, scheduling, resource management and real troubleshooting - plus CI/CD pipelines you've built and owned end to end with modern tooling and advanced deployment methodologies
• Experience implementing and customizing monitoring and observability systems (Datadog, Grafana + Loki, Prometheus, ELK)
• Software engineering craft and the judgment to review, debug and improve what an agent produces
• Hands-on use of AI tooling in your engineering work - shaping and extending it (agents, rules, integrations), not only consuming it
• A product mindset and a deep sense of responsibility for service reliability: you build internal tools from developer feedback, measure success with clear metrics and KPIs, and treat production as yours to care forCuriosity, adaptability and a bias for action in a high-velocity, changing environment - with communication that connects infrastructure work to business impact and brings R&D stakeholders together around shared goals
Our stack: GCP (primary), Azure and AWS · Linux · Kubernetes on GKE and AKS with Helm, KEDA, Argo Workflows and External Secrets · Cloud Run and Cloud Functions · Terraform and Crossplane · GitHub Actions and ArgoCD (GitOps) · Datadog, Grafana + Loki · PostgreSQL/PostGIS, MongoDB Atlas, Redis · Pub/Sub · GCS and S3