yoinka

Technical Support Engineer (L2) - Compute

Mistral AI

ParisFull TimeSenior
Sign in to applyVerified 1h ago
Location
Paris
Employment
Full Time
Work model
On-Site
Level
Senior
Posted
1h ago

Skills

AWSAnsibleDockerGCPGoGrafanaKubernetesLinuxOCIPrometheusPythonShellTerraform

About this role

About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector, co-creating customized AI systems that they can run on their terms. We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited. Mistral AI is building a new Compute Support team to ensure the reliability, performance, and scalability of our GPU clusters in a Kubernetes environment. As one of the founding members of this team, you will play a pivotal role in shaping its processes, standards, and culture. In this hybrid L1/L2 support role, you will be the first point of contact for customers and internal teams, providing technical guidance, troubleshooting, and issue resolution for compute-related inquiries. You’ll also serve as the escalation point for complex issues, leveraging your deep systems knowledge, Kubernetes expertise, and operational debugging skills to ensure our AI workloads run smoothly at scale. This is a support-focused role that blends system administration, customer-facing communication, and compute infrastructure expertise. While coding is not the primary focus, familiarity with Go for debugging and automation is a plus.

Key Responsibilities

First Customer’s Point of Contact Act as the first interlocutor for customers and internal teams, providing timely and effective responses to compute-related inquiries. Triage and prioritize incoming requests, ensuring SLAs are met for acknowledgment, first response, and resolution. Gather and analyze initial issue details (e.g., logs, error messages, system metrics) to diagnose problems efficiently. Provide clear, actionable guidance to customers and internal users, including temporary workarounds where applicable.     Technical Support Serve as the L2 escalation point for Linux, Kubernetes, and compute infrastructure issues, providing deep technical troubleshooting for: GPU/TPU workloads (e.g., CUDA errors, memory leaks, job failures). Kubernetes clusters (e.g., pod crashes, node failures, networking misconfigurations). Bare metal and cloud environments (e.g., AWS EC2, GCP VMs, HPC clusters). Diagnose and resolve performance bottlenecks, hardware failures, and resource contention in distributed systems. Analyze system metrics, logs, and traces (e.g., dmesg, journalctl, nvidia-smi, Prometheus, Grafana) to identify root causes of issues. Participate in on-call rotations to provide 24/7 support for critical compute systems Nice to have: Optimize system configurations (e.g., kernel parameters, filesystem tuning, network settings) for high-performance computing (HPC) and AI workloads.     Kubernetes & Containerization Debug Kubernetes clusters with a focus on: Pod and node issues (e.g., CrashLoopBackOff, OOMKilled, ImagePullBackOff). Networking and storage (e.g., CNI plugins, PersistentVolumes, StorageClasses). Resource management (e.g., Requests/Limits, QOS classes, node affinity). Troubleshoot container runtime issues (Docker, containerd) such as image pull failures, OCI compliance, or runtime errors. Collaborate with SRE teams to understand and improve existing IaC (Terraform, Ansible) and Go-based tooling.     Documentation & Process Improvement Create and maintain runbooks, playbooks, and internal documentation for common compute issues (e.g., GPU debugging, Kubernetes troubleshooting). Contribute to post-mortems with actionable follow-ups to prevent recurring incidents. Train internal teams on best practices for compute infrastructure and debugging. Help define and refine support processes as a founding member of

Listing verified 1h ago. Applications go through the company's official careers site.

← Back to Yoinka

Technical Support Engineer (L2) - Compute at Mistral AI, Paris | Yoinka