yoinka

Manager, AI Benchmarking and Evaluation Research (Remote, ROU)

CrowdStrike

RemoteRomania - RemoteMidH-1B sponsor company
Sign in to applyVerified 2h ago
Location
Romania - Remote
Work model
Remote
Level
Mid
H-1B history
35 approvals (FY2023)
Posted
Sep 8, 2026

Skills

CybersecurityLLM

About this role

As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed — we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We’re always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you.

About the Role

The CrowdStrike Data Science Team is looking for an experienced and driven leader to build and guide a team dedicated to designing and building evaluations for AI models that perform cybersecurity tasks. Our mission is to establish rigorous, reproducible standards for measuring how well AI and agentic systems support real-world security operations. As the Evaluation Lead, you will bring deep, hands-on SOC expertise to the table, defining what "good" looks like for AI models operating in security workflows. You will build the datasets and methodologies that ground our models in the realities of frontline defense, and lead a team that turns findings into actionable insights for engineering and product stakeholders.

What You'll Do

Lead, mentor and grow a team of researchers and engineers focused on evaluating AI models applied to cybersecurity tasks Define the strategy, roadmap and success metrics for evaluating AI/LLM and agentic systems across security use cases (e.g. incident response, threat hunting, alert triage, investigation) Design and standardize evaluation methodologies, benchmark datasets and reproducible testing pipelines rooted in real SOC workflows Assess models for accuracy, robustness, reliability and operational effectiveness in security-analyst scenarios Establish quantitative and qualitative metrics that reflect how AI models and agentic systems perform against genuine threat-detection and response challenges Collaborate cross-functionally with engineering, product and threat-research teams to translate evaluation findings into scalable improvements Communicate results, trade-offs and recommendations clearly to both technical and executive audiences What You'll Need: Hands-on SOC experience with a strong understanding of day-to-day security operations At least 3 years in a management or team-leadership position, with a proven track record of mentoring and growing technical teams (formal manager experience is a plus) Strong knowledge of incident response and threat hunting, including detection, investigation and remediation workflows Broad knowledge of the cybersecurity landscape, including attack vectors, defense mechanisms, and the analyst workflows that AI aims to augment Practical experience working with AI, including familiarity with AI/LLM capabilities and knowledge of agentic systems and how they operate Ability to define, design and standardize evaluation methodologies and reproducible testing pipelines Exceptional communication skills, including the ability to present complex technical findings clearly to both technical and non-technical audiences Demonstrated track record of delivering results, supported

Listing verified 2h ago. Applications go through the company's official careers site.

← Back to Yoinka

Manager, AI Benchmarking and Evaluation Research (Remote, ROU) at CrowdStrike, Romania - Remote | Yoinka