Data Scientist, Cybersecurity
OpenAI
- Location
- US - Remote
- Employment
- Full Time
- Work model
- Remote
- Level
- Senior
- Posted
- 1h ago
Skills
About this role
About the Team
OpenAI’s Agentic Data Science team helps shape how AI agents are built, deployed, and improved across our products. We partner with product, engineering, research, and security teams to define meaningful measures of success, understand how our systems behave in the real world, and translate evidence into better decisions. As AI agents become more capable, they can write and execute code, access sensitive systems, and complete increasingly complex tasks with greater autonomy. These capabilities create powerful opportunities to improve cybersecurity, but they also introduce risks that traditional security tools and processes were not designed to address. Meeting this moment requires new ways to measure security, evaluate defenses, and distinguish genuine risk reduction from friction that slows users down.
About the Role
We are looking for a senior data scientist to help define what effective cybersecurity looks like in the age of AI agents. You will work across OpenAI’s Security organization and cybersecurity product teams to measure emerging risks, improve internal security controls, and shape AI-powered security products. The problems are foundational: How do we know whether an agent’s security controls are effective? Which safeguards meaningfully reduce risk, and which create unnecessary friction? When an AI system identifies a potential vulnerability, how do we determine whether the finding is accurate, actionable, and ultimately resolved? How do we detect anomalous behavior or risky access when the systems themselves are changing rapidly? You will report into Data Science while partnering closely with Security, Cyber Product, Engineering, and Research. This is a high-ownership role for someone who can establish a new analytical discipline, operate across organizational boundaries, and turn ambiguous security challenges into measurable improvements. In This Role You Will Define how we measure AI-agent security. Establish metrics and evaluation frameworks for security-control coverage, agent behavior, sensitive actions, access patterns, detection quality, and emerging risks. Improve security controls without introducing unnecessary friction. Quantify the effectiveness and operational costs of safeguards, including false positives, blocked actions, escalations, approval delays, and recovery paths. Help teams make controls safer, more precise, and easier to use. Build the data foundations for security decisions. Partner with engineering and data teams to improve instrumentation, connect fragmented telemetry, establish trusted datasets, and surface important coverage and data-quality gaps. Strengthen detection and response. Identify meaningful signals of anomalous behavior, risky access, sensitive-data exposure, and other security-relevant activity. Evaluate whether interventions improve detection quality, response times, and real-world security outcomes. Shape AI-powered cybersecurity products. Partner with product, engineering, and research teams to assess how effectively AI systems identify security issues, support developer and enterprise workflows, and create measurable customer value. Develop evaluation systems for security findings. Define quality measures for findings, including accuracy, severity, actionability, duplication, resolution, and downstream impact. Connect model behavior and product changes to outcomes such as triage, remediation, and vulnerability reduction. Understand the complete security workflow. Measure how users discover, investigate, validate, prioritize, and resolve security issues. Identify opportunities to improve activation, adoption, retention, and enterprise value across customer-facing cybersecurity products. Design rigorous measurement and experimentation strategies. Evaluate new models, security controls, product features, and workflows through controlled experiments, staged rollouts, observational analyses, and other methods appropriate for high-stakes environments.