yoinka

Manager, AI Ops Site Reliability Engineer

Pfizer

Greece-Thessaloniki ChortiatisMidH-1B sponsor company
Sign in to applyVerified 2h ago
Location
Greece-Thessaloniki Chortiatis
Work model
On-Site
Level
Mid
H-1B history
9 approvals (FY2023)
Posted
Sep 7, 2026

Skills

AgileLLMPythonServiceNow

About this role

ROLE

SUMMARY The CISO Infrastructure & Cloud Services organization delivers excellence in the pursuit of breakthroughs that change patients' lives through industry-leading infrastructure operations performance. We ensure optimal performance of global application hosting and network services that power Pfizer's business processes. We strive to revolutionize service dependability by applying advanced analytics to drive predictive detection, identifying potential issues with our services and intervening before they disrupt our business. We place data at the heart of what we do and apply a relentless focus on continuous improvement to enable Pfizer's business processes and patient outcomes. We are seeking an Manager, AI Ops Site Reliability Engineer to own how our operations agents think and how well they perform, within Hosting SRE Operations, based in Greece. In this role, you design the agent workflows, prompts and context, build the evaluation harness that proves accuracy against our own estate, and partner with our existing domain experts to convert their operational know-how into reliable, measurable agent behavior. You bring depth in AI and evaluation; our domain specialists bring the deep OS, Storage, Data Protection, Database and HCI knowledge - together you turn classic operations into trustworthy AI operations.

ROLE

RESPONSIBILITIES Design agent workflows for intake and investigation: the prompts, context and retrieval, tool selection and decision logic that determine how the agent classifies, enriches and reasons about incidents. Build and own the evaluation harness: ground-truth datasets drawn from historical incidents, accuracy, precision and recall metrics, regression tests, and hallucination and error detection - so every change to an agent is measured before it ships. Define the reasoning guardrails: what an agent may conclude and propose, how confident it must be, and when it must defer to a human - the quality side of the approval gate. Turn tacit operational knowledge into machine-usable context: work with the existing domain experts to convert their runbooks and past investigations into retrievable knowledge the agent can cite. Run the transformation on Operational workflows: take a classic operational process, redesign it with the domain owner as an agent-assisted flow, pilot it, and measure toil and MTTR reduction. Continuously tune agent quality against production results, feeding misses back into prompts, context and the evaluation sets. Collaborate with SRE, Ops, Platform teams, AI Ops Platform Engineers, AI Ops Site Reliability Engineers, and other stakeholders to align operational requirements, platform capabilities, automation strategies, and execution priorities QUALIFICATIONS Bachelor's degree in a technical field or equivalent practical experience. 4 years+ in software, data or ML/AI engineering, including hands-on delivery of LLM or agent-based applications in production. Demonstrated prompt and context engineering, and experience building structured evaluations for AI systems (ground-truth sets, metrics, regression testing; LLM-as-judge a plus). Working knowledge of agent frameworks and tooling (for example Amazon Bedrock AgentCore or LangGraph) and RAG / knowledge-base design, with strong Python. IT operations literacy to design realistic operational workflows and judge the plausibility of agent output. You do not need to be a domain administrator - you will partner with domain experts who are. Strong analytical and communication skills to work with SREs, extract knowledge and present accuracy results credibly. Familiarity with ServiceNow, Dynatrace or observability data is a plus. Demonstrated experience in an agile work environment possessing qualities such as a collaborative mindset, adaptability to change, and a proactive problem-solving approach. NON-STANDARD WORK SCHEDULE, TRAVEL OR ENVIRONMENT REQUIREMENTS Periodic international and domestic travel required (less than 10%) Please

Listing verified 2h ago. Applications go through the company's official careers site.

← Back to Yoinka

Manager, AI Ops Site Reliability Engineer at Pfizer, Greece-Thessaloniki Chortiatis | Yoinka