Principal Product Manager(Work IQ)
Microsoft
- Location
- China, Jiangsu, Suzhou; China, Shanghai, Shanghai; China, Beijing, Beijing
- Work model
- On-Site
- Level
- Principal
- H-1B history
- 2,066 approvals (FY2023)
- Posted
- 2h ago
About this role
Overview
Work IQ is Microsoft 365 Copilot’s workplace intelligence layer , enabling AI agents and experiences to securely reason over enterprise data such as emails, meetings, documents, Teams messages, and people signals. Work IQ exposes this intelligence through agent-first APIs, including the Work IQ API, CLI, and Model Context Protocol (MCP) servers, while respecting Microsoft Graph permissions, tenant policies, and compliance boundaries. We, the M365 Core Asia Pacific Copilot extensibility engineering team, are looking for a Principal Product Manager for Work IQ whose primary product focus is evaluation — the load-bearing quality gate of the Work IQ agent developer lifecycle. Evaluation is where an agent earns the right to ship: it spans the committed eval suite and its correspondence to the success criteria declared at scoping, measurement across quality, accuracy, latency, cost, trust, DLP behavior, and source scoping, failure triage against the response playbook, and the advisory-board sign-off that governs what customers actually receive. You will own that gate end to end — including A/B testing with BizChat, the feedback loop from production signals back into the suite, and post-release drift detection — and you will make it fast, evidence-rich, and continuous rather than a milestone event. This role is deliberately an AI-native transformation role, and it maps directly to the two M365 core priorities. On Changing the Way We Work , you are expected to redesign the lifecycle itself rather than staff it: apply an agentic loop approach to the evaluation stage so that evidence assembly, failure classification, remediation drafting, and drift detection run zero-touch, while accountable judgment — quality-gate sign-off, accepting a known failure into production, tenant admin consent, and any irreversible or externally visible action — stays with a named human by design. On Leading with New AI Products , you will dogfood Work IQ’s own surfaces to build those loops, so the platform team experiences its activation, evaluation, and metering surfaces as a developer does, and feeds that evidence straight back into the roadmap. We are looking for a growth mindset and thought leadership measured in outcomes, not mechanism: human involvements per agent shipped trending toward zero for removable coordination toil, a rising acceptance rate for the loops’ recommendations, shrinking time from eval failure to resolution, a growing share of drift found by the platform before a customer reports it, and an absolute bar of no agent-caused false gate pass . You will set the point of view, publish it, and bring partner teams along. This is a platform PM role at the intersection of Copilot , agents , Graph , and enterprise trust . Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.
Responsibilities
Evaluation & Quality — Primary Focus Own the evaluation stage end to end: the eval suite and its baseline, measurement across quality, accuracy, latency, cost, trust, DLP behavior, and source scoping, failure triage, and the advisory-board quality gate. Make evaluation continuous rather than a milestone , keep each agent’s eval suite in correspondence with its declared success criteria, and run A/B testing with BizChat plus post-release drift detection. Protect the trust boundary: gate sign-off, accepting a known failure, and admin consent stay accountable human decisions , and every gate decision stays auditable. AI-Native Transformation & Thought Leadership Changing the Way We Work: apply an agentic loop approach to evaluation so evidence assembly, failure triage, and drift detection run zero-touch —