yoinka

AI Model Policy Trainer, Image Evaluation - Seattle Onsite

Handshake

Seattle, WAFull TimeMid$45 – $55/hr
Sign in to applyVerified 2h ago
Location
Seattle, WA
Employment
Full Time
Work model
On-Site
Level
Mid
Salary
$45 – $55/hr
Posted
2h ago

Skills

REST

About this role

About Handshake Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions. In 2025, we started Handshake AI and built the fastest-growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We've grown from $0 to ~$1B run rate and pay ~$60M to over 30K individuals every month. Why join Handshake now: Shape how every career evolves in the AI economy, at global scale, with impact your friends, family and peers can see and feel Partner hand-in-hand with world-class AI labs, Fortune 500 partners and the world's top educational institutions Work together with engineers, scientists, operators, and more from Palantir, Meta, Scale AI, and former YC founders Build a massive, fast-growing business with billions in revenue About Handshake AI Human data is the core infrastructure to AI advancement. Frontier AI labs currently improve model capabilities with various data-intensive post-training techniques. We believe that data spend for AI training will increase by 3-5x in the next few years and continue for much longer as models take on new domains. Handshake AI supports all of the frontier AI labs, working on their most complex data at the largest scale.

About the Role

As an AI Image Evaluator, you will help image generation models learn two things at once: what a good image is, and what an acceptable image is. You will look at prompts and the images a model produced from them, then answer questions like: Did the image do what the prompt asked? Is it well made, or does it have the hands, lighting, text, and anatomy problems that give generated images away? Which of two images is better, and why? Does the image violate the customer's content policy, and if so, which category and how severely? Does it depict a real person, a protected brand, or a minor in a way the policy does not allow? The interesting cases are the close ones. Two images that look nearly identical until you notice one has a logo in the background. A stylized nude that is fine as figure study and not fine with one change of pose. A prompt that asked for "a realistic photo of a senator" and a model that complied. A beautiful image that ignored half the prompt, next to an ugly one that nailed it. We are looking for people who already see images critically, whether that came from photography, illustration, design, years inside Midjourney and Stable Diffusion, or moderating visual content at scale. You do not need all of these. You need one deep, and the judgment to learn the rest. This is not rote annotation. Rubrics cannot anticipate every image, and good evaluators do not apply them mechanically. You will balance the rubric's text and intent with customer expectations, precedent, and team calibration, and you will explain your reasoning clearly enough that it can train a model.

What You Will Do

Evaluate generated images against their prompts for adherence, composition, realism, style consistency, and technical defects Compare images side by side and select the stronger one with a clear, evidence-based rationale Classify images against customer content policies covering sexual content, violence, hate symbols, real-person likeness, intellectual property, and depictions of minors Select the most defensible classification when an image is genuinely ambiguous, and write concise rationales that cite rubric language and specific visual details Distinguish "I do not like this" from "this fails the prompt" from "this violates policy," and keep those judgments separate Write and refine prompts that probe where a model's quality or safety behavior breaks down Identify rubric gaps, contradictions, and emerging edge cases, and raise them with project leads and policy teams Participate actively in

Listing verified 2h ago. Applications go through the company's official careers site.

← Back to Yoinka

AI Model Policy Trainer, Image Evaluation - Seattle Onsite at Handshake, Seattle, WA | Yoinka