Handshake · Research · Unspecified · Posted 2026-09-02
AI Safety Policy Evaluator, Violence & Threats | Seattle Onsite
Handshake · Seattle, WA
Apply on Handshake's site Watch Handshake for new roles
This is a non-engineering content-policy evaluation role. Applicants must demonstrate relevant depth in violent fiction or media, military or emergency response, crisis or threat assessment, trust and safety, content moderation, or closely related policy work. Software engineering or LLM product experience alone is not sufficient.
ABOUT HANDSHAKE
Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.
Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise.
ROLE DETAILS
Location: Onsite in Seattle, WA, Monday-Friday.
Compensation: $55-$120/hr. Placement within the range depends on experience.
Employer: TCWGlobal. This is a W-2 assignment supporting Handshake AI.
Employment: Full time, 40 hours per week, non-exempt and eligible for overtime pay.
Schedule: Monday-Friday, 8 a.m.-5 p.m. PT.
Assignment: Ongoing. Planned start date: September 21, 2026.
ABOUT THE ROLE
As an AI Safety Policy Evaluator focused on Violence & Threats, you will help AI models learn where the line falls between depicting violence and enabling it.
Violence is one of the hardest domains in AI safety because most violent content is legitimate. Novels, games, screenplays, history, journalism, self-defense, and ordinary human frustration all involve violence, and a model that refuses them is broken. A model that helps someone plan real harm is worse. Your job is to tell the difference, case by case, and to explain your reasoning clearly enough that it can train a model.
You will read user requests, model responses, and conversation history, then decide which policy category applies and whether the model's response was appropriate. The interesting cases are the close ones: a torture scene that is either a chapter of a thriller or an interrogation manual with character names; a message that reads as venting about a boss or as a plan; a "realistic" combat question from a novelist that is also a real-world capability question. One word, one contextual detail, or one shift in intent changes the answer.
We are looking for people who already have strong instincts about violence in at least one of these areas: how it works in fiction, how it works in the real world, or how it shows up in people who are struggling. You do not need all three. You need one deep and the judgment to learn the rest.
This is not rote annotation. Policies cannot anticipate every edge case, and good evaluators do not apply them mechanically. You will balance policy text and intent with customer expectations, conversation context, precedent, and team calibration.
WHAT YOU WILL DO
- Evaluate user requests and AI model responses involving violence, weapons, threats, and dark fiction within the full conversation context; maintain accuracy and consistency across repeated evaluations
- Distinguish fictional, educational, historical, and defensive violence from requests that seek real-world uplift or express real intent to harm
- Assess whether a model's response gives meaningful real-world capability, regardless of how the request was framed
- Distinguish expressions of anger, frustration, or dark humor from credible threats or crisis indicators
- Select the most defensible classification when a case is genuinely ambiguous, and write concise rationales that cite policy language and conversation details while applying customer policy consistently
- Write and refine adversarial or borderline prompts that probe where a model draws the line
- Identify policy gaps, contradictions, and emerging edge cases, and raise them with project leads and policy teams
- Participate actively in calibration discussions; challenge interpretations respectfully and update your judgment when stronger reasoning em …
More research roles at Handshake
-
Member of Technical Staff, Data AI
ResearchUS$200k–350ktoday
-
Member of Technical Staff, Post-Training
ResearchRemote US UK & Ireland$200k–350k5d
-
Member of Technical Staff, Post-Training
ResearchRemote US$200k–350k7d
See also: AI jobs in Seattle · Handshake salaries · LLMs jobs · RLHF jobs.
This listing is reproduced from Handshake's public careers feed and links to the original. AI Hiring Index is not the employer and does not accept applications. All Handshake roles · AI salaries.