AI Red Teamer (LLM Generalist)

Handshake

2h ago 0 views 0 applications
Full-time On-site
Seattle, WA
$32 - $95
Full-time

Job Description

AI Red Teamer (LLM Generalist)Location: Seattle, WA (candidates must reside in the Seattle metro area or be willing to relocate prior to start)Type: Contract, 40 hours per weekAbout the RoleAs an AI Red Teamer, you will stress-test large language models by intentionally trying to break them. Rather than checking whether an answer is correct, you will design creative, adversarial prompts that expose vulnerabilities: unsafe content, bias, broken guardrails, hallucinations, prompt injection weaknesses, and unexpected behaviors. Your work directly supports AI safety and model robustness for leading research labs.This is a generalist red teaming role. You will probe models across the full spectrum of risk categories, including content safety, CBRN (chemical, biological, radiological, nuclear), cybersecurity, persuasion and influence operations, child safety, self-harm, over-companionship, and regulatory compliance. Red teaming may span text, image, voice, and agentic model capabilities depending on project needs.This role requires creativity, curiosity, and an ability to think like an adversary while operating with strong ethical judgment.Day-to-Day ResponsibilitiesCraft creative prompts and multi-turn scenarios to stress-test AI guardrails across diverse risk categoriesDiscover ways around safety filters, restrictions, and defenses using jailbreak, evasion, and prompt injection techniquesExplore edge cases to provoke disallowed, harmful, or incorrect outputsEvaluate and score model responses against structured harm taxonomies and severity rubricsDocument experiments clearly, including what you tried, why you tried it, and what it revealedReview and refine adversarial prompts generated by other team membersContribute to harm taxonomy development, calibration exercises, and inter-rater reliability workCollaborate with engineers, data scientists, and researchers to share findings and strengthen defensesWork with potentially disturbing content on a regular basis (see Content Warning below)Stay current on jailbreaks, attack methods, and evolving model behaviorsDesired CapabilitiesCoreStrong hands-on experience using multiple LLMs (ChatGPT, Claude, Gemini, open-source models, etc.)Intuition for crafting adversarial prompts; familiarity with jailbreak or evasion techniques is a strong plusCreative, adversarial problem-solving skillsClear and thoughtful written communicationStrong ethical judgment and the ability to separate adversarial thinking from personal valuesSelf-directed, collaborative, and comfortable in feedback-heavy environmentsCuriosity, persistence, and comfort with frequent failure in experimentationNice to HaveFamiliarity with Python or other scripting languagesExperience working with LLM APIs or evaluation toolingComfort with structured data annotation and rubric-based scoringPrior work in trust and safety, content moderation, QA, or security researchSubject matter expertise in any high-risk domain (cybersecurity, chemistry, biology, medicine, law, finance, etc.)You Will Thrive Here IfYou treat every model response as a hypothesis to challengeYou can switch between creative free-association and rigorous documentation in the same sessionYou go deep into unusual interests (fandoms, niche internet cultures, gaming exploits, Wikipedia rabbit holes, etc.)You come from a creative background: writing, visual art, improv, puzzle design, or similarYou are energized by finding the thing nobody else thought to tryYou are genuinely passionate about AI and follow the space closelyContent WarningThis role involves regular and deliberate exposure to harmful content. You will encounter and intentionally generate content involving violence, self-harm, hate speech, sexually explicit material, child safety scenarios, and other categories of harmful output as part of structured adversarial testing. Candidates must be able to engage with this material professionally and sustainably. Support resources are available.About Handshake AIHandshake AI partners with leading AI research labs to make models safer and more robust. Our red teaming operations help identify vulnerabilities before they reach users, contributing directly to the responsible development of frontier AI systems.