FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in stress-testing large language models and crafting adversarial prompts while ensuring compliance with safety standards. Proficient in evaluating model responses and collaborating with cross-functional teams to enhance AI safety and robustness.
Highest-signal resume keywords
Hands-On Experience With LLMsAdversarial Prompt CraftingFamiliarity With Jailbreak TechniquesExperience With LLM APIsSubject Matter Expertise In Cybersecurity
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Stress-TestingAdversarial Prompt DesignData AnnotationRubric-Based ScoringPython ScriptingEvaluation ToolingContent Safety AssessmentModel EvaluationExperiment DocumentationHarm Taxonomy Development
Soft Skills
Creative Problem-SolvingClear Written CommunicationStrong Ethical JudgmentSelf-Directed CollaborationCuriosity and Persistence
Industry Keywords
Content ModerationTrust and SafetyCybersecurityCBRNInfluence OperationsHigh-Risk Domain ExpertiseRegulatory ComplianceHarmful Material EngagementFeedback-Heavy EnvironmentsEvasion Techniques
Tech Stack
Tools & technologiesCyber SecurityPython
About the role
Key responsibilities & impact- Stress-test large language models by intentionally trying to break them
- Design creative, adversarial prompts exposing unsafe content, bias, broken guardrails, hallucinations, prompt injection weaknesses, and unexpected behaviors
- Probe models across content safety, CBRN, cybersecurity, persuasion and influence operations, child safety, self-harm, over-companionship, and regulatory compliance
- Test text, image, voice, and agentic model capabilities as project needs require
- Craft multi-turn scenarios to stress-test AI guardrails
- Discover ways around safety filters, restrictions, and defenses using jailbreak, evasion, and prompt injection techniques
- Explore edge cases to provoke disallowed, harmful, or incorrect outputs
- Evaluate and score model responses against structured harm taxonomies and severity rubrics
- Document experiments, including methods, rationale, and findings
- Review and refine adversarial prompts from other team members
- Contribute to harm taxonomy development, calibration exercises, and inter-rater reliability work
- Collaborate with engineers, data scientists, and researchers to share findings and strengthen defenses
- Work regularly with potentially disturbing content
- Stay current on jailbreaks, attack methods, and evolving model behaviors
- Support Handshake AI's work partnering with leading AI research labs to make models safer and more robust
Requirements
What you’ll need- Strong hands-on experience using multiple LLMs, including ChatGPT, Claude, Gemini, and open-source models
- Intuition for crafting adversarial prompts; familiarity with jailbreak or evasion techniques is a strong plus
- Creative, adversarial problem-solving skills
- Clear and thoughtful written communication
- Strong ethical judgment and ability to separate adversarial thinking from personal values
- Self-directed, collaborative, and comfortable in feedback-heavy environments
- Curiosity, persistence, and comfort with frequent failure in experimentation
- Candidates must be able to engage with harmful material professionally and sustainably
- Familiarity with Python or other scripting languages
- Experience working with LLM APIs or evaluation tooling
- Comfort with structured data annotation and rubric-based scoring
- Prior work in trust and safety, content moderation, QA, or security research
- Subject matter expertise in a high-risk domain such as cybersecurity, chemistry, biology, medicine, law, or finance
- Ability to work remotely from the United States, Monday through Friday, 40 hours per week
- Must be authorized to work lawfully in the United States for Handshake
- Must not require employment visa sponsorship, as addressed in the application requirements
Benefits
Comp & perks- Support resources are available for exposure to disturbing content
- Cash compensation range of $32–$95 per hour
- Remote work from the United States
- 40 hours per week
