
AI Safety Red Teamer | $70-$84/hr Remote
Overview
This role focuses on stress-testing some of the most advanced AI systems in the world. As an AI Safety Red Teamer, you'll design tricky prompts, hunt for vulnerabilities, and push models to see how they handle dangerous or ambiguous topics. You'll work fully remotely, collaborating with researchers who care deeply about alignment and safety. If you enjoy breaking things to make them stronger, this is a great fit.
What You'll Do5
- 1Craft adversarial prompts that challenge frontier AI models in unexpected ways.
- 2Probe for jailbreaks, unsafe outputs, hallucinations, and cases where policies simply don't hold up.
- 3Test model resilience across sensitive areas like misinformation, cybersecurity, biosecurity, fraud, and political content.
- 4Document what you find and turn it into clear reports that feed safety benchmarks and red-teaming efforts.
- 5Work closely with AI researchers to tighten alignment, improve robustness, and drive safety improvements.
Requirements4
- 1A bachelor's degree (or higher) in fields like computer science, cybersecurity, journalism, communications, psychology, biology, chemistry, public policy, or a related discipline.
- 2Five or more years of hands-on experience in AI safety, red teaming, trust & safety, cybersecurity, investigative journalism, life sciences, or adjacent areas.
- 3Sharp analytical thinking, solid prompt design instincts, and strong written communication.
- 4Proven experience designing adversarial prompts or evaluating frontier AI systems.
Who Should Apply
The ideal candidate is someone who loves finding edge cases and isn't afraid to poke at sensitive or uncomfortable topics. You probably have a background in one of the listed disciplines and a track record of thinking adversarially. Familiarity with RLHF, SFT, or alignment research is a plus. You care about making AI safer and enjoy collaborating with a team that values rigor and curiosity.
Salary Insight
$70-$84 per hour, depending on experience and scope. This is a remote, hourly engagement.
Location
Required Skills
Application Tip
In your cover letter, share a specific example of a red-teaming exercise or a prompt you designed that exposed a real vulnerability in an AI system. Explain what you learned and how you'd apply that approach here.
Similar open positions
Explore active roles that match your skills and interests.
Mercor
VerifiedAI Safety Red Teamer Mercor
Mercor seeks an AI Safety Red Teamer to conduct adversarial evaluations of frontier AI models. This fully remote contract role offers competitive compensation up to $84/hour. You will design adversarial prompts identify jailbreaks evaluate model robustness and document vulnerabilities. The ideal candidate has strong analytical reasoning and experience in AI safety red teaming.

Mercor
VerifiedAI Safety Practitioner | $60-$70/hr Remote
AIUC is looking for experienced AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models on complex, policy-sensitive topics. In this role, you'll review AI-generated responses for factual accuracy, policy compliance, and overall quality, helping improve how these models handle ambiguous and high-risk grey areas. You'll apply structured evaluation rubrics and provide feedback that directly shapes model behavior. This is a fully remote, hourly freelance position that's ideal for professionals from journalism, policy, or scientific backgrounds who want to apply their expertise to AI alignment.
Mercor
VerifiedAI Safety Expert - Red Team
Lead red team efforts to improve conversational AI models and agents against jailbreaks and bias. Build human data sets and apply taxonomies to document findings. Deliver actionable insights through reports and attack case studies. This role differs by focusing on creative probing techniques and cross‑functional collaboration.
Mercor
VerifiedAI Safety Specialist (English & Dutch)
You will conduct adversarial testing on conversational AI systems, targeting jailbreaks, prompt injections, and misuse scenarios. Your findings shape model updates at leading research labs backed by Benchmark and Peter Thiel. Working fully remote, you'll collaborate asynchronously with a global team while owning the quality and consistency of security evaluations. This role emphasizes creative probing techniques, not just technical expertise. Fluent English & Dutch are required.

Mercor
VerifiedAI Safety Experts — English & Malay | $17-$25/hr Remote
Are you a bilingual English-Malay speaker with a knack for breaking things? We're building a red team of human experts to probe conversational AI models with adversarial inputs — jailbreaks, prompt injections, and tricky multi-turn conversations that expose hidden vulnerabilities. Your work will generate critical red-team data that helps our customers make their AI safer, more robust, and more trustworthy. This is a remote, text-based contract role paying $17-$25/hr, with optional exposure to sensitive topics supported by clear guidelines and wellness resources. What sets this apart: you're not just testing AI — you're shaping how the next generation of models handle bias, misinformation, and harmful behaviors.

Mercor
VerifiedAI Safety Experts — English & Indonesian | $17-$25/hr Remote
We're hiring bilingual AI safety experts to stress-test conversational AI systems. In this remote contract role, you'll probe AI models for vulnerabilities like jailbreaks, prompt injections, and bias exploits, producing data that makes AI safer. You'll work with a red team that simulates real-world adversarial attacks, using structured playbooks and taxonomies. What sets this apart: you'll shape the safety of frontier AI products while earning $17–$25 per hour.