Mercor
MercorVerified Source
Remote

AI Safety Red Teamer | $70-$84/hr Remote

70–84/hr
Remote
Posted July 25, 2026
hourly
4 openings

Overview

This role focuses on stress-testing some of the most advanced AI systems in the world. As an AI Safety Red Teamer, you'll design tricky prompts, hunt for vulnerabilities, and push models to see how they handle dangerous or ambiguous topics. You'll work fully remotely, collaborating with researchers who care deeply about alignment and safety. If you enjoy breaking things to make them stronger, this is a great fit.

What You'll Do5

  • 1Craft adversarial prompts that challenge frontier AI models in unexpected ways.
  • 2Probe for jailbreaks, unsafe outputs, hallucinations, and cases where policies simply don't hold up.
  • 3Test model resilience across sensitive areas like misinformation, cybersecurity, biosecurity, fraud, and political content.
  • 4Document what you find and turn it into clear reports that feed safety benchmarks and red-teaming efforts.
  • 5Work closely with AI researchers to tighten alignment, improve robustness, and drive safety improvements.

Requirements4

  • 1A bachelor's degree (or higher) in fields like computer science, cybersecurity, journalism, communications, psychology, biology, chemistry, public policy, or a related discipline.
  • 2Five or more years of hands-on experience in AI safety, red teaming, trust & safety, cybersecurity, investigative journalism, life sciences, or adjacent areas.
  • 3Sharp analytical thinking, solid prompt design instincts, and strong written communication.
  • 4Proven experience designing adversarial prompts or evaluating frontier AI systems.

Who Should Apply

The ideal candidate is someone who loves finding edge cases and isn't afraid to poke at sensitive or uncomfortable topics. You probably have a background in one of the listed disciplines and a track record of thinking adversarially. Familiarity with RLHF, SFT, or alignment research is a plus. You care about making AI safer and enjoy collaborating with a team that values rigor and curiosity.

Salary Insight

$70-$84 per hour, depending on experience and scope. This is a remote, hourly engagement.

Location

Typeremote
LocationRemote
This is a remote position

Required Skills

ai safetyred teamingadversarial testingprompt engineeringjailbreak testingrlhfsftai alignmenttrust & safetycybersecuritybiosecuritymisinformationfraud detectionpolitical contentevaluation methodologiesfrontier ai models

Application Tip

In your cover letter, share a specific example of a red-teaming exercise or a prompt you designed that exposed a real vulnerability in an AI system. Explain what you learned and how you'd apply that approach here.

Share:

Similar open positions

Explore active roles that match your skills and interests.

Mercor

11h agoRemotecontract

AI Safety Red Teamer Mercor

Mercor seeks an AI Safety Red Teamer to conduct adversarial evaluations of frontier AI models. This fully remote contract role offers competitive compensation up to $84/hour. You will design adversarial prompts identify jailbreaks evaluate model robustness and document vulnerabilities. The ideal candidate has strong analytical reasoning and experience in AI safety red teaming.

70K–84K
Adversarial prompt designAI safety/red teamingAnalytical reasoning+5 more
Mercor

Mercor

20d agoRemotehourly

AI Safety Practitioner | $60-$70/hr Remote

AIUC is looking for experienced AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models on complex, policy-sensitive topics. In this role, you'll review AI-generated responses for factual accuracy, policy compliance, and overall quality, helping improve how these models handle ambiguous and high-risk grey areas. You'll apply structured evaluation rubrics and provide feedback that directly shapes model behavior. This is a fully remote, hourly freelance position that's ideal for professionals from journalism, policy, or scientific backgrounds who want to apply their expertise to AI alignment.

60–70/hr
· 4 openings
ai safetyrlhfsft+13 more

Mercor

20h agoRemotecontract

AI Safety Expert - Red Team

Lead red team efforts to improve conversational AI models and agents against jailbreaks and bias. Build human data sets and apply taxonomies to document findings. Deliver actionable insights through reports and attack case studies. This role differs by focusing on creative probing techniques and cross‑functional collaboration.

48K–62K
EnglishSwedishRed teaming+5 more

Mercor

14h agoRemotepayroll

AI Safety Specialist (English & Dutch)

You will conduct adversarial testing on conversational AI systems, targeting jailbreaks, prompt injections, and misuse scenarios. Your findings shape model updates at leading research labs backed by Benchmark and Peter Thiel. Working fully remote, you'll collaborate asynchronously with a global team while owning the quality and consistency of security evaluations. This role emphasizes creative probing techniques, not just technical expertise. Fluent English & Dutch are required.

Competitive salary
pythonadversarial machine learningcybersecurity+1 more
Mercor

Mercor

3d agoRemotehourly

AI Safety Experts — English & Malay | $17-$25/hr Remote

Are you a bilingual English-Malay speaker with a knack for breaking things? We're building a red team of human experts to probe conversational AI models with adversarial inputs — jailbreaks, prompt injections, and tricky multi-turn conversations that expose hidden vulnerabilities. Your work will generate critical red-team data that helps our customers make their AI safer, more robust, and more trustworthy. This is a remote, text-based contract role paying $17-$25/hr, with optional exposure to sensitive topics supported by clear guidelines and wellness resources. What sets this apart: you're not just testing AI — you're shaping how the next generation of models handle bias, misinformation, and harmful behaviors.

17–25/hr
· 20 openings
red teamingadversarial machine learningjailbreak+14 more
Mercor

Mercor

11d agoRemotehourly

AI Safety Experts — English & Indonesian | $17-$25/hr Remote

We're hiring bilingual AI safety experts to stress-test conversational AI systems. In this remote contract role, you'll probe AI models for vulnerabilities like jailbreaks, prompt injections, and bias exploits, producing data that makes AI safer. You'll work with a red team that simulates real-world adversarial attacks, using structured playbooks and taxonomies. What sets this apart: you'll shape the safety of frontier AI products while earning $17–$25 per hour.

17–25/hr
· 10 openings
prompt injectionjailbreakred teaming+11 more