MercorVerified Source
Remote

AI Safety Expert - Red Team

48K–62K
Remote · New York, New York
Posted August 12, 2026
contract

Overview

Lead red team efforts to improve conversational AI models and agents against jailbreaks and bias. Build human data sets and apply taxonomies to document findings. Deliver actionable insights through reports and attack case studies. This role differs by focusing on creative probing techniques and cross‑functional collaboration.

What You'll Do5

  • 1Design and execute red team strategies for conversational AI models and agents to detect jailbreaks and bias exploitation
  • 2Annotate failure modes generate high quality human data classify vulnerabilities and flag systemic risks
  • 3Structure testing using taxonomies benchmarks and playbooks for consistent results
  • 4Produce reproducible documentation including reports datasets and attack cases for customer use
  • 5Work independently and asynchronously to meet deadlines while enhancing model performance

Requirements6

  • 15+ years building ETL pipelines with Spark and Airflow
  • 2Fluent in English and Swedish
  • 3Prior red teaming experience in AI adversarial work cybersecurity or socio technical probing
  • 4Strong communication skills to explain risks to technical and non technical stakeholders
  • 5Preferred experience with Adversarial ML Cybersecurity and Socio technical risk
  • 6Skills in Creative probing such as psychology acting or unconventional adversarial thinking

Salary Insight

$48 - $62k per year

Location

Typeremote
LocationNew York, New York
This is a remote position

Required Skills

EnglishSwedishRed teamingCybersecuritySocio-technical probingCommunication skillsAdversarial MLCreative probing
Share:

Similar open positions

Explore active roles that match your skills and interests.

Mercor

8h agoRemotepayroll

AI Red Team Specialist, Bilingual English Swedish

You will red team conversational AI models to identify jailbreaks, prompt injections, and bias exploitation. Collaborate with a team of elite security and AI researchers at Mercor, backed by Benchmark and General Catalyst. This remote role pays $48-$62/hr and values independent, asynchronous work. You will shape safer AI by producing actionable reports and datasets for leading labs.

Competitive salary
adversarial mlcybersecurityred teaming+2 more

Mercor

8h agoRemotepayroll

AI Safety Expert - Red Teaming

You will red team conversational AI models and agents for Mercor, an AI talent platform backed by Benchmark and General Catalyst. Working remotely and asynchronously, you will identify jailbreaks, prompt injections, and bias exploitation in English and Danish. You will generate high-quality human data by annotating failures and classifying vulnerabilities. Your reports and attack cases will directly improve model performance for leading AI research labs.

Competitive salary
aicybersecurityred teaming+2 more

Mercor

10h agoRemotepayroll

AI Safety Expert, Red Team & Adversarial ML

As an AI Safety Expert on the red team, you will probe conversational AI models and agents to uncover jailbreaks, prompt injections, and bias exploits at scale. You will join Mercor, a San Francisco-based talent network backed by Benchmark, General Catalyst, and Peter Thiel, working remotely with a team of elite technical and creative professionals. Your findings will directly shape safer AI deployments for leading research labs. This contract role offers $48–$62/hour and requires fluency in English and Finnish.

Competitive salary
red teamingadversarial mlcybersecurity+2 more
Mercor

Mercor

11d agoRemotehourly

AI Safety Experts — English & Swedish | $48-$62/hr Remote

This remote, hourly contract is for bilingual (English & Swedish) experts who want to make AI systems safer by attacking them first. You'll red team conversational AI models using adversarial techniques like jailbreaks, prompt injections, and bias exploitation, then turn your findings into structured, reproducible reports. The work is text-based, with optional higher-sensitivity projects supported by clear guidelines and wellness resources. Pay ranges from $48 to $62 per hour.

48–62/hr
· 20 openings
red teamingadversarial machine learningjailbreak datasets+12 more
Mercor

Mercor

12d agoRemotehourly

AI Safety Experts — English & Danish | $48-$62/hr Remote

This remote role puts your adversarial skills to work making AI safer. You'll stress-test conversational AI models by probing for vulnerabilities like jailbreaks, prompt injections, and bias exploits. Your findings become the data that helps customers harden their systems. We're looking for native-level fluency in English and Danish, plus a background in red teaming, cybersecurity, or socio-technical risk.

48–62/hr
· 20 openings
red teamingprompt injectionjailbreaking+16 more
Mercor

Mercor

1d agoRemotehourly

AI Safety Experts — English & Finnish | $48-$62/hr Remote

Mercor is assembling a hand-picked red team of bilingual AI safety experts to stress-test conversational AI models in English and Finnish. This remote, hourly role focuses on probing models for vulnerabilities like jailbreaks, prompt injections, and bias exploits, then turning those discoveries into structured data that helps customers harden their AI systems. You'll follow established taxonomies and playbooks, with the option to skip higher-sensitivity projects. It's a unique chance to apply adversarial thinking at the frontier of AI safety while earning a competitive hourly rate.

48–62/hr
· 20 openings
prompt injectionjailbreakadversarial ml+14 more