AI Safety Expert - Red Team
Overview
Lead red team efforts to improve conversational AI models and agents against jailbreaks and bias. Build human data sets and apply taxonomies to document findings. Deliver actionable insights through reports and attack case studies. This role differs by focusing on creative probing techniques and cross‑functional collaboration.
What You'll Do5
- 1Design and execute red team strategies for conversational AI models and agents to detect jailbreaks and bias exploitation
- 2Annotate failure modes generate high quality human data classify vulnerabilities and flag systemic risks
- 3Structure testing using taxonomies benchmarks and playbooks for consistent results
- 4Produce reproducible documentation including reports datasets and attack cases for customer use
- 5Work independently and asynchronously to meet deadlines while enhancing model performance
Requirements6
- 15+ years building ETL pipelines with Spark and Airflow
- 2Fluent in English and Swedish
- 3Prior red teaming experience in AI adversarial work cybersecurity or socio technical probing
- 4Strong communication skills to explain risks to technical and non technical stakeholders
- 5Preferred experience with Adversarial ML Cybersecurity and Socio technical risk
- 6Skills in Creative probing such as psychology acting or unconventional adversarial thinking
Salary Insight
$48 - $62k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.
Mercor
VerifiedAI Red Team Specialist, Bilingual English Swedish
You will red team conversational AI models to identify jailbreaks, prompt injections, and bias exploitation. Collaborate with a team of elite security and AI researchers at Mercor, backed by Benchmark and General Catalyst. This remote role pays $48-$62/hr and values independent, asynchronous work. You will shape safer AI by producing actionable reports and datasets for leading labs.
Mercor
VerifiedAI Safety Expert - Red Teaming
You will red team conversational AI models and agents for Mercor, an AI talent platform backed by Benchmark and General Catalyst. Working remotely and asynchronously, you will identify jailbreaks, prompt injections, and bias exploitation in English and Danish. You will generate high-quality human data by annotating failures and classifying vulnerabilities. Your reports and attack cases will directly improve model performance for leading AI research labs.
Mercor
VerifiedAI Safety Expert, Red Team & Adversarial ML
As an AI Safety Expert on the red team, you will probe conversational AI models and agents to uncover jailbreaks, prompt injections, and bias exploits at scale. You will join Mercor, a San Francisco-based talent network backed by Benchmark, General Catalyst, and Peter Thiel, working remotely with a team of elite technical and creative professionals. Your findings will directly shape safer AI deployments for leading research labs. This contract role offers $48–$62/hour and requires fluency in English and Finnish.

Mercor
VerifiedAI Safety Experts — English & Swedish | $48-$62/hr Remote
This remote, hourly contract is for bilingual (English & Swedish) experts who want to make AI systems safer by attacking them first. You'll red team conversational AI models using adversarial techniques like jailbreaks, prompt injections, and bias exploitation, then turn your findings into structured, reproducible reports. The work is text-based, with optional higher-sensitivity projects supported by clear guidelines and wellness resources. Pay ranges from $48 to $62 per hour.

Mercor
VerifiedAI Safety Experts — English & Danish | $48-$62/hr Remote
This remote role puts your adversarial skills to work making AI safer. You'll stress-test conversational AI models by probing for vulnerabilities like jailbreaks, prompt injections, and bias exploits. Your findings become the data that helps customers harden their systems. We're looking for native-level fluency in English and Danish, plus a background in red teaming, cybersecurity, or socio-technical risk.

Mercor
VerifiedAI Safety Experts — English & Finnish | $48-$62/hr Remote
Mercor is assembling a hand-picked red team of bilingual AI safety experts to stress-test conversational AI models in English and Finnish. This remote, hourly role focuses on probing models for vulnerabilities like jailbreaks, prompt injections, and bias exploits, then turning those discoveries into structured data that helps customers harden their AI systems. You'll follow established taxonomies and playbooks, with the option to skip higher-sensitivity projects. It's a unique chance to apply adversarial thinking at the frontier of AI safety while earning a competitive hourly rate.