
AI Safety Experts — English & Dutch | $48-$62/hr Remote
Overview
Mercor is assembling a remote red team of AI safety experts who speak both English and Dutch fluently. Your job is to probe conversational AI models with adversarial inputs, uncover hidden vulnerabilities, and produce the human data needed to make AI safer. This hourly role puts you at the frontier of AI safety, working on text-based projects that tackle sensitive topics like bias and misinformation.
What You'll Do4
- 1Stress-test conversational AI models and agents using techniques like jailbreaks, prompt injections, misuse cases, and multi-turn manipulation.
- 2Generate high-quality human data by annotating model failures, categorizing vulnerabilities, and flagging systemic risks.
- 3Follow structured taxonomies, benchmarks, and playbooks to keep your testing consistent and repeatable.
- 4Document every attack and finding in reproducible reports, datasets, and case studies that customers can act on.
Requirements5
- 1Hands-on experience red teaming AI systems, whether from adversarial machine learning, cybersecurity, or socio-technical probing.
- 2A naturally adversarial mindset that enjoys pushing systems to their breaking points.
- 3Comfort using frameworks or benchmarks to guide testing, not just relying on random attempts.
- 4Strong ability to explain technical risks clearly to both technical and non-technical stakeholders.
- 5Flexibility to shift across projects and customer needs as the work evolves.
Who Should Apply
You're the type of person who looks at an AI system and immediately thinks about how to break it. You have prior red teaming or security experience, you're methodical in your approach, and you're comfortable handling sensitive content like bias or misinformation. You're also independent enough to work remotely and fluent in both English and Dutch.
Salary Insight
Hourly compensation ranges from $48.00 to $62.00 per hour.
Location
Required Skills
Application Tip
When you apply, include a specific example of a time you uncovered a vulnerability that automated tests missed — detail your approach, the tools you used, and how you documented the finding.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedAI Safety Experts — English & Finnish | $48-$62/hr Remote
Mercor is assembling a hand-picked red team of bilingual AI safety experts to stress-test conversational AI models in English and Finnish. This remote, hourly role focuses on probing models for vulnerabilities like jailbreaks, prompt injections, and bias exploits, then turning those discoveries into structured data that helps customers harden their AI systems. You'll follow established taxonomies and playbooks, with the option to skip higher-sensitivity projects. It's a unique chance to apply adversarial thinking at the frontier of AI safety while earning a competitive hourly rate.
Mercor
VerifiedAI Safety Specialist (English & Dutch)
You will conduct adversarial testing on conversational AI systems, targeting jailbreaks, prompt injections, and misuse scenarios. Your findings shape model updates at leading research labs backed by Benchmark and Peter Thiel. Working fully remote, you'll collaborate asynchronously with a global team while owning the quality and consistency of security evaluations. This role emphasizes creative probing techniques, not just technical expertise. Fluent English & Dutch are required.

Mercor
VerifiedAI Safety Experts — English & Indonesian | $17-$25/hr Remote
We're hiring bilingual AI safety experts to stress-test conversational AI systems. In this remote contract role, you'll probe AI models for vulnerabilities like jailbreaks, prompt injections, and bias exploits, producing data that makes AI safer. You'll work with a red team that simulates real-world adversarial attacks, using structured playbooks and taxonomies. What sets this apart: you'll shape the safety of frontier AI products while earning $17–$25 per hour.

Mercor
VerifiedAI Safety Experts — English & Norwegian | $48-$62/hr Remote
This remote contract role gives you a chance to stress-test conversational AI systems before they reach the public. You'll use your native-level English and Norwegian to probe models for security flaws, bias, and harmful outputs, then turn your findings into structured datasets and reports. You'll be working at the frontier of AI safety at Mercor, where the philosophy is that the safest AI is the one that has already been attacked. The pay is $48–$62 per hour, and you'll have the flexibility to work from anywhere.

Mercor
VerifiedAI Safety Experts — English & Danish | $48-$62/hr Remote
This remote role puts your adversarial skills to work making AI safer. You'll stress-test conversational AI models by probing for vulnerabilities like jailbreaks, prompt injections, and bias exploits. Your findings become the data that helps customers harden their systems. We're looking for native-level fluency in English and Danish, plus a background in red teaming, cybersecurity, or socio-technical risk.
Mercor
VerifiedAI Safety Expert - Red Teaming
You will red team conversational AI models and agents for Mercor, an AI talent platform backed by Benchmark and General Catalyst. Working remotely and asynchronously, you will identify jailbreaks, prompt injections, and bias exploitation in English and Danish. You will generate high-quality human data by annotating failures and classifying vulnerabilities. Your reports and attack cases will directly improve model performance for leading AI research labs.