
AI Safety Experts — English & Danish | $48-$62/hr Remote
Overview
This remote role puts your adversarial skills to work making AI safer. You'll stress-test conversational AI models by probing for vulnerabilities like jailbreaks, prompt injections, and bias exploits. Your findings become the data that helps customers harden their systems. We're looking for native-level fluency in English and Danish, plus a background in red teaming, cybersecurity, or socio-technical risk.
What You'll Do4
- 1Probe conversational AI systems with adversarial inputs to uncover weaknesses, including jailbreaks, prompt injections, and multi-turn manipulation.
- 2Generate high-quality human data by classifying failure modes, flagging systemic risks, and annotating model responses.
- 3Work within defined taxonomies, benchmarks, and playbooks to keep testing structured and repeatable.
- 4Write clear reports and build datasets that show customers exactly how an attack was executed and what to fix.
Requirements4
- 1Native fluency in both English and Danish — you'll be reviewing and producing content in both languages.
- 2Hands-on red teaming experience, whether through AI adversarial work, cybersecurity, or socio-technical probing.
- 3A structured approach: you're comfortable using frameworks or benchmarks rather than just trying random attacks.
- 4The ability to explain complex security risks to both technical and non-technical audiences.
Who Should Apply
You're the kind of person who instinctively looks for the cracks in a system. You're curious, persistent, and think like an adversary but can document your work with precision. You're excited to work at the frontier of AI safety, and you thrive when you're challenged across different projects and customers.
Salary Insight
$48–$62 per hour, paid for hourly engagement.
Location
Required Skills
Application Tip
Don't just say you're good at breaking things — show it. Include a specific example of a vulnerability you discovered, the method you used, and how you reported it. If you've ever written a jailbreak prompt or found a bias exploit, describe that.
Similar open positions
Explore active roles that match your skills and interests.
Mercor
VerifiedAI Safety Expert - Red Teaming
You will red team conversational AI models and agents for Mercor, an AI talent platform backed by Benchmark and General Catalyst. Working remotely and asynchronously, you will identify jailbreaks, prompt injections, and bias exploitation in English and Danish. You will generate high-quality human data by annotating failures and classifying vulnerabilities. Your reports and attack cases will directly improve model performance for leading AI research labs.

Mercor
VerifiedAI Safety Experts — English & Swedish | $48-$62/hr Remote
This remote, hourly contract is for bilingual (English & Swedish) experts who want to make AI systems safer by attacking them first. You'll red team conversational AI models using adversarial techniques like jailbreaks, prompt injections, and bias exploitation, then turn your findings into structured, reproducible reports. The work is text-based, with optional higher-sensitivity projects supported by clear guidelines and wellness resources. Pay ranges from $48 to $62 per hour.

Mercor
VerifiedAI Safety Experts — English & Indonesian | $17-$25/hr Remote
We're hiring bilingual AI safety experts to stress-test conversational AI systems. In this remote contract role, you'll probe AI models for vulnerabilities like jailbreaks, prompt injections, and bias exploits, producing data that makes AI safer. You'll work with a red team that simulates real-world adversarial attacks, using structured playbooks and taxonomies. What sets this apart: you'll shape the safety of frontier AI products while earning $17–$25 per hour.

Mercor
VerifiedAI Safety Experts — English & Finnish | $48-$62/hr Remote
Mercor is assembling a hand-picked red team of bilingual AI safety experts to stress-test conversational AI models in English and Finnish. This remote, hourly role focuses on probing models for vulnerabilities like jailbreaks, prompt injections, and bias exploits, then turning those discoveries into structured data that helps customers harden their AI systems. You'll follow established taxonomies and playbooks, with the option to skip higher-sensitivity projects. It's a unique chance to apply adversarial thinking at the frontier of AI safety while earning a competitive hourly rate.

Mercor
VerifiedAI Safety Experts — English & Dutch | $48-$62/hr Remote
Mercor is assembling a remote red team of AI safety experts who speak both English and Dutch fluently. Your job is to probe conversational AI models with adversarial inputs, uncover hidden vulnerabilities, and produce the human data needed to make AI safer. This hourly role puts you at the frontier of AI safety, working on text-based projects that tackle sensitive topics like bias and misinformation.

Mercor
VerifiedAI Safety Experts — English & Norwegian | $48-$62/hr Remote
This remote contract role gives you a chance to stress-test conversational AI systems before they reach the public. You'll use your native-level English and Norwegian to probe models for security flaws, bias, and harmful outputs, then turn your findings into structured datasets and reports. You'll be working at the frontier of AI safety at Mercor, where the philosophy is that the safest AI is the one that has already been attacked. The pay is $48–$62 per hour, and you'll have the flexibility to work from anywhere.