
AI Safety Experts — English & Swedish | $48-$62/hr Remote
Overview
This remote, hourly contract is for bilingual (English & Swedish) experts who want to make AI systems safer by attacking them first. You'll red team conversational AI models using adversarial techniques like jailbreaks, prompt injections, and bias exploitation, then turn your findings into structured, reproducible reports. The work is text-based, with optional higher-sensitivity projects supported by clear guidelines and wellness resources. Pay ranges from $48 to $62 per hour.
What You'll Do5
- 1Probe conversational AI models and agents for vulnerabilities using jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation tactics.
- 2Generate high-quality human data by annotating model failures, classifying vulnerability types, and flagging systemic risks.
- 3Follow established taxonomies, benchmarks, and playbooks to keep testing consistent and repeatable.
- 4Document your work reproducibly—write detailed reports, build datasets, and describe attack cases that customers can act on.
- 5Expand evaluation coverage by uncovering issues that automated tests miss and collaborating across projects.
Requirements6
- 1Prior hands-on experience in AI red teaming, adversarial machine learning, cybersecurity, or socio-technical probing.
- 2A naturally curious and adversarial mindset—you instinctively push systems to their breaking points.
- 3Ability to work in a structured way, applying frameworks or benchmarks rather than random hacking.
- 4Clear communication skills to explain risks to both technical and non-technical stakeholders.
- 5Comfortable adapting to multiple projects and customer environments.
- 6Native fluency in English and Swedish for both written and spoken work.
Who Should Apply
You're an experienced red teamer who thinks like an attacker—whether your background is in AI safety, cybersecurity, or creative adversarial probing. You're bilingual in English and Swedish, methodical in your testing, and able to translate technical vulnerabilities into clear, actionable findings. If you're comfortable with text-based work and thrive on evaluating models across different projects and customers, this contract is for you.
Salary Insight
Hourly rate between $48 and $62, based on experience and project scope.
Location
Required Skills
Application Tip
Include a concrete example of a vulnerability you discovered and how you documented it—showing your structured, reproducible approach will help you stand out.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedAI Safety Experts — English & Danish | $48-$62/hr Remote
This remote role puts your adversarial skills to work making AI safer. You'll stress-test conversational AI models by probing for vulnerabilities like jailbreaks, prompt injections, and bias exploits. Your findings become the data that helps customers harden their systems. We're looking for native-level fluency in English and Danish, plus a background in red teaming, cybersecurity, or socio-technical risk.

Mercor
VerifiedAI Safety Experts — English & Vietnamese | $17-$25/hr Remote
Join a remote team of AI safety specialists tasked with stress-testing conversational AI systems before they reach the public. This contract role pairs native English and Vietnamese fluency with hands-on adversarial probing — jailbreaks, prompt injections, and bias exploitation — to uncover weaknesses automated checks miss. You'll generate structured red-team data, document reproducible attack cases, and help customers build safer, more trustworthy AI. Compensation is $17–$25 per hour, with all work text-based and optional exposure to sensitive topics supported by clear guidelines.

Mercor
VerifiedAI Safety Experts — English & Thai | $24-$35/hr Remote
This remote role puts your linguistic and adversarial skills to work stress-testing conversational AI systems. You'll probe models for biases, injection flaws, and other hidden vulnerabilities, then document what you find so customers can build safer products. Fluency in both English and Thai is essential, and you'll have the option to engage with higher-sensitivity topics under clear guidelines. It's a chance to shape AI safety from the front lines, with pay starting at $24–$35 per hour.

Mercor
VerifiedAI Safety Experts — English & Finnish | $48-$62/hr Remote
Mercor is assembling a hand-picked red team of bilingual AI safety experts to stress-test conversational AI models in English and Finnish. This remote, hourly role focuses on probing models for vulnerabilities like jailbreaks, prompt injections, and bias exploits, then turning those discoveries into structured data that helps customers harden their AI systems. You'll follow established taxonomies and playbooks, with the option to skip higher-sensitivity projects. It's a unique chance to apply adversarial thinking at the frontier of AI safety while earning a competitive hourly rate.

Mercor
VerifiedAI Safety Experts — English & Malay | $17-$25/hr Remote
Are you a bilingual English-Malay speaker with a knack for breaking things? We're building a red team of human experts to probe conversational AI models with adversarial inputs — jailbreaks, prompt injections, and tricky multi-turn conversations that expose hidden vulnerabilities. Your work will generate critical red-team data that helps our customers make their AI safer, more robust, and more trustworthy. This is a remote, text-based contract role paying $17-$25/hr, with optional exposure to sensitive topics supported by clear guidelines and wellness resources. What sets this apart: you're not just testing AI — you're shaping how the next generation of models handle bias, misinformation, and harmful behaviors.

Mercor
VerifiedAI Safety Experts — English & Indonesian | $17-$25/hr Remote
We're hiring bilingual AI safety experts to stress-test conversational AI systems. In this remote contract role, you'll probe AI models for vulnerabilities like jailbreaks, prompt injections, and bias exploits, producing data that makes AI safer. You'll work with a red team that simulates real-world adversarial attacks, using structured playbooks and taxonomies. What sets this apart: you'll shape the safety of frontier AI products while earning $17–$25 per hour.