Mercor
MercorVerified Source
Remote

AI Safety Experts — English & Malay | $17-$25/hr Remote

17–25/hr
Remote
Posted August 9, 2026
hourly
20 openings

Overview

Are you a bilingual English-Malay speaker with a knack for breaking things? We're building a red team of human experts to probe conversational AI models with adversarial inputs — jailbreaks, prompt injections, and tricky multi-turn conversations that expose hidden vulnerabilities. Your work will generate critical red-team data that helps our customers make their AI safer, more robust, and more trustworthy. This is a remote, text-based contract role paying $17-$25/hr, with optional exposure to sensitive topics supported by clear guidelines and wellness resources. What sets this apart: you're not just testing AI — you're shaping how the next generation of models handle bias, misinformation, and harmful behaviors.

What You'll Do4

  • 1Attack conversational AI systems using techniques like jailbreaks, prompt injections, and multi-turn manipulation to uncover misuse cases and bias exploitation.
  • 2Create high-quality human-annotated data by classifying model failures, flagging systemic risks, and documenting vulnerability patterns.
  • 3Apply structured testing frameworks — including taxonomies, benchmarks, and playbooks — to keep your red-team efforts consistent and measurable.
  • 4Write clear, reproducible reports and datasets that customers can use to strengthen their AI defenses and inform future safety improvements.

Requirements5

  • 1Native-level fluency in English and Malay (both written and spoken) is mandatory.
  • 2Hands-on experience with AI red teaming — whether that's adversarial ML work, cybersecurity probing, or socio-technical testing of AI systems.
  • 3A curious, adversarial mindset: you instinctively look for edge cases and love pushing systems to their breaking points.
  • 4Strong structure and communication skills — you use frameworks and benchmarks, not just random hacks, and you can explain risks clearly to both technical and non-technical stakeholders.
  • 5Adaptability to rotate across different projects and customers, picking up new testing playbooks quickly.

Who Should Apply

You're a bilingual (English/Malay) security-minded thinker who loves exploring the dark corners of AI. You have prior red teaming experience—maybe from adversarial ML, cybersecurity, or social risk analysis—and you enjoy structured, creative probing that goes beyond automated tests. You're comfortable writing detailed reports, working independently in a remote setting, and optionally engaging with sensitive content (which is always clearly flagged before exposure). If you want to turn your curiosity into real-world AI safety impact, this role is for you.

Salary Insight

Hourly compensation of $17.00–$25.00/hour, depending on experience and project.

Location

Typeremote
LocationRemote
This is a remote position

Required Skills

red teamingadversarial machine learningjailbreakprompt injectionrlhfdpomodel extractioncybersecuritypenetration testingexploit developmentreverse engineeringsocio-technical riskabuse analysisconversational aidata annotationenglishmalay

Application Tip

Stand out by including a concrete, specific example of a time you successfully jailbroke or manipulated an AI model — even a personal side project counts. Show the exact prompts you used and what you learned. That kind of proof speaks louder than any generic cover letter.

Share:

Similar open positions

Explore active roles that match your skills and interests.

Mercor

Mercor

11d agoRemotehourly

AI Safety Experts — English & Indonesian | $17-$25/hr Remote

We're hiring bilingual AI safety experts to stress-test conversational AI systems. In this remote contract role, you'll probe AI models for vulnerabilities like jailbreaks, prompt injections, and bias exploits, producing data that makes AI safer. You'll work with a red team that simulates real-world adversarial attacks, using structured playbooks and taxonomies. What sets this apart: you'll shape the safety of frontier AI products while earning $17–$25 per hour.

17–25/hr
· 10 openings
prompt injectionjailbreakred teaming+11 more
Mercor

Mercor

11d agoRemotehourly

AI Safety Experts — English & Thai | $24-$35/hr Remote

This remote role puts your linguistic and adversarial skills to work stress-testing conversational AI systems. You'll probe models for biases, injection flaws, and other hidden vulnerabilities, then document what you find so customers can build safer products. Fluency in both English and Thai is essential, and you'll have the option to engage with higher-sensitivity topics under clear guidelines. It's a chance to shape AI safety from the front lines, with pay starting at $24–$35 per hour.

24–35/hr
· 20 openings
red teamingprompt injectionjailbreak datasets+15 more
Mercor

Mercor

7d agoRemotehourly

AI Safety Experts — English & Vietnamese | $17-$25/hr Remote

Join a remote team of AI safety specialists tasked with stress-testing conversational AI systems before they reach the public. This contract role pairs native English and Vietnamese fluency with hands-on adversarial probing — jailbreaks, prompt injections, and bias exploitation — to uncover weaknesses automated checks miss. You'll generate structured red-team data, document reproducible attack cases, and help customers build safer, more trustworthy AI. Compensation is $17–$25 per hour, with all work text-based and optional exposure to sensitive topics supported by clear guidelines.

17–25/hr
· 10 openings
red teamingadversarial machine learningprompt injection+12 more
Mercor

Mercor

11d agoRemotehourly

AI Safety Experts — English & Swedish | $48-$62/hr Remote

This remote, hourly contract is for bilingual (English & Swedish) experts who want to make AI systems safer by attacking them first. You'll red team conversational AI models using adversarial techniques like jailbreaks, prompt injections, and bias exploitation, then turn your findings into structured, reproducible reports. The work is text-based, with optional higher-sensitivity projects supported by clear guidelines and wellness resources. Pay ranges from $48 to $62 per hour.

48–62/hr
· 20 openings
red teamingadversarial machine learningjailbreak datasets+12 more
Mercor

Mercor

1d agoRemotehourly

AI Safety Experts — English & Finnish | $48-$62/hr Remote

Mercor is assembling a hand-picked red team of bilingual AI safety experts to stress-test conversational AI models in English and Finnish. This remote, hourly role focuses on probing models for vulnerabilities like jailbreaks, prompt injections, and bias exploits, then turning those discoveries into structured data that helps customers harden their AI systems. You'll follow established taxonomies and playbooks, with the option to skip higher-sensitivity projects. It's a unique chance to apply adversarial thinking at the frontier of AI safety while earning a competitive hourly rate.

48–62/hr
· 20 openings
prompt injectionjailbreakadversarial ml+14 more
Mercor

Mercor

12d agoRemotehourly

AI Safety Experts — English & Danish | $48-$62/hr Remote

This remote role puts your adversarial skills to work making AI safer. You'll stress-test conversational AI models by probing for vulnerabilities like jailbreaks, prompt injections, and bias exploits. Your findings become the data that helps customers harden their systems. We're looking for native-level fluency in English and Danish, plus a background in red teaming, cybersecurity, or socio-technical risk.

48–62/hr
· 20 openings
red teamingprompt injectionjailbreaking+16 more