
LLM Red-Teamer | $40-$65/hr Remote
Overview
A remote contractor role focused on stress-testing and improving frontier language models. You’ll craft challenging multi-turn prompts, build precise evaluation rubrics, and document findings to guide AI training. Your domain knowledge and clear written communication drive the quality of next‑gen AI systems, without needing prior AI experience. The work revolves around real-world input, rigorous evaluation, and independent delivery.
What You'll Do7
- 1Design intricate adversarial conversations and task-based scenarios aligned with project briefs
- 2Create unambiguous evaluation rubrics to measure model behavior against target criteria
- 3Iteratively push conversations and tasks against cutting-edge LLMs to reach defined quality thresholds
- 4Deliver complete task packs with transcripts, behavior targets, rubrics, and supporting evidence
- 5Assess LLM outputs, highlighting strengths and failure modes relative to the spec
- 6Coordinate with leads to stay aligned as project needs evolve
- 7Work autonomously to produce high‑quality deliverables at a steady pace
Requirements7
- 1Exceptional written English with clear structure and precision
- 2Experience in AI human data workflows (RLHF, SFT, evaluations, annotation, or prompt engineering) is a plus
- 3Strong familiarity with large language models and common failure patterns
- 4Ability to understand complex specs and translate them into actionable tasks without heavy oversight
- 5Attention to detail and strong critical thinking
- 6Experience designing evaluation items or rubrics is advantageous
- 7Background in writing-intensive or analysis-focused fields is beneficial
Who Should Apply
Ideal candidates are self-driven individuals who excel at precise writing and problem solving. You bring a curious mindset about how LLMs think and perform, can interpret detailed project briefs, and work effectively without heavy supervision. Experience in AI data work is helpful, but a strong command of clear language and rigorous evaluation is what truly matters.
Salary Insight
Pay is task-based and dependent on output quality and completion, with minimum submission requirements. Rates fall within a higher hourly band for contractors, but exact figures vary by task and pace.
Location
Required Skills
Application Tip
Show a concrete example in your application: describe a complex prompt you would design for a frontier model and outline how you would evaluate its response using a clear rubric.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedLLM Red Team Specialist — Failure Modes & Edge Cases | $60-$90/hr Remote
This role puts you on the front lines of AI safety, working with a top-tier lab to stress-test frontier models and uncover where they fail. You'll design complex, multi-step tasks that probe coding, ML, and reasoning abilities, then transform the weaknesses you find into rigorous evaluation benchmarks. The work is fully remote, runs about 35 hours per week, and pays $60-$90 per hour as a W-2 employee through Cincinnatus LLC. What makes this unique is the tight feedback loop with researchers, giving you direct influence over how the next generation of models is measured and improved.

Micro1
VerifiedAI Evaluation Analyst | $20-$30/hr Remote
The AI Evaluation Analyst helps advance frontier language models from a remote setting. You’ll craft detailed, task-based conversations and evaluation rubrics, test them against leading LLMs, and produce clear, evidence-backed assets that guide model learning. This contract role centers on written clarity, multi-turn design, and meticulous analysis to shape how AI systems reason and respond.

Mercor
VerifiedLLM Research Scientist (Pre-training & Computer Vision & Adversarial Robustness) | $100-$120/hr Remote
This remote contract role is for an experienced machine learning researcher who wants to dig into empirical, open-ended research problems across both vision and language. You'll train and fine-tune deep learning models end-to-end, from image classifiers to open-weight LLMs, while working within strict compute and data budgets. The work focuses on making models genuinely robust — to adversarial inputs, tricky conversations, and real-world constraints — rather than just chasing benchmark numbers. If you enjoy hands-on experimentation and have a strong research track record, this is a chance to collaborate with leading AI scientists on high-impact projects.

Wise Skulls Corp.
VerifiedAI Engineer (LLM Agents & Data Engineering)
Lead design and delivery of AI solutions that scale across multiple platforms. Own the end-to-end pipeline from concept to production while driving innovation in large language models. This role shapes how our systems learn and adapt.

Micro1
VerifiedLinguistics Specialists | $40-$50/hr Remote
Linguistics Specialists wanted for a remote, contractor role. You’ll leverage language expertise to gauge AI-generated and human-written content, ensuring accuracy, clarity, and consistency. Work with detailed guidelines to drive improvements in next generation AI systems, without needing AI experience. This position emphasizes precise analysis, constructive feedback, and reliable, guideline-driven evaluation using a remote setup across the US and Western Europe.

Micro1
VerifiedEnglish Speaking Generalist | $20-$30/hr Remote
This remote contractor role focuses on helping train and assess next generation AI systems. You’ll craft diverse prompts, compare outputs from leading models, and provide clear, structured feedback in American English. Your domain knowledge and attention to detail drive model improvement, with no AI background required beyond your expertise in your field.