
AI Evaluation Analyst | $20-$30/hr Remote
Overview
The AI Evaluation Analyst helps advance frontier language models from a remote setting. You’ll craft detailed, task-based conversations and evaluation rubrics, test them against leading LLMs, and produce clear, evidence-backed assets that guide model learning. This contract role centers on written clarity, multi-turn design, and meticulous analysis to shape how AI systems reason and respond.
What You'll Do6
- 1Create thorough multi-turn dialogues and corresponding rubrics aligned to project needs
- 2Test drafts with frontier LLMs and refine them to meet quality and complexity targets
- 3Produce evaluation artifacts such as transcripts, behavior targets, binary rubrics, and supporting documentation
- 4Maintain strict alignment with evolving project specs while delivering high throughput and accuracy
- 5Validate outputs with leads and QA teams as guidelines update
- 6Work independently to consistently meet delivery milestones and task quotas
Requirements6
- 1Native-level written English with exceptional clarity and structure
- 2Experience with data annotation, RLHF, SFT, evaluation, or prompt engineering for AI systems
- 3Solid understanding of frontier LLM behaviors and common failure modes
- 4Ability to interpret and apply highly detailed specs with minimal supervision
- 5Strong critical thinking and analytical writing skills
- 6Experience creating evaluation items or rubrics and analyzing technology outputs
Who Should Apply
Ideal candidates are self-driven professionals who excel at turning complex instructions into precise, testable evaluation content. You should enjoy rigorous analysis, have a keen eye for detail, and bring hands-on experience with AI data workflows or prompt design. Remote workers who can manage their time and deliver consistently will fit well here.
Salary Insight
Pay is task-based and tied to project specifications; compensation varies with experience and workflow, with minimum submission requirements.
Location
Required Skills
Application Tip
Showcase a sample evaluation artifact or rubric you crafted for an AI task to demonstrate your ability to translate specs into actionable items.
Similar open positions
Explore active roles that match your skills and interests.

Micro1
VerifiedLLM Red-Teamer | $40-$65/hr Remote
A remote contractor role focused on stress-testing and improving frontier language models. You’ll craft challenging multi-turn prompts, build precise evaluation rubrics, and document findings to guide AI training. Your domain knowledge and clear written communication drive the quality of next‑gen AI systems, without needing prior AI experience. The work revolves around real-world input, rigorous evaluation, and independent delivery.
YO AI Labs
VerifiedAI Data Science Expert, Model Evaluation & Prompt Engineering
As an AI Data Science Domain Expert, you own the evaluation and refinement of AI-generated technical content, directly shaping the reasoning and accuracy of next-generation AI systems. You work remotely with a cross-functional team of data scientists, engineers, and product leads, delivering rubric-based assessments and structured feedback that drive model improvements. No prior AI experience is required; your expertise in data science, analytical thinking, and communication is what matters most. This contract role offers the chance to contribute to cutting-edge model development while honing skills in prompt engineering and RLHF.

Micro1
VerifiedAI Consulting Domain Expert | $100-$200/hr Remote
AI Consulting Domain Expert lets you shape how AI systems learn and perform. This remote contractor role focuses on evaluating and refining AI outputs, drafting technical analyses, and crafting prompts to guide large language models. You’ll rely on sharp domain knowledge and strong writing to influence real-world AI behavior without requiring prior AI experience. Your work stands out by combining rigorous analysis with practical, customer-facing communication.

Micro1
VerifiedEnglish Speaking Generalist | $20-$30/hr Remote
This remote contractor role focuses on helping train and assess next generation AI systems. You’ll craft diverse prompts, compare outputs from leading models, and provide clear, structured feedback in American English. Your domain knowledge and attention to detail drive model improvement, with no AI background required beyond your expertise in your field.

Micro1
VerifiedAI Software Engineering Domain Expert | $100-$200/hr Remote
An AI software engineering domain expert who works remotely on a part time contract. You’ll help shape how next generation AI systems learn and perform by refining technical content, prompts, and evaluation rubrics. Your domain knowledge will drive data-driven improvements, even if you’re new to AI itself. This role centers on high-quality engineering input, documentation, and rigorous review to guide model behavior.

Wise Skulls Corp.
VerifiedAI Engineer (LLM Agents & Data Engineering)
Lead design and delivery of AI solutions that scale across multiple platforms. Own the end-to-end pipeline from concept to production while driving innovation in large language models. This role shapes how our systems learn and adapt.