YO AI LabsVerified Source
Remote

AI Data Science Expert, Model Evaluation & Prompt Engineering

Remote · New York, New York
Posted August 12, 2026
contract

Overview

As an AI Data Science Domain Expert, you own the evaluation and refinement of AI-generated technical content, directly shaping the reasoning and accuracy of next-generation AI systems. You work remotely with a cross-functional team of data scientists, engineers, and product leads, delivering rubric-based assessments and structured feedback that drive model improvements. No prior AI experience is required; your expertise in data science, analytical thinking, and communication is what matters most. This contract role offers the chance to contribute to cutting-edge model development while honing skills in prompt engineering and RLHF.

What You'll Do7

  • 1Review and edit AI-generated content for accuracy, clarity, and technical relevance, focusing on data-driven insights and statistical claims.
  • 2Design and run prompt engineering experiments to optimize model outputs for specific analytical tasks.
  • 3Conduct rubric-based evaluations of AI responses, scoring for logical consistency, factual correctness, and completeness.
  • 4Verify technical information through independent research and fact-checking, citing credible sources where needed.
  • 5Annotate datasets with labels and feedback to support RLHF and model fine-tuning.
  • 6Interpret complex datasets and produce clear technical summaries that inform model development.
  • 7Collaborate with remote teams to refine evaluation metrics and improve AI workflows.

Requirements7

  • 13+ years in Data Science, Machine Learning, Applied AI, Statistics, Quantitative Analytics, or Data Analytics.
  • 2Proven experience producing or reviewing research papers, analytical reports, technical documentation, or experiment summaries.
  • 3Strong analytical reasoning and critical thinking skills, with attention to detail in evaluating complex information.
  • 4Excellent written communication skills to deliver structured feedback and technical documentation.
  • 5Experience with data annotation, content review, or rubric-based evaluation is preferred.
  • 6Familiarity with prompt engineering, AI output evaluation, fact-checking, or RLHF is a plus.
  • 7Advanced degree (Master's, MBA, PhD) preferred.

Salary Insight

Salary not disclosed in listing

Location

Typeremote
LocationNew York, New York
This is a remote position

Required Skills

data sciencemachine learningprompt engineeringrlhfdata annotation
Share:

Similar open positions

Explore active roles that match your skills and interests.

Mercor

Mercor

7d agoRemotehourly

Data Science Expert | $120-$170/hr Remote

This is a remote, hourly contract role for experienced data scientists who want to shape how AI systems are evaluated. You won't be building models or dashboards yourself — instead, you'll define what great data science work looks like and score AI-generated output against your criteria. The work sits at the intersection of data science and AI evaluation, making it a great fit for someone who enjoys judgment calls and rigorous reasoning. You'll collaborate with a leading AI research organization through Mercor, with pay at $120–$170 per hour.

120–170/hr
· 29 openings
sqlpythonexperiment design+10 more

Mercor

15h agoRemotepayroll

ML Research Expert, AI Evaluation

You will evaluate AI-generated content to strengthen reasoning and rigor in model outputs. You will review complex machine learning research for alignment with domain principles. You will provide structured feedback to AI teams to improve training data and downstream performance. You will collaborate with subject matter experts to ensure consistency across datasets. You will work independently and asynchronously to meet deadlines. Your published work at ICML, NeurIPS, or ICLR will anchor your credibility.

Competitive salary
machine learningreinforcement learningmeta-learning+2 more
Mercor

Mercor

20d agoRemotehourly

AI Safety Practitioner | $60-$70/hr Remote

AIUC is looking for experienced AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models on complex, policy-sensitive topics. In this role, you'll review AI-generated responses for factual accuracy, policy compliance, and overall quality, helping improve how these models handle ambiguous and high-risk grey areas. You'll apply structured evaluation rubrics and provide feedback that directly shapes model behavior. This is a fully remote, hourly freelance position that's ideal for professionals from journalism, policy, or scientific backgrounds who want to apply their expertise to AI alignment.

60–70/hr
· 4 openings
ai safetyrlhfsft+13 more

Mercor

17h agoRemotepayroll

AI Content Evaluator, Humanities & Culture

You will evaluate AI-generated artifacts for 5+ years of professional experience in Humanities, arts, or culture. You will identify factual, aesthetic, and presentation errors in documents, spreadsheets, and slide decks, and provide clear, structured written feedback. You will work independently and asynchronously with AI research teams to enhance training data quality. This role offers direct impact on AI model outputs at a leading AI research company, with flexible remote work and competitive compensation.

Competitive salary
microsoft officegoogle workspaceslides+1 more

Mercor

17h agoRemotepayroll

Data Evaluator, AI Quality & Tech Assessment

You will evaluate AI-generated artifacts against domain-specific quality rubrics, owning accuracy and presentation standards for documents, spreadsheets, and slide decks. Reporting to a decentralized team at Mercor, you'll work independently asynchronously with leading AI research labs. With 5+ years in Software, AI, or IT, you'll apply deep expertise to grade outputs with rigor. Contract pays $80–$120/hour remotely.

Competitive salary
microsoft officegoogle workspaceslides+2 more
Mercor

Mercor

20d agoRemotehourly

Data Science and Analytics Experts | $60-$70/hr Remote

This remote, hourly role puts your senior data science expertise to work building evaluation tasks that test AI systems against the realities of Fortune 500 enterprise data operations. You'll design realistic scenarios, craft reference outputs, and author scoring rubrics that capture how top analytics leaders think. If you've lived the complexity of enterprise pipelines, governance, and model deployment, this is a chance to shape AI evaluation at the highest level.

60–70/hr
· 25 openings
snowflakedatabrickspython+17 more