Mercor
MercorVerified Source
Remote

Applied Health & Medicine Benchmark Specialist | $94-$119/hr Remote

94–119/hr
Remote
Posted August 14, 2026
hourly
21 openings

Overview

We’re looking for seasoned health science experts to create and review high quality assessment content for an AI research effort. You’ll craft and validate multiple choice questions across core medical domains, assess solution quality, and help set gold standard benchmarks to push AI capabilities forward. Work is remote and on an hourly basis, with flexible, asynchronous collaboration. The role blends subject matter expertise with rigorous evaluation to strengthen AI benchmarks in medicine and health care.

What You'll Do7

  • 1Author original, conceptually deep health and medicine MCQs and rate their difficulty
  • 2Review prewritten questions for accuracy, clarity, and rigor and implement edits
  • 3Rate question difficulty levels and ensure all necessary context is included in each prompt
  • 4Provide a correct answer plus 9 plausible distractors to challenge advanced solvers
  • 5Draft clear step-by-step reasoning explanations and document any changes for verification
  • 6Suggest 1–5 credible references per question from reputable medical literature
  • 7Flag issues related to clarity, completeness, precision, or solvability and justify edits during verification

Requirements5

  • 1MD, DO, PhD, or doctoral candidate in Medicine or related field
  • 2Strong background in graduate level medical knowledge and clinical reasoning
  • 3Experience with biomedical research methods and writing for scholarly audiences
  • 4Excellent written English and ability to convey complex ideas concisely
  • 5Board certification, clinical experience, or research publications in health fields is a plus

Who Should Apply

Ideal candidates are medical or biomedical PhDs or clinicians with a passion for education and AI, who can craft precise questions, verify content quickly, and communicate complex concepts clearly in writing while maintaining rigorous standards.

Salary Insight

Pay range is $94.00 to $119.00 per hour and work is fully remote on a flexible, asynchronous schedule.

Location

Typeremote
LocationRemote
This is a remote position

Required Skills

md / do / phd or doctoral candidatemedical knowledge across core domainsresearch methodology and academic writingclinical reasoning and problem solvingenglish communication and editing

Application Tip

Prepare a concise portfolio snippet showing 2 examples: one original MCQ set and one verification edit, with brief justification for difficulty rating and references.

Share:

Similar open positions

Explore active roles that match your skills and interests.

Mercor

Mercor

4h agoRemotehourly

Applied Biology Benchmark Specialist | $60-$75/hr Remote

One sentence on its own. We’re looking for seasoned biologists to craft and review high quality academic assessment content for an AI research project. You’ll develop and verify rigorous biology multiple-choice items, assess solution quality, and help set gold-standard benchmarks that push AI capabilities forward. The role spans authoring original questions and validating existing ones across several biology domains using a remote, asynchronous setup and competitive hourly pay.

60–75/hr
· 17 openings
biologymolecular biologybiochemistry+4 more
Mercor

Mercor

4h agoRemotehourly

Applied Computer Science Benchmark Specialist | $66-$84/hr Remote

A remote, hourly role focused on building and reviewing high quality AI research benchmark content for computer science. You’ll craft and validate rigorous multiple choice questions across core CS topics, assess answer quality, and contribute to gold-standard datasets that push AI capabilities forward. Expect a split between creating new questions and verifying existing ones, with opportunities to shape benchmarks used by researchers worldwide.

66–84/hr
· 27 openings
phd or doctoral candidatecomputer science theoryalgorithms+3 more
Mercor

Mercor

4h agoRemotehourly

Applied Mathematics Benchmark Specialist | $61-$77/hr Remote

A remote role focused on creating and validating rigorous math assessment content for an AI research initiative. You’ll write and review advanced multiple choice questions across key mathematics domains, assess solution quality, and help set benchmarks that push AI capabilities forward. The work blends mathematical depth with clear, instructional problem design and documentation. This position stands out for its blend of academic rigor and practical benchmarking in AI research.

61–77/hr
· 19 openings
question designformal proof writingacademic referencing+2 more
Mercor

Mercor

4h agoRemotehourly

Applied Business & Commerce Benchmark Specialist | $77-$98/hr Remote

We are in search of seasoned business and commerce experts to craft and review high quality assessment content for an AI research program. You’ll develop and validate rigorous multiple choice items across core business areas, assess solution quality, and help define gold standard benchmarks that push AI capabilities forward. The role offers remote, hourly engagement with flexible scheduling and the chance to influence how AI understands business concepts.

77–98/hr
· 19 openings
logisticssupply chain managementoperations research+5 more
Mercor

Mercor

4h agoRemotehourly

Applied Engineering Benchmark Specialist | $61-$77/hr Remote

The role focuses on creating and validating rigorous academic assessment content for an AI research effort. You will author and review multiple-choice questions across core engineering topics, judge solution quality, and help build gold-standard benchmarks to push AI capabilities forward. This is a remote, hourly engagement, with tasks that blend technical evaluation and high-quality writing using a clear, structured approach.

61–77/hr
· 26 openings
semiconductor design & manufacturingcontrol sciencemechatronics+5 more
Mercor

Mercor

4h agoRemotehourly

Applied Psychology Benchmark Specialist | $50-$63/hr Remote

A remote role for seasoned psychologists to author and review rigorous assessment items for an AI research effort. You’ll craft original multiple-choice questions and validate existing items to help define benchmarks that advance AI understanding of psychology. The work spans areas like psychometrics, digital health, and consumer psychology, with both creation and verification responsibilities. This opportunity blends scholarly rigor with flexible, asynchronous engagement and compensation on an hourly basis.

50–63/hr
· 17 openings
psychometricsassessment designgraduate-level psychology+2 more