
Applied History & Political Science Benchmark Specialist | $44-$56/hr Remote
Overview
This remote role focuses on creating and validating academic assessment content for an AI research initiative. You’ll craft and review high quality multiple choice items in history and political science, help determine gold standard benchmarks, and contribute to improving AI understanding in these fields. The work blends scholarly rigor with practical QA to support advanced AI systems. Expect a flexible, asynchronous arrangement that taps deep expertise in your domains.
What You'll Do7
- 1Draft original, challenging multiple-choice questions that probe deep understanding of history and political science.
- 2Assess questions for clarity, self containment, and precise framing, ensuring all needed information is included in the prompt.
- 3Assign a difficulty level such as Medium, Hard, or Expert to each item and provide accurate rationale.
- 4Create one correct answer plus 9 plausible distractors that differentiate expert solvers.
- 5Develop step-by-step Chain-of-Thought reasoning in a clear, concise format for each item.
- 6Provide 1 to 5 scholarly references from reputable sources per question.
- 7For verification tasks, identify and justify any issues related to clarity, completeness, or solvability and document edits.
Requirements5
- 1PhD or current doctoral candidate in History, Political Science, International Relations, or a closely related field
- 2Master's degree considered for applicants with exceptional depth in a subfield
- 3Strong skill set in historiography, political theory, and comparative analysis
- 4Track record of research publications or policy-related experience is advantageous
- 5Excellent command of written English with the ability to convey complex ideas clearly
Who Should Apply
Ideal candidates are scholars with deep expertise in history or political science who enjoy turning complex concepts into precise assessment items. You should be comfortable evaluating questions for accuracy, crafting rigorous answer options, and documenting references, all in a remote, self paced setting.
Salary Insight
Pay range: $44.00 - $56.00/hr with remote, hourly engagement.
Location
Required Skills
Application Tip
Show concrete examples of past multiple choice or exam content you created, and briefly describe how you ensured question validity and fairness.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedApplied Philosophy Benchmark Specialist | $50-$63/hr Remote
A remote role focused on creating and vetting high quality philosophy assessment content for an AI research initiative. You will craft and critique multiple choice questions across core philosophy domains, while helping set gold standard benchmarks for advancing AI capabilities. The work blends deep philosophical expertise with rigorous evaluation, and offers flexible, asynchronous engagement.

Mercor
VerifiedApplied Computer Science Benchmark Specialist | $66-$84/hr Remote
A remote, hourly role focused on building and reviewing high quality AI research benchmark content for computer science. You’ll craft and validate rigorous multiple choice questions across core CS topics, assess answer quality, and contribute to gold-standard datasets that push AI capabilities forward. Expect a split between creating new questions and verifying existing ones, with opportunities to shape benchmarks used by researchers worldwide.

Mercor
VerifiedApplied Psychology Benchmark Specialist | $50-$63/hr Remote
A remote role for seasoned psychologists to author and review rigorous assessment items for an AI research effort. You’ll craft original multiple-choice questions and validate existing items to help define benchmarks that advance AI understanding of psychology. The work spans areas like psychometrics, digital health, and consumer psychology, with both creation and verification responsibilities. This opportunity blends scholarly rigor with flexible, asynchronous engagement and compensation on an hourly basis.

Mercor
VerifiedApplied Physics Benchmark Specialist | $61-$77/hr Remote
We are looking for seasoned physicists to craft and review advanced academic assessment content for an AI research project. You’ll develop and validate high quality multiple-choice questions across key physics areas, assess solution quality, and help establish benchmark materials that push AI capabilities forward. The role is remote and hourly, with flexible, asynchronous collaboration. You’ll work in two modes: authoring original questions and verifying pre-written items for accuracy and rigor.

Mercor
VerifiedApplied Mathematics Benchmark Specialist | $61-$77/hr Remote
A remote role focused on creating and validating rigorous math assessment content for an AI research initiative. You’ll write and review advanced multiple choice questions across key mathematics domains, assess solution quality, and help set benchmarks that push AI capabilities forward. The work blends mathematical depth with clear, instructional problem design and documentation. This position stands out for its blend of academic rigor and practical benchmarking in AI research.

Mercor
VerifiedApplied Engineering Benchmark Specialist | $61-$77/hr Remote
The role focuses on creating and validating rigorous academic assessment content for an AI research effort. You will author and review multiple-choice questions across core engineering topics, judge solution quality, and help build gold-standard benchmarks to push AI capabilities forward. This is a remote, hourly engagement, with tasks that blend technical evaluation and high-quality writing using a clear, structured approach.