
Applied Computer Science Benchmark Specialist | $66-$84/hr Remote
Overview
A remote, hourly role focused on building and reviewing high quality AI research benchmark content for computer science. You’ll craft and validate rigorous multiple choice questions across core CS topics, assess answer quality, and contribute to gold-standard datasets that push AI capabilities forward. Expect a split between creating new questions and verifying existing ones, with opportunities to shape benchmarks used by researchers worldwide.
What You'll Do7
- 1Craft original computer science questions that probe deep understanding and avoid surface recall
- 2Review and refine pre written questions for accuracy, clarity and rigor, making precise edits as needed
- 3Assign difficulty levels such as Medium, Hard, or Expert to each item
- 4Provide a correct answer plus 9 challenging distractors to test expert solvers
- 5Document clear, concise step by step explanations for solutions in markdown format
- 6Suggest 1 to 5 scholarly references per item from reputable sources
- 7For verification tasks, identify and justify issues related to clarity, completeness or solvability
Requirements3
- 1PhD or current doctoral student in Computer Science or a closely related field
- 2Strong grasp of advanced CS theory, algorithms, systems design, or machine learning
- 3Proven ability to express complex ideas clearly in English written form
Who Should Apply
Ideal candidates are researchers or highly experienced practitioners who enjoy building rigorous assessments, have a track record in CS research or competitive programming, and can work independently in a remote, asynchronous setup.
Salary Insight
Compensation ranges from $66.00 to $84.00 per hour with remote, hourly engagement.
Location
Required Skills
Application Tip
Submit a concise sample question and verification task demonstrating your ability to craft precise problem statements and clear solutions, including a difficulty rating and references.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedApplied Engineering Benchmark Specialist | $61-$77/hr Remote
The role focuses on creating and validating rigorous academic assessment content for an AI research effort. You will author and review multiple-choice questions across core engineering topics, judge solution quality, and help build gold-standard benchmarks to push AI capabilities forward. This is a remote, hourly engagement, with tasks that blend technical evaluation and high-quality writing using a clear, structured approach.

Mercor
VerifiedApplied Mathematics Benchmark Specialist | $61-$77/hr Remote
A remote role focused on creating and validating rigorous math assessment content for an AI research initiative. You’ll write and review advanced multiple choice questions across key mathematics domains, assess solution quality, and help set benchmarks that push AI capabilities forward. The work blends mathematical depth with clear, instructional problem design and documentation. This position stands out for its blend of academic rigor and practical benchmarking in AI research.

Mercor
VerifiedApplied Physics Benchmark Specialist | $61-$77/hr Remote
We are looking for seasoned physicists to craft and review advanced academic assessment content for an AI research project. You’ll develop and validate high quality multiple-choice questions across key physics areas, assess solution quality, and help establish benchmark materials that push AI capabilities forward. The role is remote and hourly, with flexible, asynchronous collaboration. You’ll work in two modes: authoring original questions and verifying pre-written items for accuracy and rigor.

Mercor
VerifiedApplied Biology Benchmark Specialist | $60-$75/hr Remote
One sentence on its own. We’re looking for seasoned biologists to craft and review high quality academic assessment content for an AI research project. You’ll develop and verify rigorous biology multiple-choice items, assess solution quality, and help set gold-standard benchmarks that push AI capabilities forward. The role spans authoring original questions and validating existing ones across several biology domains using a remote, asynchronous setup and competitive hourly pay.

Mercor
VerifiedApplied Psychology Benchmark Specialist | $50-$63/hr Remote
A remote role for seasoned psychologists to author and review rigorous assessment items for an AI research effort. You’ll craft original multiple-choice questions and validate existing items to help define benchmarks that advance AI understanding of psychology. The work spans areas like psychometrics, digital health, and consumer psychology, with both creation and verification responsibilities. This opportunity blends scholarly rigor with flexible, asynchronous engagement and compensation on an hourly basis.

Mercor
VerifiedApplied Business & Commerce Benchmark Specialist | $77-$98/hr Remote
We are in search of seasoned business and commerce experts to craft and review high quality assessment content for an AI research program. You’ll develop and validate rigorous multiple choice items across core business areas, assess solution quality, and help define gold standard benchmarks that push AI capabilities forward. The role offers remote, hourly engagement with flexible scheduling and the chance to influence how AI understands business concepts.