
STEM Researcher — Computational Fields | $60-$90/hr Remote
Overview
A leading AI lab is assembling a team of computational researchers to create the next generation of evaluation benchmarks for frontier AI systems. In this role, you'll turn your everyday research skills — designing studies, testing hypotheses, and scrutinizing results — into complex, multi-step tasks that current models struggle to complete. You'll work in a tight loop with lab researchers to surface methodological gaps that only a working scientist would catch. This is a fully remote, full-time W-2 position (about 35 hours/week) paying $60–$90 per hour, with employment administered by Cincinnatus LLC.
What You'll Do5
- 1Convert your daily research workflows (study design, hypothesis testing, result evaluation) into engaging, multi-step benchmark tasks.
- 2Solve your own tasks in Python and notebook environments, producing reference solutions at the level of rigor expected from a careful colleague.
- 3Articulate the line between sound scientific reasoning and superficially plausible reasoning, and document that standard for evaluators.
- 4Review model-generated attempts at your tasks, flagging the same kinds of errors a working researcher would immediately notice.
- 5Collaborate with other researchers and lab staff to keep evaluations consistent, accurate, and aligned with the lab's goals.
Requirements8
- 1MSc or PhD in a computational STEM field (or a computationally heavy social science/humanities discipline), or equivalent hands-on experience in a research-heavy role requiring coding and data analysis.
- 2At least 1+ years of active research experience in academia, industry, or a national lab.
- 3Your own research involves substantial computational work: Python-based analysis, simulation, modeling, or building data pipelines.
- 4Solid grounding in experimental design, hypothesis testing, and rigorous evaluation of results.
- 5Working comfort with Git, IDEs, and notebook environments like Jupyter or Colab.
- 6Past exposure to AI training, model evaluation, or writing evaluation tasks/benchmarks is a plus.
- 7Detail-obsessed, creative with task design, and able to clearly communicate complex ideas in writing; comfortable working autonomously on ambiguous problems.
- 8Ability to reliably commit approximately 35 hours per week.
Who Should Apply
The ideal candidate is a working researcher — MSc/PhD level — whose day-to-day work involves heavy computational analysis and a serious methodology. You're the kind of person who instinctively pokes holes in a flimsy conclusion and can build a rigorous, multi-step task that tests real reasoning. Experience with AI evaluation or benchmark authoring is a bonus, and you're comfortable being paid W-2 through a staffing partner while embedding with a top AI lab.
Salary Insight
Hourly rate $60–$90 for a full-time W-2 role (≈35 hrs/week), fully remote within the US. Benefits and payroll are handled by the employer of record.
Location
Required Skills
Application Tip
Before applying, craft a short portfolio showing one research task you've designed or one model evaluation you've run — including the rubric you used to score it. That concrete example will separate you from candidates who only list credentials.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedData Science & Quantitative Analysis Expert | $60-$90/hr Remote
A premier AI research organization is assembling a team of quantitative experts to help design the next generation of evaluation benchmarks for frontier models. In this role, you'll create realistic, hands-on data analysis challenges that push models to their limits — cleaning messy datasets, comparing statistical methods, and verifying whether model outputs hold up to scrutiny. This is a fully remote, W-2 position with Cincinnatus LLC, working about 35 hours per week and collaborating closely with the lab's researchers. If you're passionate about the science behind AI and want your analytical expertise to directly shape how models are measured, this is a unique opportunity.

Mercor
VerifiedSoftware Engineering Expert | $60-$90/hr Remote
This role puts you on the front lines of AI development, working with a top-tier AI lab to design the next generation of evaluation benchmarks for frontier models. As a Software Engineering Expert, you'll craft complex, multi-step engineering challenges that push the limits of today's most advanced AI coding agents. You'll work closely with researchers, using your Python skills and hands-on engineering experience to identify exactly where AI models stumble. It's a fully remote, full-time W-2 position with Cincinnatus LLC, offering the chance to shape how the best AI systems are tested and improved.

Mercor
VerifiedComputational Astrophysics & Cosmology Expert | $70-$100/hr Remote
This role is for a computational astrophysics and cosmology expert who will design graduate-level problems that test AI models' ability to use real scientific software. You'll craft challenging tasks, run them against state-of-the-art AI models, and iterate until the difficulty is calibrated. This is not data labeling — it's about creating problems that require deep domain expertise, strategic thinking, and hands-on use of tools like astropy. The work is fully remote, hourly, and flexible within a 15-20 hour weekly commitment.

Mercor
VerifiedMachine Learning Engineer — Model Evaluation & Experimentation | $60-$90/hr Remote
Join a leading AI lab's Generative AI team as a Machine Learning Engineer specializing in model evaluation and experimentation. In this role, you'll turn ambitious research ideas into concrete, multi-step ML tasks—like adjusting an RL reward function, implementing the change, running training experiments, and analyzing the results—to reveal exactly where today's frontier models stumble. This is a fully remote, W-2 position through Cincinnatus LLC, offering $60–$90 per hour at roughly 35 hours per week, with direct collaboration alongside top AI researchers.

Mercor
VerifiedComputational Chemistry & Electronic Structure Expert | $70-$100/hr Remote
Join a project that measures how well advanced AI systems tackle real scientific and engineering challenges. As a task designer, you'll craft original, graduate-level computational problems centered on real research workflows, then test them against leading AI models and fine-tune the difficulty until each one lands just right. We're currently seeking experts in computational chemistry and electronic structure—especially those with hands-on PySCF experience—to help build this large-scale benchmark. This is a remote, hourly role that blends deep scientific knowledge with puzzle-like problem design.

Mercor
VerifiedSoftware Engineer, Full Stack (Python, Java, Rust, C#, C++) | $50-$65/hr Remote
This is a full-time remote contract role with a leading AI lab's GenAI team, focused on building and stress-testing software on top of frontier large language models. You'll work directly with the engineering manager to integrate pre-release model APIs, build full-stack tools, and document failure modes that shape future model development. The stack spans Python, Java, Rust, C#, C++, and TypeScript, with production experience in at least one and working proficiency in a second. What sets this apart is the chance to work on cutting-edge AI infrastructure before it ships, with a fast-paced, high-impact environment. Compensation is $50–$65 per hour, and you'll be employed via Cincinnatus LLC.