Mercor
MercorVerified Source
Remote

STEM Researcher — Computational Fields | $60-$90/hr Remote

60–90/hr
Remote · United States
Posted July 31, 2026
full-time
10 openings

Overview

A leading AI lab is assembling a team of computational researchers to create the next generation of evaluation benchmarks for frontier AI systems. In this role, you'll turn your everyday research skills — designing studies, testing hypotheses, and scrutinizing results — into complex, multi-step tasks that current models struggle to complete. You'll work in a tight loop with lab researchers to surface methodological gaps that only a working scientist would catch. This is a fully remote, full-time W-2 position (about 35 hours/week) paying $60–$90 per hour, with employment administered by Cincinnatus LLC.

What You'll Do5

  • 1Convert your daily research workflows (study design, hypothesis testing, result evaluation) into engaging, multi-step benchmark tasks.
  • 2Solve your own tasks in Python and notebook environments, producing reference solutions at the level of rigor expected from a careful colleague.
  • 3Articulate the line between sound scientific reasoning and superficially plausible reasoning, and document that standard for evaluators.
  • 4Review model-generated attempts at your tasks, flagging the same kinds of errors a working researcher would immediately notice.
  • 5Collaborate with other researchers and lab staff to keep evaluations consistent, accurate, and aligned with the lab's goals.

Requirements8

  • 1MSc or PhD in a computational STEM field (or a computationally heavy social science/humanities discipline), or equivalent hands-on experience in a research-heavy role requiring coding and data analysis.
  • 2At least 1+ years of active research experience in academia, industry, or a national lab.
  • 3Your own research involves substantial computational work: Python-based analysis, simulation, modeling, or building data pipelines.
  • 4Solid grounding in experimental design, hypothesis testing, and rigorous evaluation of results.
  • 5Working comfort with Git, IDEs, and notebook environments like Jupyter or Colab.
  • 6Past exposure to AI training, model evaluation, or writing evaluation tasks/benchmarks is a plus.
  • 7Detail-obsessed, creative with task design, and able to clearly communicate complex ideas in writing; comfortable working autonomously on ambiguous problems.
  • 8Ability to reliably commit approximately 35 hours per week.

Who Should Apply

The ideal candidate is a working researcher — MSc/PhD level — whose day-to-day work involves heavy computational analysis and a serious methodology. You're the kind of person who instinctively pokes holes in a flimsy conclusion and can build a rigorous, multi-step task that tests real reasoning. Experience with AI evaluation or benchmark authoring is a bonus, and you're comfortable being paid W-2 through a staffing partner while embedding with a top AI lab.

Salary Insight

Hourly rate $60–$90 for a full-time W-2 role (≈35 hrs/week), fully remote within the US. Benefits and payroll are handled by the employer of record.

Location

Typeremote
LocationUnited States
This is a remote position

Required Skills

pythonjupytercolabgitdata analysissimulationmodelingdata pipelinesexperimental designhypothesis testingai model evaluationbenchmark designagentic evaluationresearch methodology

Application Tip

Before applying, craft a short portfolio showing one research task you've designed or one model evaluation you've run — including the rubric you used to score it. That concrete example will separate you from candidates who only list credentials.

Share:

Similar open positions

Explore active roles that match your skills and interests.

Mercor

Mercor

2d agoRemotefull-time

Data Science & Quantitative Analysis Expert | $60-$90/hr Remote

A premier AI research organization is assembling a team of quantitative experts to help design the next generation of evaluation benchmarks for frontier models. In this role, you'll create realistic, hands-on data analysis challenges that push models to their limits — cleaning messy datasets, comparing statistical methods, and verifying whether model outputs hold up to scrutiny. This is a fully remote, W-2 position with Cincinnatus LLC, working about 35 hours per week and collaborating closely with the lab's researchers. If you're passionate about the science behind AI and want your analytical expertise to directly shape how models are measured, this is a unique opportunity.

60–90/hr
· 10 openings
pythonpandasnumpy+11 more
Mercor

Mercor

10d agoRemotefull-time

Software Engineering Expert | $60-$90/hr Remote

This role puts you on the front lines of AI development, working with a top-tier AI lab to design the next generation of evaluation benchmarks for frontier models. As a Software Engineering Expert, you'll craft complex, multi-step engineering challenges that push the limits of today's most advanced AI coding agents. You'll work closely with researchers, using your Python skills and hands-on engineering experience to identify exactly where AI models stumble. It's a fully remote, full-time W-2 position with Cincinnatus LLC, offering the chance to shape how the best AI systems are tested and improved.

60–90/hr
· 10 openings
pythongitdebugging+9 more
Mercor

Mercor

8d agoRemotehourly

Computational Astrophysics & Cosmology Expert | $70-$100/hr Remote

This role is for a computational astrophysics and cosmology expert who will design graduate-level problems that test AI models' ability to use real scientific software. You'll craft challenging tasks, run them against state-of-the-art AI models, and iterate until the difficulty is calibrated. This is not data labeling — it's about creating problems that require deep domain expertise, strategic thinking, and hands-on use of tools like astropy. The work is fully remote, hourly, and flexible within a 15-20 hour weekly commitment.

70–100/hr
astropypythonlinux+7 more
Mercor

Mercor

19d agoRemotefull-time

Machine Learning Engineer — Model Evaluation & Experimentation | $60-$90/hr Remote

Join a leading AI lab's Generative AI team as a Machine Learning Engineer specializing in model evaluation and experimentation. In this role, you'll turn ambitious research ideas into concrete, multi-step ML tasks—like adjusting an RL reward function, implementing the change, running training experiments, and analyzing the results—to reveal exactly where today's frontier models stumble. This is a fully remote, W-2 position through Cincinnatus LLC, offering $60–$90 per hour at roughly 35 hours per week, with direct collaboration alongside top AI researchers.

60–90/hr
· 10 openings
pythongitmachine learning+6 more
Mercor

Mercor

16d agoRemotehourly

Computational Chemistry & Electronic Structure Expert | $70-$100/hr Remote

Join a project that measures how well advanced AI systems tackle real scientific and engineering challenges. As a task designer, you'll craft original, graduate-level computational problems centered on real research workflows, then test them against leading AI models and fine-tune the difficulty until each one lands just right. We're currently seeking experts in computational chemistry and electronic structure—especially those with hands-on PySCF experience—to help build this large-scale benchmark. This is a remote, hourly role that blends deep scientific knowledge with puzzle-like problem design.

70–100/hr
· 10 openings
pyscfpythonlinux+13 more
Mercor

Mercor

7d agoRemotefull-time

Software Engineer, Full Stack (Python, Java, Rust, C#, C++) | $50-$65/hr Remote

This is a full-time remote contract role with a leading AI lab's GenAI team, focused on building and stress-testing software on top of frontier large language models. You'll work directly with the engineering manager to integrate pre-release model APIs, build full-stack tools, and document failure modes that shape future model development. The stack spans Python, Java, Rust, C#, C++, and TypeScript, with production experience in at least one and working proficiency in a second. What sets this apart is the chance to work on cutting-edge AI infrastructure before it ships, with a fast-paced, high-impact environment. Compensation is $50–$65 per hour, and you'll be employed via Cincinnatus LLC.

50–65/hr
· 10 openings
pythonjavarust+12 more