Mercor
MercorVerified Source
Remote

Software Engineering Expert | $60-$90/hr Remote

60–90/hr
Remote · United States
Posted August 2, 2026
full-time
10 openings

Overview

This role puts you on the front lines of AI development, working with a top-tier AI lab to design the next generation of evaluation benchmarks for frontier models. As a Software Engineering Expert, you'll craft complex, multi-step engineering challenges that push the limits of today's most advanced AI coding agents. You'll work closely with researchers, using your Python skills and hands-on engineering experience to identify exactly where AI models stumble. It's a fully remote, full-time W-2 position with Cincinnatus LLC, offering the chance to shape how the best AI systems are tested and improved.

What You'll Do5

  • 1Develop realistic, multi-step software engineering tasks that reflect genuine on-the-job challenges, specifically designed to test the boundaries of cutting-edge AI coding agents.
  • 2Solve your own tasks using Python, including setting up environments and creating verification checks so every problem has a clear, reproducible answer.
  • 3Integrate AI coding assistants into your daily workflow, tracking where they succeed and where they hit limitations.
  • 4Review tasks submitted by other experts and provide constructive feedback on clarity, correctness, and overall difficulty.
  • 5Analyze how AI agents approach your tasks and collaborate with the research team to understand failure patterns and improve future benchmarks.

Requirements8

  • 1A master's or PhD in computer science or a related STEM field, or equivalent hands-on experience in a research-focused role with substantial coding and data analysis.
  • 2At least one year of professional experience in software engineering, research engineering, or a similar role.
  • 3Strong Python scripting and debugging abilities, with a clean-code mindset and emphasis on readability.
  • 4Fluency with Git, IDEs, and standard software development workflows in an everyday setting.
  • 5Familiarity with AI coding assistants, prompt engineering, or agent-based workflows is a plus.
  • 6Previous experience in AI training, model evaluation, or creating benchmarks/tasks is preferred.
  • 7A detail-oriented, creative approach to problem-solving, strong written communication skills, and the ability to work independently on ambiguous challenges.
  • 8Reliable availability for roughly 35 hours per week.

Who Should Apply

You're an experienced software engineer with deep Python expertise and a passion for AI — especially model evaluation and benchmarking. You thrive on designing hard problems, have a knack for identifying weak spots in AI systems, and enjoy working autonomously within a collaborative research environment.

Salary Insight

$60.00 - $90.00 per hour

Location

Typeremote
LocationUnited States
This is a remote position

Required Skills

pythongitdebuggingscriptingai coding assistantsprompt engineeringagent workflowsmodel evaluationbenchmarkingtask authoringresearch engineeringsoftware engineering

Application Tip

Before applying, prepare a short portfolio of Python-heavy projects or tasks you've authored. Highlight any experience you have with AI evaluation, prompt engineering, or building test suites — concrete examples will make your application stand out.

Share:

Similar open positions

Explore active roles that match your skills and interests.

Mercor

Mercor

7d agoRemotefull-time

Software Engineer, Full Stack (Python, Java, Rust, C#, C++) | $50-$65/hr Remote

This is a full-time remote contract role with a leading AI lab's GenAI team, focused on building and stress-testing software on top of frontier large language models. You'll work directly with the engineering manager to integrate pre-release model APIs, build full-stack tools, and document failure modes that shape future model development. The stack spans Python, Java, Rust, C#, C++, and TypeScript, with production experience in at least one and working proficiency in a second. What sets this apart is the chance to work on cutting-edge AI infrastructure before it ships, with a fast-paced, high-impact environment. Compensation is $50–$65 per hour, and you'll be employed via Cincinnatus LLC.

50–65/hr
· 10 openings
pythonjavarust+12 more
Mercor

Mercor

6d agoRemotefull-time

Senior Software Engineer, Full Stack (Python, Java, Rust, C#, C++) | $90-$110/hr Remote

Here's a chance to work on the cutting edge of AI directly with a leading AI lab's GenAI team. As a Senior Full-Stack Engineer, you'll help build and test foundational large language models by creating real applications that push their limits. Your stack will span Python, Java, Rust, C#, C++, and TypeScript, with a focus on quick integration of pre-release APIs. This is a remote, full-time W-2 role through Cincinnatus LLC, offering $90-$110 per hour and the opportunity to collaborate directly with an engineering manager on fast-moving projects.

90–110/hr
· 10 openings
pythonjavarust+4 more
Mercor

Mercor

14d agoRemotefull-time

Software Engineer, Full Stack — India | $25-$30/hr Remote

This role puts you on a frontier AI lab's GenAI team, building and testing real software on top of large language models before they're released. You'll work across the full stack—Python, Java, Rust, C#, C++, or TypeScript—to integrate model APIs, build evaluation harnesses, and diagnose failure modes. The pace is fast, requirements shift, and you're expected to contribute working code within your first week. It's a fully remote, 40-hour-per-week engagement paying $25–$30/hr.

25–30/hr
· 10 openings
pythonjavarust+12 more
Mercor

Mercor

2d agoRemotefull-time

QA/Test Engineer | $60-$90/hr Remote

A leading AI lab is looking for a meticulous QA/Test Engineer to help build evaluation benchmarks for cutting-edge agentic AI models. Your work ensures complex multi-step tasks are unambiguous, correctly graded, and free of shortcuts, directly impacting how frontier models are measured. You'll dive into each task, run it, probe edge cases, and debug environments, working in a tight loop with researchers and task authors. This is a fully remote, W-2 role with Cincinnatus LLC, embedded within the AI lab's extended workforce, offering a competitive hourly rate of $60–$90.

60–90/hr
· 10 openings
pythongittest engineering+6 more
Mercor

Mercor

13d agoRemotefull-time

STEM Researcher — Computational Fields | $60-$90/hr Remote

A leading AI lab is assembling a team of computational researchers to create the next generation of evaluation benchmarks for frontier AI systems. In this role, you'll turn your everyday research skills — designing studies, testing hypotheses, and scrutinizing results — into complex, multi-step tasks that current models struggle to complete. You'll work in a tight loop with lab researchers to surface methodological gaps that only a working scientist would catch. This is a fully remote, full-time W-2 position (about 35 hours/week) paying $60–$90 per hour, with employment administered by Cincinnatus LLC.

60–90/hr
· 10 openings
pythonjupytercolab+11 more
Mercor

Mercor

8d agoRemotehourly

Data Engineer (Coding Agent Experience) | $80/hr Remote

Mercor is teaming up with a leading AI research lab to support a Frontier Code Agents project, and we need data engineers to help push the boundaries of AI-assisted coding. You'll use frontier AI coding tools like Cursor, Claude Code, or Codex to execute and evaluate complex data engineering tasks, then assess how well these models handle real-world infrastructure challenges. Your insights will directly improve the next generation of AI coding models, making this a rare chance to shape the future of software development. With limited spots and a fast-filling roster, this sprint-based, remote gig is ideal for engineers who enjoy dissecting code and spotting subtle flaws.

80–80/hr
· 18 openings
ai coding agentscursorclaude code+9 more