
ML Engineer (Coding Agent Experience) | $85/hr Remote
Overview
We're looking for an ML Engineer to support a frontier code agents project in partnership with a leading AI research lab. You'll use advanced AI coding tools to evaluate and improve cutting-edge coding models, working through realistic machine learning workflows. This is a remote, sprint-based contract role with competitive hourly compensation — and spots are filling fast.
What You'll Do5
- 1Leverage frontier AI coding agents to complete complex ML and AI engineering tasks and assess model outputs.
- 2Review model-generated implementations covering model training, inference systems, MLOps, and LLM applications.
- 3Pinpoint bugs, edge cases, and performance bottlenecks in generated code and model behavior.
- 4Compare outputs from multiple frontier models, documenting their relative strengths and weaknesses.
- 5Apply professional engineering judgment to realistic production scenarios to determine whether model solutions are viable.
Requirements5
- 1At least 2 years of professional experience in machine learning engineering.
- 2Hands-on experience building production ML systems, deployment infrastructure, LLM applications, or AI-powered products.
- 3Regularly use AI coding agents like Cursor, Claude Code, Codex, Windsurf, or Gemini CLI.
- 4Ability to critically evaluate model-generated ML implementations and understand technical tradeoffs.
- 5Experience deploying ML systems to production is a plus.
Who Should Apply
The ideal candidate is a hands-on ML engineer who spends as much time in an AI coding agent as in a Jupyter notebook. You're comfortable evaluating model-generated code, have a sharp eye for failure modes, and can make sound judgment calls on realistic ML engineering problems. You're looking for a flexible, sprint-based contract that fits around your schedule and want to help shape the next generation of frontier coding models.
Salary Insight
$85/hour, paid per accepted task — typically $400 per task that takes 2–3 hours after initial ramp-up. Compensation is tied to accepted work.
Location
Required Skills
Application Tip
In your application, give concrete examples of how you've used AI coding agents on real ML projects — including any bugs you caught in generated code and how you evaluated model output for correctness.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedData Engineer (Coding Agent Experience) | $80/hr Remote
Mercor is teaming up with a leading AI research lab to support a Frontier Code Agents project, and we need data engineers to help push the boundaries of AI-assisted coding. You'll use frontier AI coding tools like Cursor, Claude Code, or Codex to execute and evaluate complex data engineering tasks, then assess how well these models handle real-world infrastructure challenges. Your insights will directly improve the next generation of AI coding models, making this a rare chance to shape the future of software development. With limited spots and a fast-filling roster, this sprint-based, remote gig is ideal for engineers who enjoy dissecting code and spotting subtle flaws.
Mercor
VerifiedBackend Engineer, AI Coding Agent
You will use AI coding agents like Cursor, Claude Code, and Codex to complete and evaluate complex backend engineering tasks. Your assessments directly improve frontier model performance. You will work remotely with a team backed by Benchmark and General Catalyst, reviewing code for correctness, quality, and edge cases. This contract role pays $400 per accepted task.

Mercor
VerifiedDevOps / SRE / Cloud Engineer (Coding Agent Experience) | $85/hr Remote
You'll join a select group of engineers working with a leading AI research lab to test and improve frontier AI coding models. This remote role centers on DevOps, SRE, and cloud engineering tasks, using AI coding agents to evaluate realistic infrastructure workflows. You'll review model-generated code, spot flaws, and help shape the next generation of coding tools. Spots are limited, so if this sounds like you, apply now.
David Joseph & Company
VerifiedApplied AI Engineer, Agent Harness & LLM Systems
You will own the core agent harness that turns raw model capability into dependable product for real users. You'll build the execution loop, tool-use strategies, and context construction, and shape agent behavior across real customer workflows. You'll work with Python or TypeScript, modern AI tooling, and sandboxed microVM execution. The role sits on a small founding team at a seed-stage startup backed by strong operators, with design partners already live. You'll define the boundary between runtime model decisions and deterministic, compiled code.

Wise Skulls Corp.
VerifiedAI Engineer (LLM Agents & Data Engineering)
Lead design and delivery of AI solutions that scale across multiple platforms. Own the end-to-end pipeline from concept to production while driving innovation in large language models. This role shapes how our systems learn and adapt.

Mercor
VerifiedSoftware Engineering Expert | $60-$90/hr Remote
This role puts you on the front lines of AI development, working with a top-tier AI lab to design the next generation of evaluation benchmarks for frontier models. As a Software Engineering Expert, you'll craft complex, multi-step engineering challenges that push the limits of today's most advanced AI coding agents. You'll work closely with researchers, using your Python skills and hands-on engineering experience to identify exactly where AI models stumble. It's a fully remote, full-time W-2 position with Cincinnatus LLC, offering the chance to shape how the best AI systems are tested and improved.