Mercor
MercorVerified Source
Remote

CUDA Engineering Expert | $500 One-Time Remote

500–500/hr
Remote
Posted August 10, 2026
task-based
92 openings

Overview

This is a short-term, remote contracting role for GPU engineers who live and breathe kernel-level performance. You'll join a project backed by a leading AI lab, working hands-on with CUDA/HIP kernels and using profiler-driven analysis to squeeze more speed out of modern hardware. You don't need deep algorithmic context — just strong C++, GPU programming instincts, and the ability to find bottlenecks quickly. The engagement is task-based, pays a flat $500, and is ideal for specialists who enjoy focused optimization challenges.

What You'll Do6

  • 1Profile and optimize GPU kernels to improve throughput, efficiency, and hardware utilization.
  • 2Use profiler metrics such as L2 cache hit rate, L2 throughput, and occupancy to drive kernel improvements.
  • 3Identify performance bottlenecks in GPU kernel implementations without needing deep background in the underlying algorithm.
  • 4Write, modify, and reason about code in C++17, Python, and GPU programming languages like CUDA or HIP.
  • 5Apply advanced techniques like inline PTX assembly or tensor core tuning to squeeze out extra performance.
  • 6Document optimization decisions clearly, including when specific profiler metrics were helpful or not.

Requirements8

  • 1Available to commit at least 20 hours per week to the project.
  • 2Fluent in core C++ features through C++17.
  • 3Working knowledge of Python and Git for day-to-day development.
  • 4Proficiency in at least one GPU programming model, such as CUDA, HIP, HLSL, GLSL, Slang, or similar.
  • 5At least 1 year of professional or graduate-level research experience working with GPUs.
  • 6Strong understanding of GPU profiler performance metrics and how to use them for kernel optimization.
  • 7Bonus: experience with CUDA C++ Core Libraries, inline PTX, tensor cores, NVIDIA Blackwell, Nsight Compute, or previous work at GPU hardware vendors like NVIDIA, AMD, or Qualcomm.
  • 8Bonus: open-source contributions related to GPU kernel optimization.

Who Should Apply

You're a freelance GPU specialist who loves the nitty-gritty of kernel tuning and can reason confidently about low-level hardware behavior. You're comfortable stepping into unfamiliar code, using profiler output to guide your hypotheses, and explaining your optimization choices clearly. You're ready for a short, focused contract and can work independently from anywhere.

Salary Insight

The posting mentions a one-time payment of $500 for this remote task-based engagement.

Location

Typeremote
LocationRemote
This is a remote position

Required Skills

cudahipc++pythongithlslglslslangnsight-computenvidia-blackwelltensor-codesinline-ptxl2-cacheoccupancygpu-kernel-optimizationprofiling

Application Tip

Be ready to talk through a specific kernel optimization you've done — especially which profiler metrics you relied on and why. Since the role emphasizes profiler-guided analysis, a quick refresher on Nsight Compute workflows will help you stand out.

Share:

Similar open positions

Explore active roles that match your skills and interests.

Bright Vision Technologies

Bright Vision Technologies

13h agoRemotepayroll

CUDA Developer

Remote CUDA Developer role at Bright Vision Technologies offers full-time opportunities with competitive compensation. This position provides extensive career growth within a leading technology consulting firm specializing in cloud AI and enterprise solutions.

85K–110K
CUDA C/C++GPU programmingC+++5 more
Mercor

Mercor

18d agoRemotehourly

Performance Engineer (C++, Python, Rust) | $70-$110/hr Remote

This is a fully remote Performance Engineer role supporting a leading AI lab's GenAI team, where your work directly helps improve how large language models are trained and evaluated. You'll bring deep expertise in low-level systems optimization across C++, Python, or Rust to design and assess performance engineering tasks that generate high-quality training data. As a W-2 contractor through Cincinnatus LLC, you'll join a full-time, 40-hour-per-week engagement and collaborate with researchers building some of the most advanced AI systems in the world.

70–110/hr
· 10 openings
c++pythonrust+8 more
Mercor

Mercor

10d agoRemotefull-time

Software Engineering Expert | $60-$90/hr Remote

This role puts you on the front lines of AI development, working with a top-tier AI lab to design the next generation of evaluation benchmarks for frontier models. As a Software Engineering Expert, you'll craft complex, multi-step engineering challenges that push the limits of today's most advanced AI coding agents. You'll work closely with researchers, using your Python skills and hands-on engineering experience to identify exactly where AI models stumble. It's a fully remote, full-time W-2 position with Cincinnatus LLC, offering the chance to shape how the best AI systems are tested and improved.

60–90/hr
· 10 openings
pythongitdebugging+9 more
Intel

Intel

14h agoSan Jose, Californiapayroll

AI Infrastructure Engineer Intel

Performance‑obsessed AI Infrastructure Engineer at Intel in San Jose, California. You will drive inference performance and redefine peak performance on Intel’s next‑generation GPU architectures.

170K–315K
C++PythonGPU Computing+8 more
NVIDIA

NVIDIA

10h agoAustin, Texaspayroll

Senior Compiler Engineer Infrastructure at NVIDIA

You will align NVIDIA's compiler codebases with open-source ecosystems, focusing on LLVM, Clang, and MLIR. Own the reconciliation of downstream repositories with upstream projects, building tooling that scales across thousands of engineers. You will join the Compute Compiler Team, reporting to the director of compiler infrastructure, and collaborate with compiler developers and CI teams. This role stands out for its dual focus on open-source stewardship and developer productivity, with access to NVIDIA's GPU stack.

152K–242K
C++Open-source compiler frameworksLLVM+9 more
Celestica

Celestica

9h agoDallas, Texaspayroll

Senior Lead Software Engineer, GPU Data Centers

You will architect and validate a full stack application for next-generation data centers with GPU/AI compute elements. Build orchestration software for the entire rack, integrated visualization tools, and diagnostics to optimize GPU utilization. Collaborate with cross-functional teams to ship production-ready code and mentor engineers. This role stands out through its focus on Cloud Native methods, Kubernetes deployments, and GenAI tool adoption for development efficiency.

Competitive salary
PythonGoKubernetes+7 more