Machine Learning Performance Engineer, Offboard Training & Inference
Overview
You'll optimize large-scale ML training and inference in the datacenter, driving throughput and cost efficiency for petabytes of autonomy data. Your work directly cuts GPU waste and accelerates iteration across the company. You'll own profiling, performance modeling, and stack-rank optimizations, partnering with teams that build the compute stack. This role targets cluster goodput and cost per unit of data, not vehicle latency. You'll shape tooling and influence technical decisions in a fast-moving environment.
What You'll Do7
- 1Profile end-to-end distributed training with NCCL and PyTorch, fixing data loader stalls, kernel inefficiencies, and checkpointing bottlenecks.
- 2Optimize batch inference on Triton and TensorRT, tuning batching, quantization, and graph execution to saturate accelerators.
- 3Build roofline models for workloads, measuring the gap between theoretical and achieved performance, and rank optimizations by impact.
- 4Improve multi-node scaling with FSDP and DeepSpeed, addressing communication and memory-bandwidth issues.
- 5Drive cluster goodput by reducing GPU idle time from I/O, scheduling, stragglers, and failure recovery.
- 6Develop benchmarking and observability tools to catch performance regressions as models evolve.
- 7Collaborate with cross-functional teams to solve complex compute and data challenges.
Requirements7
- 15+ years in ML performance engineering, profiling production systems with PyTorch Profiler or Nsight.
- 23+ years with distributed training at scale using FSDP, DeepSpeed, or Megatron.
- 3Deep knowledge of GPU concepts: memory bandwidth, kernel launch, occupancy, quantization.
- 4Experience with Triton, TensorRT, ONNX Runtime, or Ray for high-throughput inference.
- 5Fluency in Python and proficiency in C++.
- 6Strong debugging and analytical skills for root-cause investigations.
- 7Solid understanding of ML foundations and ability to solve novel problems.
Salary Insight
$215 - $285k per year
Similar open positions
Explore active roles that match your skills and interests.
zoox
VerifiedMachine Learning Engineer Multi-Modality Foundation Model
Design and manage system capacity for perception models on autonomous vehicles. Lead initiatives to share perception stack components for efficient inference. Focus on delivering production-ready large-scale models for on-vehicle stacks. Seek experts experienced in compressing accelerating deploying complex models for power-constrained SoCs.
Intel
VerifiedAI Infrastructure Engineer Intel
Performance‑obsessed AI Infrastructure Engineer at Intel in San Jose, California. You will drive inference performance and redefine peak performance on Intel’s next‑generation GPU architectures.

Technogen, Inc.
VerifiedSenior LLM Inference & GPU Systems Engineer
Technogen, Inc. seeks a Senior On-Premise LLM Inference & GPU Systems Engineer to architect and optimize high-performance inference platforms in Charlotte, NC. You will own the deployment and tuning of large language models on on-premise GPU clusters, ensuring low-latency, high-throughput serving for enterprise workloads. Collaborate with data scientists and infrastructure teams to build scalable ML pipelines. This role offers direct impact on production AI systems within a Woman-Owned Small Business with 15+ years of IT services.
Watney
VerifiedStaff ML Infrastructure Engineer, Robotics
Watney builds autonomous robots for data center construction, and you will own the ML infrastructure that turns fleet data into better models. You will drive training and inference systems handling live video and telemetry data from real work sites. Python and PyTorch or TensorFlow power your daily work. The role sits on the ML platform team, collaborating with robotics engineers and data scientists. You will scale distributed training clusters and optimize GPU utilization as the fleet grows.

Shree Narayani Networking Solutions LLC
VerifiedSenior On-Premise LLM Inference & GPU Systems Engineer
You will build and maintain large-scale on-prem LLM infrastructure on NVIDIA H200 GPU clusters, powering enterprise private GenAI workloads. You will manage production inference, including self-hosting Llama and other open-source models, within an OpenShift AI deployment ecosystem. Your stack includes Kubernetes, Docker, and GPU orchestration, with a focus on reliability and performance. This role is pivotal for scaling inference in a hybrid cloud environment.
Security-Level-5
VerifiedMachine Learning Engineer, AI Security Datacenter
Design experiments that measure the impact of Security Level 5 controls on real ML workflows at a Bay Area AI security nonprofit. Work with AI labs and US intelligence agencies to defend frontier models against nation-state threats. Build the reference tech stack for the first SL5 datacenter, shipping in 2-3 years. Forecast 2028 workloads and run experiments at tractable scales, interpreting results for frontier-scale audiences.