WatneyVerified Source

Staff ML Infrastructure Engineer, Robotics

Onsite · San Francisco, California
Posted August 13, 2026
payroll

Overview

Watney builds autonomous robots for data center construction, and you will own the ML infrastructure that turns fleet data into better models. You will drive training and inference systems handling live video and telemetry data from real work sites. Python and PyTorch or TensorFlow power your daily work. The role sits on the ML platform team, collaborating with robotics engineers and data scientists. You will scale distributed training clusters and optimize GPU utilization as the fleet grows.

What You'll Do7

  • 1Design and own training and inference infrastructure that supports large-scale robotic learning
  • 2Build data pipelines that feed training runs with field-collected video and telemetry data
  • 3Launch experiments with fast reproducibility, cutting iteration time for model improvements
  • 4Scale distributed training systems across clusters, handling growing data volumes
  • 5Contribute to core training code, focusing on PyTorch or TensorFlow model execution
  • 6Troubleshoot and resolve GPU performance bottlenecks in production training environments
  • 7Drive infrastructure improvements to increase cluster utilization and throughput

Requirements6

  • 15+ years building ML infrastructure that carried real production training runs
  • 2Experience scaling distributed training systems with large data volumes
  • 3Strong proficiency in Python and PyTorch or TensorFlow
  • 4Proven ability to identify and fix GPU performance bottlenecks in large-scale environments
  • 5Familiarity with data pipelines for video and telemetry data
  • 6Experience with cluster management and scheduling tools like Kubernetes or Slurm

Salary Insight

Salary not disclosed in listing

Location

Typeonsite
LocationSan Francisco, California

Required Skills

pythonpytorchtensorflowgpukubernetes
Share:

Similar open positions

Explore active roles that match your skills and interests.

Finoit Inc.

Finoit Inc.

8d agoSan Francisco, Californiapayroll

Senior ML Infrastructure Engineer PyTorch Kubernetes GPU Training

Design and scale infrastructure for large-scale machine learning training workloads. Build high-performance GPU training platforms and optimize distributed training pipelines. Improve developer experience for ML researchers. This role focuses on cloud-native solutions using AWS and Kubernetes.

250K–320K
pythonawskubernetes+2 more
THE TILTED CIRCLE LLC

THE TILTED CIRCLE LLC

14h agoSan Francisco, Californiapayroll

Senior Software Engineer, ML Platform

You will own the evolution of our ML Platform, building reliable systems for model experimentation, training, evaluation, inference, and retraining that underwrite small businesses. You will join the Infrastructure team, collaborating with data scientists and product engineers to ship developer-friendly tooling. This role stands out because you will set the technical direction for our ML infrastructure, impacting every model-driven decision in the company.

Competitive salary
pythonkubernetesaws+2 more

Applied

15h agoSan Jose, Californiapayroll

Machine Learning Performance Engineer, Offboard Training & Inference

You'll optimize large-scale ML training and inference in the datacenter, driving throughput and cost efficiency for petabytes of autonomy data. Your work directly cuts GPU waste and accelerates iteration across the company. You'll own profiling, performance modeling, and stack-rank optimizations, partnering with teams that build the compute stack. This role targets cluster goodput and cost per unit of data, not vehicle latency. You'll shape tooling and influence technical decisions in a fast-moving environment.

215K–285K
pythonc++tensorrt+2 more
Unisoft Technology Inc

Unisoft Technology Inc

1d agoWashington, District of Columbiacontract

Sr Lead AI Data Engineer

Lead design and delivery of AI infrastructure to drive scalable machine learning solutions across enterprise platforms. Own development of end-to-end ML pipelines and foster collaboration between data science and engineering teams. This role differs by focusing on cross-functional leadership and production-grade MLOps implementation.

Competitive salary
PythonPyTorchTensorFlow+4 more

motional

1d agoRemotepayroll

Senior Software Engineer ML Infrastructure

Design and implement core components of ML infrastructure platform on Kubernetes. Develop scalable services supporting ML lifecycle handling petabytes of data. Own end-to-end features from design to production. Build robust high-throughput systems for AV innovation.

159K–207K
kubernetespythongo+2 more

Friendliai

14h agoRemotepayroll

Cloud Infrastructure Engineer, Kubernetes & AWS

FriendliAI seeks a Cloud Infrastructure Engineer to own the architecture and evolution of the GPU-accelerated AI inference cloud. You will design multi-cluster,Kubernetes fleets, extend the scheduler, and own the network path for latency-sensitive traffic. Work with the inference engine, platform, SRE, and security teams to turn serving demands into platform capabilities. This role offers a hands-on architecture position for an engineer ready to push large clusters further.

Competitive salary
kubernetesawsterraform+2 more