
Senior ML Infrastructure Engineer PyTorch Kubernetes GPU Training
Overview
Design and scale infrastructure for large-scale machine learning training workloads. Build high-performance GPU training platforms and optimize distributed training pipelines. Improve developer experience for ML researchers. This role focuses on cloud-native solutions using AWS and Kubernetes.
What You'll Do5
- 1Design and scale distributed ML training infrastructure for large GPU clusters
- 2Build high-performance GPU training platforms using PyTorch and Kubernetes
- 3Optimize distributed training pipelines for efficiency
- 4Improve developer experience for ML researchers
- 5Ship scalable solutions that support rapid model iteration
Requirements5
- 15+ years building ETL pipelines with Spark and Airflow
- 2Experience with PyTorch and Kubernetes
- 3Knowledge of AWS services for distributed computing
- 4Proficiency in container orchestration with Kubernetes
- 5Understanding of GPU training frameworks and performance tuning
Salary Insight
$250 - $320k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.
motional
VerifiedSenior Software Engineer ML Infrastructure
Design and implement core components of ML infrastructure platform on Kubernetes. Develop scalable services supporting ML lifecycle handling petabytes of data. Own end-to-end features from design to production. Build robust high-throughput systems for AV innovation.

Unisoft Technology Inc
VerifiedSr Lead AI Data Engineer
Lead design and delivery of AI infrastructure to drive scalable machine learning solutions across enterprise platforms. Own development of end-to-end ML pipelines and foster collaboration between data science and engineering teams. This role differs by focusing on cross-functional leadership and production-grade MLOps implementation.

Finoit Inc.
VerifiedSenior Software Engineer Data Infrastructure
Design and lead development of scalable data pipelines for AI/ML platforms. Own design and implementation of distributed systems using Python and cloud services. Drive improvements in data quality and visualization. Lead cross-functional teams to deliver high-performance solutions.

ClifyX
VerifiedSenior Software Engineer, Machine Learning Platform
You will own the ML platform roadmap, building distributed systems that process terabytes of training data daily. You will join a platform team of 8 engineers, working alongside data scientists and infrastructure specialists. Your work will directly impact model iteration speed for the company's core ranking and recommendation products. This contract role in Sunnyvale offers a chance to design systems from scratch, not just maintain existing pipelines.

VDart, Inc.
VerifiedSenior AI DevOps Engineer AI Ops Platform Engineering
Design and implement scalable AI-driven CI/CD pipelines using Python AWS React Kubernetes Spark Airflow. Own automated workflows that accelerate software delivery while integrating generative AI and model context protocol solutions.

Intellectt INC
VerifiedAI Data Center Architect
Design and architect modern AI infrastructure at enterprise or hyperscale scale. Lead GPU infrastructure and AI factories solutions. Differentiate from traditional data center roles.