
Site Reliability Engineer, SRE & ML Systems
Overview
You will own the reliability, scalability, and performance of production systems at a global scale. You will work alongside platform and data engineering teams, integrating ML workloads with Python and TensorFlow into robust SRE practices. This role demands a blend of software engineering and systems thinking, operating in a hybrid environment with clear ownership of uptime and incident response.
What You'll Do8
- 1Design and implement automated monitoring and alerting for critical production services using Prometheus and Grafana.
- 2Build and deploy self-healing infrastructure with Kubernetes and Terraform to reduce manual intervention.
- 3Lead incident response and postmortems, driving root cause analysis and preventive actions across the organization.
- 4Optimize system performance through load testing and capacity planning, using Python scripts to analyze bottlenecks.
- 5Scale ML pipelines by integrating TensorFlow and PyTorch models into production, ensuring low-latency inference.
- 6Collaborate with development teams to embed reliability practices into the CI/CD pipeline using GitLab CI and Jenkins.
- 7Automate database failover and backup verification for PostgreSQL and MongoDB to ensure data durability.
- 8Drive service level objectives (SLOs) and error budgets, reporting on reliability metrics to stakeholders.
Requirements8
- 15+ years in site reliability engineering or related field with strong Linux system administration.
- 23+ years programming in Python or R, with production experience in ML libraries such as scikit-learn, TensorFlow, or PyTorch.
- 33+ years with data analysis and visualization using Pandas, NumPy, and Power BI or Tableau.
- 43+ years with SQL and database management, including performance tuning and schema design.
- 5Hands-on experience with cloud platforms (AWS or GCP) and infrastructure as code (Terraform, CloudFormation).
- 6Proven track record of designing and operating Kubernetes clusters in production.
- 7Strong scripting skills in Bash and Python for automation and tooling.
- 8Excellent communication and collaboration skills, with the ability to work effectively in a hybrid onsite environment.
Salary Insight
$104 - $114k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

PRIMUS Global Services Inc.
VerifiedSite Reliability Engineer Hybrid
Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

Xoriant Corporation
VerifiedSenior/Staff SRE, AI/ML Platform Infrastructure
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.

SRI Tech Solutions
VerifiedLead Site Reliability Engineer SRE
Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Dice Talent & Staffing Solutions
VerifiedSite Reliability Engineer, SRE & Cloud Infrastructure
You will own the reliability and scalability of mission-critical systems for a cutting-edge government contractor in San Francisco. You will design, build, and run large-scale distributed systems with AWS, Kubernetes, and Terraform. You will work with a tight-knit team of engineers to ensure 99.99% uptime for services that matter. This role demands a hands-on engineer who thrives on incident response and automation. You will directly shape the infrastructure strategy from day one.

JC CORPORATIONS
VerifiedAI SRE Engineer, AI/ML Platforms & Reliability
You will own the reliability, scalability, and observability of AI/ML production systems at JC CORPORATIONS. You will design and maintain highly reliable environments, monitor performance and latency, and build SLIs/SLOs with error budget tracking. You will work with Kubernetes, Terraform, and Prometheus to automate and optimize infrastructure. This role combines SRE best practices with AI/ML infrastructure to ensure seamless operation of models and pipelines.
Uipath
VerifiedSenior Software Engineer SRE
Design and engineer SRE platform systems using AI while leading cross‑team initiatives. You will identify gaps across teams, design solutions, build, ship, and adopt them, and drive measurable improvements in reliability scalability and performance. This role focuses on ownership accountability and fostering a culture of continuous iteration without relying on generic statements.