Incorporan Inc
Incorporan IncVerified Source

Site Reliability Engineer, SRE & ML Systems

104K–114K
Onsite · Chicago, Illinois
Posted August 13, 2026
payroll

Overview

You will own the reliability, scalability, and performance of production systems at a global scale. You will work alongside platform and data engineering teams, integrating ML workloads with Python and TensorFlow into robust SRE practices. This role demands a blend of software engineering and systems thinking, operating in a hybrid environment with clear ownership of uptime and incident response.

What You'll Do8

  • 1Design and implement automated monitoring and alerting for critical production services using Prometheus and Grafana.
  • 2Build and deploy self-healing infrastructure with Kubernetes and Terraform to reduce manual intervention.
  • 3Lead incident response and postmortems, driving root cause analysis and preventive actions across the organization.
  • 4Optimize system performance through load testing and capacity planning, using Python scripts to analyze bottlenecks.
  • 5Scale ML pipelines by integrating TensorFlow and PyTorch models into production, ensuring low-latency inference.
  • 6Collaborate with development teams to embed reliability practices into the CI/CD pipeline using GitLab CI and Jenkins.
  • 7Automate database failover and backup verification for PostgreSQL and MongoDB to ensure data durability.
  • 8Drive service level objectives (SLOs) and error budgets, reporting on reliability metrics to stakeholders.

Requirements8

  • 15+ years in site reliability engineering or related field with strong Linux system administration.
  • 23+ years programming in Python or R, with production experience in ML libraries such as scikit-learn, TensorFlow, or PyTorch.
  • 33+ years with data analysis and visualization using Pandas, NumPy, and Power BI or Tableau.
  • 43+ years with SQL and database management, including performance tuning and schema design.
  • 5Hands-on experience with cloud platforms (AWS or GCP) and infrastructure as code (Terraform, CloudFormation).
  • 6Proven track record of designing and operating Kubernetes clusters in production.
  • 7Strong scripting skills in Bash and Python for automation and tooling.
  • 8Excellent communication and collaboration skills, with the ability to work effectively in a hybrid onsite environment.

Salary Insight

$104 - $114k per year

Location

Typeonsite
LocationChicago, Illinois

Required Skills

pythontensorflowpytorchscikit-learnkubernetes
Share:

Similar open positions

Explore active roles that match your skills and interests.

PRIMUS Global Services Inc.

PRIMUS Global Services Inc.

1d agoChicago, Illinoiscontract

Site Reliability Engineer Hybrid

Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

45K–50K
pythonawskubernetes+2 more
Xoriant Corporation

Xoriant Corporation

1d agoSan Jose, Californiapayroll

Senior/Staff SRE, AI/ML Platform Infrastructure

Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.

Competitive salary
Production on-callIncident commandBlameless postmortem+15 more
SRI Tech Solutions

SRI Tech Solutions

1d agoOrlando, Floridapayroll

Lead Site Reliability Engineer SRE

Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Competitive salary
Site Reliability EngineeringCloud InfrastructureKubernetes+3 more
Dice Talent & Staffing Solutions

Dice Talent & Staffing Solutions

11h agoSan Francisco, Californiapayroll

Site Reliability Engineer, SRE & Cloud Infrastructure

You will own the reliability and scalability of mission-critical systems for a cutting-edge government contractor in San Francisco. You will design, build, and run large-scale distributed systems with AWS, Kubernetes, and Terraform. You will work with a tight-knit team of engineers to ensure 99.99% uptime for services that matter. This role demands a hands-on engineer who thrives on incident response and automation. You will directly shape the infrastructure strategy from day one.

180K–265K
awskubernetesterraform+2 more
JC CORPORATIONS

JC CORPORATIONS

11h agoSan Jose, Californiapayroll

AI SRE Engineer, AI/ML Platforms & Reliability

You will own the reliability, scalability, and observability of AI/ML production systems at JC CORPORATIONS. You will design and maintain highly reliable environments, monitor performance and latency, and build SLIs/SLOs with error budget tracking. You will work with Kubernetes, Terraform, and Prometheus to automate and optimize infrastructure. This role combines SRE best practices with AI/ML infrastructure to ensure seamless operation of models and pipelines.

146K–208K
pythonawskubernetes+2 more

Uipath

1d agoDenver, Coloradopayroll

Senior Software Engineer SRE

Design and engineer SRE platform systems using AI while leading cross‑team initiatives. You will identify gaps across teams, design solutions, build, ship, and adopt them, and drive measurable improvements in reliability scalability and performance. This role focuses on ownership accountability and fostering a culture of continuous iteration without relying on generic statements.

160K–210K
pythonawsreact+2 more