JC CORPORATIONS
JC CORPORATIONSVerified Source

AI SRE Engineer, AI/ML Platforms & Reliability

146K–208K
Onsite · San Jose, California
Posted August 13, 2026
payroll

Overview

You will own the reliability, scalability, and observability of AI/ML production systems at JC CORPORATIONS. You will design and maintain highly reliable environments, monitor performance and latency, and build SLIs/SLOs with error budget tracking. You will work with Kubernetes, Terraform, and Prometheus to automate and optimize infrastructure. This role combines SRE best practices with AI/ML infrastructure to ensure seamless operation of models and pipelines.

What You'll Do8

  • 1Design and maintain highly reliable and scalable AI/ML production environments on AWS.
  • 2Monitor availability, performance, latency, and reliability of AI/ML applications using Prometheus and Grafana.
  • 3Build and maintain SLIs, SLOs, and SLAs with error budget tracking and alerting.
  • 4Automate incident response and runbook execution using Python and Terraform.
  • 5Optimize model serving infrastructure with Kubernetes and KServe for low latency and high throughput.
  • 6Drive capacity planning and performance tuning for GPU clusters and model inference.
  • 7Collaborate with data scientists to integrate MLflow for model lifecycle management and deployment.
  • 8Implement chaos engineering and load testing to validate system resilience and scalability.

Requirements10

  • 15+ years in SRE, DevOps, or AI/ML infrastructure roles.
  • 2Strong experience with Kubernetes and container orchestration in production.
  • 3Proficiency in Python and at least one infrastructure-as-code tool like Terraform.
  • 4Hands-on experience with monitoring and observability tools: Prometheus, Grafana, Datadog, or similar.
  • 5Deep understanding of SLI/SLO concepts and error budget policies.
  • 6Experience with AWS or other cloud providers, and GPU computing environments.
  • 7Familiarity with model serving frameworks like KServe, TorchServe, or TensorFlow Serving.
  • 8Experience with MLflow or similar MLOps tools for model lifecycle management.
  • 9Strong scripting skills in Bash and Python for automation.
  • 10Excellent troubleshooting and incident management skills in production environments.

Salary Insight

$146 - $208k per year

Location

Typeonsite
LocationSan Jose, California

Required Skills

pythonawskubernetesterraformprometheus
Share:

Similar open positions

Explore active roles that match your skills and interests.

Xoriant Corporation

Xoriant Corporation

1d agoSan Jose, Californiapayroll

Senior/Staff SRE, AI/ML Platform Infrastructure

Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.

Competitive salary
Production on-callIncident commandBlameless postmortem+15 more
JC CORPORATIONS

JC CORPORATIONS

9h agoSan Jose, Californiapayroll

DevOps AI Engineer

You will own the AI/ML infrastructure and deployment pipelines for JC CORPORATIONS' AI applications. You will design, build, and automate scalable systems that deliver models to production. Your stack spans AWS, Azure, Google Cloud, Kubernetes, and MLOps tools. You will collaborate with data scientists and software engineers to streamline model deployment. This role combines deep DevOps expertise with AI/ML platform engineering, driving the backbone of our AI services.

146K–208K
kubernetesawsdocker+2 more
Incorporan Inc

Incorporan Inc

9h agoChicago, Illinoispayroll

Site Reliability Engineer, SRE & ML Systems

You will own the reliability, scalability, and performance of production systems at a global scale. You will work alongside platform and data engineering teams, integrating ML workloads with Python and TensorFlow into robust SRE practices. This role demands a blend of software engineering and systems thinking, operating in a hybrid environment with clear ownership of uptime and incident response.

104K–114K
pythontensorflowpytorch+2 more
SRI Tech Solutions

SRI Tech Solutions

1d agoOrlando, Floridapayroll

Lead Site Reliability Engineer SRE

Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Competitive salary
Site Reliability EngineeringCloud InfrastructureKubernetes+3 more

Hi Marley

1d agoBoston, Massachusettspayroll

Sr. Software Engineer II DevOps

Lead design and operation of cloud infrastructure on AWS to support core SaaS platform and agentic AI services. Build AI/ML infrastructure and monitoring for LLM-powered services. Establish IaC standards using Terraform. Implement observability beyond availability. Support big data pipelines, warehousing, and analytics. Drive disaster recovery and improve infrastructure parity. Lead architecture reviews and innovate on developer experience. This role differs by focusing on scaling autonomous AI agents in regulated insurance workflows.

119K–221K
AWSTerraformPython+4 more

Uipath

1d agoDenver, Coloradopayroll

Senior Software Engineer SRE

Design and engineer SRE platform systems using AI while leading cross‑team initiatives. You will identify gaps across teams, design solutions, build, ship, and adopt them, and drive measurable improvements in reliability scalability and performance. This role focuses on ownership accountability and fostering a culture of continuous iteration without relying on generic statements.

160K–210K
pythonawsreact+2 more