
AI SRE Engineer, AI/ML Platforms & Reliability
Overview
You will own the reliability, scalability, and observability of AI/ML production systems at JC CORPORATIONS. You will design and maintain highly reliable environments, monitor performance and latency, and build SLIs/SLOs with error budget tracking. You will work with Kubernetes, Terraform, and Prometheus to automate and optimize infrastructure. This role combines SRE best practices with AI/ML infrastructure to ensure seamless operation of models and pipelines.
What You'll Do8
- 1Design and maintain highly reliable and scalable AI/ML production environments on AWS.
- 2Monitor availability, performance, latency, and reliability of AI/ML applications using Prometheus and Grafana.
- 3Build and maintain SLIs, SLOs, and SLAs with error budget tracking and alerting.
- 4Automate incident response and runbook execution using Python and Terraform.
- 5Optimize model serving infrastructure with Kubernetes and KServe for low latency and high throughput.
- 6Drive capacity planning and performance tuning for GPU clusters and model inference.
- 7Collaborate with data scientists to integrate MLflow for model lifecycle management and deployment.
- 8Implement chaos engineering and load testing to validate system resilience and scalability.
Requirements10
- 15+ years in SRE, DevOps, or AI/ML infrastructure roles.
- 2Strong experience with Kubernetes and container orchestration in production.
- 3Proficiency in Python and at least one infrastructure-as-code tool like Terraform.
- 4Hands-on experience with monitoring and observability tools: Prometheus, Grafana, Datadog, or similar.
- 5Deep understanding of SLI/SLO concepts and error budget policies.
- 6Experience with AWS or other cloud providers, and GPU computing environments.
- 7Familiarity with model serving frameworks like KServe, TorchServe, or TensorFlow Serving.
- 8Experience with MLflow or similar MLOps tools for model lifecycle management.
- 9Strong scripting skills in Bash and Python for automation.
- 10Excellent troubleshooting and incident management skills in production environments.
Salary Insight
$146 - $208k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

Xoriant Corporation
VerifiedSenior/Staff SRE, AI/ML Platform Infrastructure
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.

JC CORPORATIONS
VerifiedDevOps AI Engineer
You will own the AI/ML infrastructure and deployment pipelines for JC CORPORATIONS' AI applications. You will design, build, and automate scalable systems that deliver models to production. Your stack spans AWS, Azure, Google Cloud, Kubernetes, and MLOps tools. You will collaborate with data scientists and software engineers to streamline model deployment. This role combines deep DevOps expertise with AI/ML platform engineering, driving the backbone of our AI services.

Incorporan Inc
VerifiedSite Reliability Engineer, SRE & ML Systems
You will own the reliability, scalability, and performance of production systems at a global scale. You will work alongside platform and data engineering teams, integrating ML workloads with Python and TensorFlow into robust SRE practices. This role demands a blend of software engineering and systems thinking, operating in a hybrid environment with clear ownership of uptime and incident response.

SRI Tech Solutions
VerifiedLead Site Reliability Engineer SRE
Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.
Hi Marley
VerifiedSr. Software Engineer II DevOps
Lead design and operation of cloud infrastructure on AWS to support core SaaS platform and agentic AI services. Build AI/ML infrastructure and monitoring for LLM-powered services. Establish IaC standards using Terraform. Implement observability beyond availability. Support big data pipelines, warehousing, and analytics. Drive disaster recovery and improve infrastructure parity. Lead architecture reviews and innovate on developer experience. This role differs by focusing on scaling autonomous AI agents in regulated insurance workflows.
Uipath
VerifiedSenior Software Engineer SRE
Design and engineer SRE platform systems using AI while leading cross‑team initiatives. You will identify gaps across teams, design solutions, build, ship, and adopt them, and drive measurable improvements in reliability scalability and performance. This role focuses on ownership accountability and fostering a culture of continuous iteration without relying on generic statements.