Sr. Software Engineer II DevOps
Overview
Lead design and operation of cloud infrastructure on AWS to support core SaaS platform and agentic AI services. Build AI/ML infrastructure and monitoring for LLM-powered services. Establish IaC standards using Terraform. Implement observability beyond availability. Support big data pipelines, warehousing, and analytics. Drive disaster recovery and improve infrastructure parity. Lead architecture reviews and innovate on developer experience. This role differs by focusing on scaling autonomous AI agents in regulated insurance workflows.
What You'll Do11
- 1Design and operate cloud infrastructure on AWS supporting core SaaS and agentic AI services ensuring reliability scalability and cost efficiency
- 2Build and maintain AI/ML infrastructure and monitoring for LLM-powered agentic services
- 3Establish IaC standards with Terraform defining environment parity drift detection and automated compliance validation
- 4Implement observability beyond availability including data integrity monitoring SLO frameworks and automated regression detection
- 5Develop deployment automation with pre-deployment verification migration script validation and codified rollback procedures
- 6Support big data infrastructure such as data pipelines warehousing Redshift and analytics tooling for reporting BI and AI training
- 7Implement security and compliance controls for AI workloads in regulated carrier environments covering audit logging access governance and configuration management
- 8Drive environment parity across all infrastructure with automated drift detection and remediation
- 9Improve disaster recovery capabilities through documented DR procedures defined RTOs RPOs and tested recovery runbooks
- 10Lead architecture reviews for new services integrations and AI agent deployments partnering with engineering product and security
- 11Innovate on developer experience reducing friction in testing environments CI/CD pipelines and local development workflows
Requirements18
- 16+ years of DevOps SRE or Platform Engineering experience
- 22+ years building or operating AI/ML infrastructure including model serving inference LLM orchestration or agentic systems
- 3Bachelor's degree in Computer Science Engineering or equivalent
- 4Experience building infrastructure for traditional and AI/ML workloads at a SaaS company
- 5Strong Terraform skills managing state modules and multi-environment configurations
- 6Deep understanding of AWS services such as ECS Lambda SageMaker Bedrock S3 DynamoDB Redshift
- 7Proficiency in infrastructure-as-code with Terraform and knowledge of state management modules
- 8Experience with data pipelines warehousing ETL ELT and analytics at scale
- 9Understanding of observability beyond dashboards focusing on data integrity SLOs error budgets and silent failure detection
- 10Compliance-sensitive environment experience with audit trails access governance and change management
- 11Comfortable operating in fast-moving environments with evolving AI capabilities and regulatory implications
- 12Strong communication skills with engineering and non-technical stakeholders
- 13Track record of leading cross-team technical initiatives and mentoring engineers
- 14Proficiency in Python Go TypeScript or similar language
- 15Experience with container orchestration ECS EKS monitoring platforms Datadog CloudWatch
- 16Familiarity with data infrastructure tools like Redshift Airflow Dagster
- 17Experience in regulated industries such as insurance financial services or healthcare
- 18Curiosity about AI and emerging technologies with judgment to apply them responsibly
Salary Insight
$119 - $221k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

VDart, Inc.
VerifiedSenior AI DevOps Engineer AI Ops Platform Engineering
Design and implement scalable AI-driven CI/CD pipelines using Python AWS React Kubernetes Spark Airflow. Own automated workflows that accelerate software delivery while integrating generative AI and model context protocol solutions.

Xoriant Corporation
VerifiedSenior/Staff SRE, AI/ML Platform Infrastructure
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.
Ntt-Data-Aivista
VerifiedMember of Technical Staff Applied AI Engineering
Design and deploy agentic systems that empower enterprise AI products to operate reliably at scale. Own the end-to-end lifecycle from prototype to production while collaborating with cross‑functional teams to solve complex governance challenges. This builder role offers ownership of outcomes and drives impact in a fast‑moving environment.
TetraScience
VerifiedLead Software Platform Engineer MLOps TetraScience
We seek a Lead Software Platform Engineer at the intersection of distributed systems and MLOps to own and scale AI and data infrastructure for customers. This role involves architecting cloud-based services and MLOps platforms enabling production-grade AI workflows for pharmaceutical clients while ensuring security and compliance. The ideal candidate thrives in regulated environments and drives technical strategy for multi-tenant AI products.
SS&C Technologies
VerifiedDirector AI Platform Engineering SS&C Technologies
SS&C Advent`s financial solutions are evolving into an Agentic AI Framework. We seek a Director to lead the design and scaling of a production-grade multi-tenant platform enabling AI agent workflows across finance and healthcare. This hybrid role blends flexibility with strategic impact.
grailbio
VerifiedEnterprise AI Infrastructure Engineer at GRAIL
Senior Staff Software Development Engineer leads design and scaling of enterprise AI platform using AWS and Kubernetes. This role drives technical excellence in cloud infrastructure and AI governance while mentoring teams and shaping enterprise AI strategy. Based in Sunnyvale with potential visits to Menlo Park and Durham NC.