
SRE Architect Info Way Solutions Seattle
Overview
Lead enterprise-wide reliability engineering, observability, and operational excellence initiatives. Drive SRE transformation define reliability strategies establish SLO governance and lead reliability engineering adoption across large-scale enterprise environments.
What You'll Do11
- 1Design Build and Scale observability platforms using Prometheus and Grafana
- 2Own and Operate distributed systems ensuring high availability and performance
- 3Define and enforce Service Level Objectives and error budgets
- 4Lead cross-functional teams to implement reliability best practices
- 5Ship and maintain robust incident response playbooks
- 6Drive automation of monitoring alerting and deployment processes
- 7Debug complex production issues and resolve root causes
- 8Drive cultural change toward reliability and ownership
- 9Scale solutions across multiple cloud environments using AWS and Azure
- 10Collaborate with product managers to align reliability with business goals
- 11Mentor junior engineers and foster knowledge sharing
Requirements10
- 15+ years building SRE architectures with Kubernetes and Docker
- 23-5 seasons managing complex distributed systems
- 3AWS Certified Solutions Architect professional level
- 4CPC Card certification preferred
- 55+ years developing and enforcing SLOs and SLIs
- 6Proven track record implementing observability stacks
- 7Experience with CI/CD pipelines using Jenkins or GitLab CI
- 8Strong scripting skills in Python or Shell
- 9Understanding of container orchestration with K8s
- 10Familiarity with chaos engineering principles
Salary Insight
$125 - $135k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

PRIMUS Global Services Inc.
VerifiedSite Reliability Engineer Hybrid
Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

SRI Tech Solutions
VerifiedLead Site Reliability Engineer SRE
Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.
Uipath
VerifiedSenior Software Engineer SRE
Design and engineer SRE platform systems using AI while leading cross‑team initiatives. You will identify gaps across teams, design solutions, build, ship, and adopt them, and drive measurable improvements in reliability scalability and performance. This role focuses on ownership accountability and fostering a culture of continuous iteration without relying on generic statements.

Xoriant Corporation
VerifiedSenior/Staff SRE, AI/ML Platform Infrastructure
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.

Yoh - A Day & Zimmerman Company
VerifiedSRE and Platform Engineer Yoh - A Day & Zimmerman Company
Lead technical reliability performance security observability operations for enterprise platforms and modern web applications. Provide technical leadership as a senior engineer technical lead or architect. Drive solutions mentor engineers influence technical direction.
Cisco Systems
VerifiedStaff Site Reliability Engineer SRE Cisco Systems Hybrid
Lead technical roadmap for platform reliability scalability and operational excellence. Drive architecture evolution for cloud and air-gapped environments. Establish SLOs and resilience reviews. Lead major initiatives across Kubernetes databases and networking. Mentor engineers and collaborate with cross-functional teams to deliver secure highly reliable deployments.