
SRE and Platform Engineer Yoh - A Day & Zimmerman Company
Overview
Lead technical reliability performance security observability operations for enterprise platforms and modern web applications. Provide technical leadership as a senior engineer technical lead or architect. Drive solutions mentor engineers influence technical direction.
What You'll Do11
- 1Design Build and Scale observability pipelines using Prometheus Grafana and ELK stack
- 2Implement CI CD pipelines with Jenkins GitHub Actions to automate deployments
- 3Own incident response and root cause analysis for production incidents
- 4Mentor junior engineers through code reviews and pair programming sessions
- 5Scale Kubernetes clusters on AWS using EKS and manage cluster lifecycle
- 6Develop and enforce platform standards for security compliance and governance
- 7Drive automation of infrastructure provisioning with Terraform and CloudFormation
- 8Collaborate with product teams to define platform capabilities and roadmap
- 9Optimize database queries and storage solutions for high throughput workloads
- 10Establish monitoring dashboards for real time system health metrics
- 11Champion best practices for disaster recovery and business continuity planning
Requirements11
- 15+ years building ETL pipelines with Spark and Airflow
- 23-5 seasons leading cloud native platform engineering initiatives
- 3AWS Certified Solutions Architect Professional
- 4Python proficiency for scripting automation tasks
- 5React experience for internal tooling development
- 6Kubernetes expertise for container orchestration
- 7CI/CD pipeline design using Jenkins GitHub Actions
- 8Terraform infrastructure as code implementation
- 9Prometheus monitoring and alerting configuration
- 10ELK Stack log aggregation and visualization
- 11Security compliance frameworks like CIS benchmarks
Salary Insight
Salary not disclosed in listing
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.
Uipath
VerifiedSenior Software Engineer SRE
Design and engineer SRE platform systems using AI while leading cross‑team initiatives. You will identify gaps across teams, design solutions, build, ship, and adopt them, and drive measurable improvements in reliability scalability and performance. This role focuses on ownership accountability and fostering a culture of continuous iteration without relying on generic statements.

Info Way Solutions
VerifiedSRE Architect Info Way Solutions Seattle
Lead enterprise-wide reliability engineering, observability, and operational excellence initiatives. Drive SRE transformation define reliability strategies establish SLO governance and lead reliability engineering adoption across large-scale enterprise environments.

PRIMUS Global Services Inc.
VerifiedSite Reliability Engineer Hybrid
Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

SRI Tech Solutions
VerifiedLead Site Reliability Engineer SRE
Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Xoriant Corporation
VerifiedSenior/Staff SRE, AI/ML Platform Infrastructure
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.

BCforward
VerifiedTechnology & Information Architectures - Performance & Reliability Engineer
Own performance and reliability engineering initiatives scaling across hybrid cloud environments. Deliver high availability solutions leveraging Python AWS React Kubernetes and Spark. Drive observability and automation using industry best practices. Impact multiple global teams.