
Site Reliability Engineer, Production Operations
Overview
You own production reliability for critical internal and external applications across Pittsburgh, Cleveland, and Dallas sites. You drive incident response, performance tuning, and automation to maintain high availability. You collaborate with production support and engineering teams to reduce downtime and improve system resilience. This role focuses on proactive monitoring and reliability engineering, not just firefighting.
What You'll Do8
- 1Lead incident response and root cause analysis for production outages using PagerDuty and Datadog.
- 2Build automation scripts in Python to handle routine operational tasks and reduce manual intervention.
- 3Design and implement monitoring dashboards in Grafana to track system health and performance.
- 4Drive performance tuning and capacity planning for Kubernetes clusters and AWS workloads.
- 5Collaborate with engineering teams to improve deployment pipelines using Jenkins and GitLab CI.
- 6Document runbooks and on-call procedures to streamline incident resolution.
- 7Conduct post-incident reviews and implement changes to prevent recurrence.
- 8Scale infrastructure with Terraform to support growing application demands.
Requirements8
- 15+ years in site reliability engineering or similar role, with production support experience.
- 23+ years working with AWS (EC2, S3, RDS, Lambda) and container orchestration via Kubernetes.
- 3Strong Python scripting skills for automation and tooling.
- 4Experience with infrastructure as code using Terraform or CloudFormation.
- 5Hands-on with monitoring tools like Prometheus, Grafana, or Datadog.
- 6Familiarity with CI/CD pipelines using Jenkins or GitLab CI.
- 7Solid understanding of networking, DNS, load balancing, and security best practices.
- 8Excellent troubleshooting and communication skills for on-call rotations.
Salary Insight
Salary not disclosed in listing
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

Blue Rose Technologies LLC
VerifiedSite Reliability Engineer, Production Support
You will own production support for enterprise systems across Dallas, Pittsburgh, or Cleveland, ensuring uptime and rapid incident resolution at scale. You will join a team of engineers focused on system management and analytics, reporting directly to technical leads. This role centers on client-facing reliability, where you propose improvements and drive stability from day one.

System One
VerifiedSite Reliability Engineer, Monitoring & Incident Response
Own SRE practices for distribution systems across Cleveland, Pittsburgh, and Dallas. You'll monitor system health, respond to incidents, and drive automation improvements. Work with a team of engineers and collaborate with automation specialists. This role stands out for its focus on incident analysis and proactive monitoring.

Incorporan Inc
VerifiedSite Reliability Engineer, SRE & ML Systems
You will own the reliability, scalability, and performance of production systems at a global scale. You will work alongside platform and data engineering teams, integrating ML workloads with Python and TensorFlow into robust SRE practices. This role demands a blend of software engineering and systems thinking, operating in a hybrid environment with clear ownership of uptime and incident response.

PRIMUS Global Services Inc.
VerifiedSite Reliability Engineer Hybrid
Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

Dice Talent & Staffing Solutions
VerifiedSite Reliability Engineer, SRE & Cloud Infrastructure
You will own the reliability and scalability of mission-critical systems for a cutting-edge government contractor in San Francisco. You will design, build, and run large-scale distributed systems with AWS, Kubernetes, and Terraform. You will work with a tight-knit team of engineers to ensure 99.99% uptime for services that matter. This role demands a hands-on engineer who thrives on incident response and automation. You will directly shape the infrastructure strategy from day one.

TSQ Systems Inc
VerifiedSite Reliability Engineer (SRE) - Philadelphia, PA
SRE at TSQ Systems Inc in Philadelphia drives reliability, performance, and scalability across mission-critical systems. They own incident response, automate operational workflows, and collaborate with development teams on reliability best practices. This contract role demands a hands-on engineer who thrives in high-stakes environments. They will modernize monitoring with Prometheus and Grafana and reduce toil through Python automation.