System One
System OneVerified Source

Site Reliability Engineer, Monitoring & Incident Response

58K
Onsite · Dallas, Texas
Posted August 13, 2026
contract

Overview

Own SRE practices for distribution systems across Cleveland, Pittsburgh, and Dallas. You'll monitor system health, respond to incidents, and drive automation improvements. Work with a team of engineers and collaborate with automation specialists. This role stands out for its focus on incident analysis and proactive monitoring.

What You'll Do6

  • 1Monitor distribution systems and alert stakeholders to potential issues.
  • 2Troubleshoot system issues during on-call rotations.
  • 3Manage and track incidents from detection to resolution, including outages.
  • 4Lead incident analysis meetings to identify root causes and follow-up actions.
  • 5Identify automation opportunities and coordinate with automation specialists to implement fixes.
  • 6Watch over in-scope and related systems to ensure stability and performance.

Requirements7

  • 15+ years in SRE or related field.
  • 2Strong experience with Linux and Kubernetes.
  • 3Hands-on experience with Python and Bash scripting.
  • 4Familiarity with AWS or GCP cloud environments.
  • 5Experience with monitoring tools like Prometheus and Grafana.
  • 6Proven ability to lead incident response and postmortems.
  • 7Excellent communication and collaboration skills.

Salary Insight

$58k per year

Location

Typeonsite
LocationDallas, Texas

Required Skills

srelinuxkubernetespythonaws
Share:

Similar open positions

Explore active roles that match your skills and interests.

System One

System One

12h agoPittsburgh, Pennsylvaniapayroll

Site Reliability Engineer, Production Operations

You own production reliability for critical internal and external applications across Pittsburgh, Cleveland, and Dallas sites. You drive incident response, performance tuning, and automation to maintain high availability. You collaborate with production support and engineering teams to reduce downtime and improve system resilience. This role focuses on proactive monitoring and reliability engineering, not just firefighting.

Competitive salary
pythonawskubernetes+2 more
PRIMUS Global Services Inc.

PRIMUS Global Services Inc.

1d agoChicago, Illinoiscontract

Site Reliability Engineer Hybrid

Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

45K–50K
pythonawskubernetes+2 more
TSQ Systems Inc

TSQ Systems Inc

1d agoPhiladelphia, Pennsylvaniacontract

Site Reliability Engineer (SRE) - Philadelphia, PA

SRE at TSQ Systems Inc in Philadelphia drives reliability, performance, and scalability across mission-critical systems. They own incident response, automate operational workflows, and collaborate with development teams on reliability best practices. This contract role demands a hands-on engineer who thrives in high-stakes environments. They will modernize monitoring with Prometheus and Grafana and reduce toil through Python automation.

Competitive salary
MonitoringCapacity PlanningAutomation+1 more
Blue Rose Technologies LLC

Blue Rose Technologies LLC

12h agoDallas, Texaspayroll

Site Reliability Engineer, Production Support

You will own production support for enterprise systems across Dallas, Pittsburgh, or Cleveland, ensuring uptime and rapid incident resolution at scale. You will join a team of engineers focused on system management and analytics, reporting directly to technical leads. This role centers on client-facing reliability, where you propose improvements and drive stability from day one.

Competitive salary
pythonsplunknewrelic+2 more
Dice Talent & Staffing Solutions

Dice Talent & Staffing Solutions

12h agoSan Francisco, Californiapayroll

Site Reliability Engineer, SRE & Cloud Infrastructure

You will own the reliability and scalability of mission-critical systems for a cutting-edge government contractor in San Francisco. You will design, build, and run large-scale distributed systems with AWS, Kubernetes, and Terraform. You will work with a tight-knit team of engineers to ensure 99.99% uptime for services that matter. This role demands a hands-on engineer who thrives on incident response and automation. You will directly shape the infrastructure strategy from day one.

180K–265K
awskubernetesterraform+2 more
Incorporan Inc

Incorporan Inc

12h agoChicago, Illinoispayroll

Site Reliability Engineer, SRE & ML Systems

You will own the reliability, scalability, and performance of production systems at a global scale. You will work alongside platform and data engineering teams, integrating ML workloads with Python and TensorFlow into robust SRE practices. This role demands a blend of software engineering and systems thinking, operating in a hybrid environment with clear ownership of uptime and incident response.

104K–114K
pythontensorflowpytorch+2 more