Dice Talent & Staffing Solutions
Dice Talent & Staffing SolutionsVerified Source

Site Reliability Engineer, SRE & Cloud Infrastructure

180K–265K
Onsite · San Francisco, California
Posted August 13, 2026
payroll

Overview

You will own the reliability and scalability of mission-critical systems for a cutting-edge government contractor in San Francisco. You will design, build, and run large-scale distributed systems with AWS, Kubernetes, and Terraform. You will work with a tight-knit team of engineers to ensure 99.99% uptime for services that matter. This role demands a hands-on engineer who thrives on incident response and automation. You will directly shape the infrastructure strategy from day one.

What You'll Do8

  • 1Design and implement robust monitoring and alerting for Kubernetes clusters and AWS services.
  • 2Build and maintain CI/CD pipelines with Jenkins and GitHub Actions to deploy code securely.
  • 3Automate infrastructure provisioning and configuration using Terraform and Ansible.
  • 4Lead incident response efforts, performing root cause analysis and implementing corrective actions.
  • 5Scale backend services and databases to handle growing traffic and data volumes.
  • 6Optimize system performance through load testing and capacity planning.
  • 7Collaborate with development teams to ensure applications are designed for reliability and scalability.
  • 8Document system architectures and runbooks to improve team knowledge and response time.

Requirements8

  • 15+ years in site reliability, DevOps, or infrastructure engineering roles.
  • 2Hands-on experience with AWS (EC2, S3, RDS, Lambda) and Kubernetes production clusters.
  • 3Proficiency in Terraform, Ansible, or similar IaC tools.
  • 4Solid scripting skills in Python or Go for automation and tooling.
  • 5Experience with monitoring and log management tools like Prometheus, Grafana, and ELK stack.
  • 6Strong understanding of networking, load balancing, and security best practices.
  • 7Bachelor's degree in Computer Science, Engineering, or equivalent work experience.
  • 8US citizenship required due to government contract requirements.

Salary Insight

$180 - $265k per year

Location

Typeonsite
LocationSan Francisco, California

Required Skills

awskubernetesterraformpythonprometheus
Share:

Similar open positions

Explore active roles that match your skills and interests.

Incorporan Inc

Incorporan Inc

13h agoChicago, Illinoispayroll

Site Reliability Engineer, SRE & ML Systems

You will own the reliability, scalability, and performance of production systems at a global scale. You will work alongside platform and data engineering teams, integrating ML workloads with Python and TensorFlow into robust SRE practices. This role demands a blend of software engineering and systems thinking, operating in a hybrid environment with clear ownership of uptime and incident response.

104K–114K
pythontensorflowpytorch+2 more
SRI Tech Solutions

SRI Tech Solutions

1d agoOrlando, Floridapayroll

Lead Site Reliability Engineer SRE

Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Competitive salary
Site Reliability EngineeringCloud InfrastructureKubernetes+3 more
PRIMUS Global Services Inc.

PRIMUS Global Services Inc.

1d agoChicago, Illinoiscontract

Site Reliability Engineer Hybrid

Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

45K–50K
pythonawskubernetes+2 more
Cisco Systems

Cisco Systems

12d agoSan Francisco, Californiapayroll

Staff Site Reliability Engineer SRE Cisco Systems Hybrid

Lead technical roadmap for platform reliability scalability and operational excellence. Drive architecture evolution for cloud and air-gapped environments. Establish SLOs and resilience reviews. Lead major initiatives across Kubernetes databases and networking. Mentor engineers and collaborate with cross-functional teams to deliver secure highly reliable deployments.

187K–268K
KubernetesAWSGCP+8 more
TSQ Systems Inc

TSQ Systems Inc

1d agoPhiladelphia, Pennsylvaniacontract

Site Reliability Engineer (SRE) - Philadelphia, PA

SRE at TSQ Systems Inc in Philadelphia drives reliability, performance, and scalability across mission-critical systems. They own incident response, automate operational workflows, and collaborate with development teams on reliability best practices. This contract role demands a hands-on engineer who thrives in high-stakes environments. They will modernize monitoring with Prometheus and Grafana and reduce toil through Python automation.

Competitive salary
MonitoringCapacity PlanningAutomation+1 more
System One

System One

13h agoDallas, Texascontract

Site Reliability Engineer, Monitoring & Incident Response

Own SRE practices for distribution systems across Cleveland, Pittsburgh, and Dallas. You'll monitor system health, respond to incidents, and drive automation improvements. Work with a team of engineers and collaborate with automation specialists. This role stands out for its focus on incident analysis and proactive monitoring.

58K
srelinuxkubernetes+2 more