GallatinVerified Source

Senior Site Reliability Engineer - Defense Logistics

80K–210K
Onsite · Austin, Texas
Posted August 13, 2026
payroll

Overview

Gallatin seeks a Senior Site Reliability Engineer to own the reliability, availability, and performance of production systems supporting a national security logistics platform, including in classified and air-gapped environments. You will join a team of AWS, Azure, and Kubernetes experts building AI-driven decision systems from factory to foxhole. This role requires an active Secret clearance and a willingness to travel >50% to government sites. You will lead incident response, automate toil, and harden systems to DoD standards, making an immediate impact on mission-critical uptime.

What You'll Do9

  • 1Own reliability, availability, and performance for production systems, including services in classified and air-gapped environments, using Prometheus, Grafana, and Datadog.
  • 2Lead incident response from detection through triage, resolution, and blameless postmortem, implementing fixes that prevent recurrence.
  • 3Design and operate CI/CD pipelines with Terraform and Ansible across cloud and classified networks, enabling safe, frequent deployments.
  • 4Automate manual operational work: deployments, scaling, failover, and credential rotation to eliminate single points of failure.
  • 5Harden systems to meet RMF, STIGs, and ATO requirements without slowing delivery.
  • 6Partner with software engineers to define SLOs/SLIs and embed reliability into service design.
  • 7Work directly with government customers to translate mission constraints into system requirements.
  • 8Write runbooks and architecture docs so the team can operate what you build.
  • 9Travel >50% to customer and government sites for hands-on deployments and support.

Requirements10

  • 1Active Secret clearance required; eligibility for higher clearance.
  • 23-5 years in SRE, DevOps, or production infrastructure roles.
  • 3Strong technical background in Linux systems, networking, and cloud infrastructure (AWS, Azure, or GovCloud).
  • 4Hands-on with infrastructure-as-code (Terraform, Ansible or similar) and container orchestration (Kubernetes, Docker).
  • 5Experience building and maintaining CI/CD pipelines and automated deployment systems.
  • 6Proficiency with monitoring and observability tooling (Prometheus, Grafana, Datadog, ELK or similar).
  • 7Track record of owning production incidents end-to-end.
  • 8Comfort operating in classified, air-gapped, or network-constrained environments.
  • 9Clear communicator for both engineers and government stakeholders.
  • 10Willingness to travel >50% of the time.

Salary Insight

$80 - $210k per year

Location

Typeonsite
LocationAustin, Texas

Required Skills

awsazurekubernetesterraformprometheus
Share:

Similar open positions

Explore active roles that match your skills and interests.

Space Ground System Solutions

1d agoWashington, District of Columbiapayroll

Site Reliability Engineer, Cloud Infrastructure & AWS GovCloud

You will own the reliability and expansion of satellite ground system software on hybrid and private cloud infrastructure for Space Ground System Solutions (SGSS), a Parsons subsidiary. You will architect and automate AWS GovCloud environments, manage Kubernetes and OpenStack platforms, and embed security into CI/CD pipelines. Your team supports developers building systems with government-off-the-shelf, commercial, and open-source software. You will drive a 98% uptime mandate while performing root cause analysis and building infrastructure as code with Terraform and Ansible. This role demands a rare blend of software development and operations expertise, with daily focus on network engineering and security hardening.

89K–135K
DoD Secret security clearanceTerraformAnsible+2 more
Dice Talent & Staffing Solutions

Dice Talent & Staffing Solutions

17h agoSan Francisco, Californiapayroll

Site Reliability Engineer, SRE & Cloud Infrastructure

You will own the reliability and scalability of mission-critical systems for a cutting-edge government contractor in San Francisco. You will design, build, and run large-scale distributed systems with AWS, Kubernetes, and Terraform. You will work with a tight-knit team of engineers to ensure 99.99% uptime for services that matter. This role demands a hands-on engineer who thrives on incident response and automation. You will directly shape the infrastructure strategy from day one.

180K–265K
awskubernetesterraform+2 more

A3 Technology, Inc.

5d agoWashington, District of Columbiapayroll

Senior Windows Server DevOps Engineer (AWS GovCloud)

You will own Windows Server production operations for mission-critical applications in AWS GovCloud. This role centers on ~95% Windows environments, with hands-on support, patching, and recovery during business hours and on-call rotation. You also take ownership of Kubernetes (EKS) clusters, resolving complex platform issues. You lead incident response and root-cause analysis across Windows, cloud, and containers, while designing secure cloud infrastructure. You will automate with Terraform and Ansible, build CI/CD pipelines, and enforce federal security standards.

Competitive salary
Windows Server administrationKubernetesTerraform+5 more
00100 LEIDOS, INC.

00100 LEIDOS, INC.

9d agoWashington, District of Columbiapayroll

Senior Cloud Engineer Leidos Intelligence Group

We seek a senior Cloud Engineer to lead secure, scalable infrastructure modernization at Reston, Virginia. This hybrid role bridges unclassified and classified environments ensuring compliance with DoD mandates while driving automation and cloud-native transformation. The ideal candidate excels in AWS GovCloud and Azure delivering high‑availability solutions across hybrid networks.

92K–167K
AWSAzureTerraform+9 more

palantir

6d agoWashington, District of Columbiapayroll

Forward Deployed Site Reliability Engineer - US Government

Palantir seeks a Forward Deployed Site Reliability Engineer to build and maintain high performance scalable reliable services for US Government on-prem production infrastructure. This role involves traveling to various locations where you become the expert for Palantir’s infrastructure supporting partner teams. The position emphasizes autonomy collaboration and ownership of operations.

Competitive salary
Linux system administrationRHELPhysical server hardware setup+8 more

System One

18d agoWashington, District of Columbiapayroll

Senior Windows Server Cloud DevOps Engineer

Own Windows Server production operations and Kubernetes environment while maintaining availability. Lead incident response and design secure scalable cloud architectures. Develop and manage Infrastructure as Code using Terraform and Ansible. Create CI/CD pipelines and DevSecOps automation. Design monitoring logging and alerting solutions. Enforce security best practices and produce technical documentation. Mentor junior engineers and provide technical leadership in DevOps practices. This is a senior role where you will own systems and drive improvements. The position offers a unique blend of on-premises and cloud expertise.

Competitive salary
windows serverkubernetesterraform+2 more