Senior Software Engineer SRE
Overview
Design and engineer SRE platform systems using AI while leading cross‑team initiatives. You will identify gaps across teams, design solutions, build, ship, and adopt them, and drive measurable improvements in reliability scalability and performance. This role focuses on ownership accountability and fostering a culture of continuous iteration without relying on generic statements.
What You'll Do16
- 1Design and engineer SRE platform systems and capabilities with cutting‑edge AI treating them as products that other teams depend on in their critical path
- 2Participate in livesite monitoring rotations handle escalations and drive effective mitigations aggressively reducing customer impact through detailed postmortems
- 3Drive availability scalability and performance improvements based on livesite learnings
- 4Generate or codify best practices embed them into built systems not by publishing guidance but by shaping the platforms themselves
- 5Onboard other teams onto new platforms by writing integrations pairing with engineers and removing friction rather than handing off documentation
- 6Ship early seek feedback relentlessly and iterate fast treating every user complaint as a design input
- 7Plan tasks estimate schedule and staff while influencing process improvements and best practices across the engineering organization
- 8Architect and maintain large‑scale distributed commercial applications that have stood the test of time
- 9Maintain complex AI‑powered applications in production
- 10Proficiency in object‑oriented languages such as C# C++ Go or Python with strong computer science fundamentals
- 11Deep understanding of data structures algorithms multithreading synchronization asynchronous patterns and cloud programming
- 12Experience with service‑oriented and microservice architectures HTTP applications and web services development
- 13Familiarity with agile development CI/CD and DevOps practices
- 14Ability to collaborate with globally distributed teams
- 15Experience managing production Kubernetes infrastructure is a plus
- 16Experience with major cloud providers and managed services is required
Requirements6
- 15+ years building ETL pipelines with Spark and Airflow
- 23-5 seasons leading cross‑functional teams in large scale system design
- 36+ years proven track record of architecting and engineering world‑class distributed applications
- 4Proficiency in Python or other object oriented languages
- 5Experience with Kubernetes and cloud platforms such as AWS Azure or GCP
- 6Expertise in designing and deploying AI driven reliability solutions
Salary Insight
$160 - $210k per year
Similar open positions
Explore active roles that match your skills and interests.

SRI Tech Solutions
VerifiedLead Site Reliability Engineer SRE
Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Xoriant Corporation
VerifiedSenior/Staff SRE, AI/ML Platform Infrastructure
Own the reliability of a large-scale AI/ML platform serving millions of requests daily. Kubernetes and Docker are your primary tools. Join a team of 8 SREs supporting 50+ microservices on AWS and GCP. This role focuses on incident command, automation, and platform improvements.
Lambda
VerifiedIT Systems Engineer - Internal Platforms & SRE
Design and deliver software to improve availability scalability reliability and efficiency of internal IT systems. Solve problems related to mission critical services and build automation to prevent problem recurrence. Influence architecture standards for large scale distributed systems. Engage in capacity planning and performance analysis. Be an excellent communicator producing documentation. You have a keen interest in system design and experience with AWS GCP Azure. Think carefully about edge cases failure modes and specific implementations.

SilverSearch, Inc.
VerifiedDevOps SRE AI Platform SilverSearch Inc Philadelphia
We seek a skilled DevOps Site Reliability Engineer to build scalable infrastructure for AI workloads at SilverSearch. This role partners with security and engineering teams to operationalize an AI-driven security platform identifying code vulnerabilities and infrastructure risks. The ideal candidate enjoys solving complex challenges and driving innovation.

Intraedge
VerifiedSenior Lead Site Reliability Engineer Principal Engineer Intraedge Phoenix
We seek a senior technical leader to drive reliability and scalability for cloud-native applications. This role focuses on enhancing observability and operational readiness. The ideal candidate will lead initiatives that improve deployment safety and performance.

Finoit Inc.
VerifiedSenior Software Engineer Data Infrastructure
Design and lead development of scalable data pipelines for AI/ML platforms. Own design and implementation of distributed systems using Python and cloud services. Drive improvements in data quality and visualization. Lead cross-functional teams to deliver high-performance solutions.