
Site Reliability Engineer SRE TechSpace Solutions Inc San Francisco CA
Overview
Own application performance monitoring using AppDynamics and build scalable observability solutions with Splunk. Lead scaling of infrastructure and application monitoring systems. Differentiate by delivering high availability across distributed services.
What You'll Do5
- 1Design Build and Scale observability platforms leveraging AppDynamics and Splunk to monitor application health.
- 2Develop and maintain dashboards alerts and performance reports for production environments.
- 3Lead debugging and resolving critical incidents within 24 hours.
- 4Drive automation of monitoring workflows to reduce manual effort.
- 5Collaborate with cross functional teams to define SLIs SLOs and error budgets.
Requirements5
- 15+ years experience building ETL pipelines with Spark and Airflow
- 23-5 seasons leading cloud infrastructure projects
- 3Expertise in AppDynamics and Splunk
- 4Proven background in infrastructure application and production monitoring
- 5Strong skills in dashboard creation alert configuration and reporting
Salary Insight
Salary not disclosed in listing
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

PRIMUS Global Services Inc.
VerifiedSite Reliability Engineer Hybrid
Own and scale high availability production systems across hybrid environments. Lead development of monitoring solutions and automation scripts. Ensure system resilience and performance optimization. Collaborate with cross functional teams to drive reliability initiatives. Differentiate by focusing on real world SRE practices.

SRI Tech Solutions
VerifiedLead Site Reliability Engineer SRE
Own the design and scaling of high availability cloud infrastructure for a Generative AI platform. Drive reliability and operational excellence while leading technical initiatives. Shape architecture and mentor engineers across SRE practices.

Info Way Solutions
VerifiedSRE Architect Info Way Solutions Seattle
Lead enterprise-wide reliability engineering, observability, and operational excellence initiatives. Drive SRE transformation define reliability strategies establish SLO governance and lead reliability engineering adoption across large-scale enterprise environments.
Uipath
VerifiedSenior Software Engineer SRE
Design and engineer SRE platform systems using AI while leading cross‑team initiatives. You will identify gaps across teams, design solutions, build, ship, and adopt them, and drive measurable improvements in reliability scalability and performance. This role focuses on ownership accountability and fostering a culture of continuous iteration without relying on generic statements.

TSQ Systems Inc
VerifiedSite Reliability Engineer (SRE) - Philadelphia, PA
SRE at TSQ Systems Inc in Philadelphia drives reliability, performance, and scalability across mission-critical systems. They own incident response, automate operational workflows, and collaborate with development teams on reliability best practices. This contract role demands a hands-on engineer who thrives in high-stakes environments. They will modernize monitoring with Prometheus and Grafana and reduce toil through Python automation.

BCforward
VerifiedTechnology & Information Architectures - Performance & Reliability Engineer
Own performance and reliability engineering initiatives scaling across hybrid cloud environments. Deliver high availability solutions leveraging Python AWS React Kubernetes and Spark. Drive observability and automation using industry best practices. Impact multiple global teams.