Infrastructure Engineer III, Monitoring & Observability
Overview
Owns enterprise observability platforms at scale to ensure 24/7 service availability. Leads integration of monitoring tools across cloud native architectures and microservices. Collaborates with SRE, engineering, and security teams to implement reliable, secure, and automated monitoring solutions. Focuses on performance, incident response, and root cause analysis to improve service quality. This role is unique in combining hands-on tooling with cross-team governance and security considerations.
What You'll Do5
- 1Build and maintain Splunk and Dynatrace configurations to support scalable monitoring and rapid incident response
- 2Design and implement Grafana dashboards and integrate data sources such as Prometheus and Elasticsearch to drive observability insights
- 3Lead performance tuning and clustering optimization for Splunk to improve search, indexing, and data retention
- 4Collaborate with cloud, Kubernetes, and microservices teams to instrument services for proactive monitoring and AI-driven root cause analysis
- 5Own automation and scripting for observability workflows using Python and Terraform to reduce toil and improve repeatability
Requirements5
- 15+ years building and maintaining monitoring and observability in large-scale environments
- 2Hands-on experience with Splunk indexers, forwarders, search heads, clustering, and performance optimization
- 3Proficiency with Grafana dashboard development and data sources such as Prometheus and Elasticsearch
- 4Strong experience with Dynatrace OneAgent deployment and service monitoring
- 5Scripting and automation experience using Python, Shell, REST APIs, Terraform, and Ansible
Salary Insight
$104 - $175k per year
Location
Required Skills
Similar open positions
Explore active roles that match your skills and interests.

Apex Systems
VerifiedInfrastructure Engineer Principal Apex Systems Remote
Design implement administration enhancement and operational support of enterprise monitoring ecosystem. Serve as subject matter expert for enterprise monitoring technologies including Dynatrace and related platforms. Own observability initiatives at scale.

Arkhya Tech
VerifiedPlatform Architect, Observability & Grafana
You will own the design and delivery of observability platforms for enterprise customers, scaling across Grafana and related stacks. You will lead client requirement discussions independently, translating functional and security needs into robust solutions. Working with a team of engineers, you will ensure production-grade implementations with a customer-first mindset. This role demands hands-on troubleshooting and the ability to drive architecture decisions end-to-end.

Dolphin Solutions Inc
VerifiedSenior Observability Engineer, Elastic Cloud & Monitoring
You will own the observability strategy for a high-traffic distributed platform, analyzing the current Elastic Cloud deployment and closing monitoring gaps. You will work with platform and application teams to implement end-to-end telemetry across cloud-native services. This contract role in McLean, VA requires deep Elastic expertise and a track record of scaling monitoring solutions.

Recruitment.ai
VerifiedPrincipal Observability Platform Engineer, GPU & AI Infrastructure
You will own the architectural roadmap for our observability platform, covering GPU clusters, AI workloads, and the underlying infrastructure. You will define and build the platform that provides deep visibility into Kubernetes, Prometheus, and Grafana stacks, and lead a team of engineers to raise the engineering bar. This role shapes the technical direction and scales infrastructure ahead of business needs, directly impacting the reliability of AI and ML pipelines.
Manifest Solutions
VerifiedSenior Enterprise Monitoring Engineer, Dynatrace
You will own the enterprise monitoring and observability platform using Dynatrace, supporting critical applications across on-prem, cloud, and hybrid environments. You will design monitoring standards, tune alerting, and drive incident response for a large-scale infrastructure. The role sits within the infrastructure engineering team, working with application owners, cloud architects, and operations. You will lead onboarding and automation, reducing false positives and improving visibility. This role stands out due to its impact on regulated utilities and NERC/CIP compliance.

SmallArc, Inc
VerifiedPlatform Architect Senior Observability Platform Architect
Lead design and delivery of observability platforms for large scale systems. Own technical strategy and drive cross functional collaboration. Shape product roadmap based on insights from monitoring and logging data. Differentiate by delivering scalable solutions that improve system reliability and performance.