
Platform Architect, Observability & Grafana
Overview
You will own the design and delivery of observability platforms for enterprise customers, scaling across Grafana and related stacks. You will lead client requirement discussions independently, translating functional and security needs into robust solutions. Working with a team of engineers, you will ensure production-grade implementations with a customer-first mindset. This role demands hands-on troubleshooting and the ability to drive architecture decisions end-to-end.
What You'll Do8
- 1Drive requirements discussions with customers independently, covering functional, non-functional, and security aspects.
- 2Design end-to-end observability solutions using Grafana, including dashboards, alerts, and data sources.
- 3Build platform components with a focus on scalability and reliability, integrating Prometheus, Loki, and Tempo.
- 4Lead the implementation of monitoring pipelines, ensuring seamless data flow and visualization.
- 5Debug production issues in distributed systems, utilizing Grafana and ELK stack for root cause analysis.
- 6Collaborate with engineering teams to embed observability practices into the platform lifecycle.
- 7Own the technical roadmap for observability, evaluating new tools and open-source integrations.
- 8Provide technical guidance and mentorship to junior engineers, fostering a culture of quality.
Requirements8
- 112+ years in platform architecture or software engineering, with a focus on observability.
- 22+ years of production experience with Grafana, including dashboard design and alerting.
- 3Hands-on experience with Prometheus, Loki, Tempo, or similar monitoring tools.
- 4Proven ability to drive customer requirement discussions and propose solutions independently.
- 5Strong troubleshooting and debugging skills in cloud-native environments (AWS, Kubernetes).
- 6Experience building platforms with end-to-end customer focus, from concept to deployment.
- 7Solid understanding of security requirements and how they impact platform design.
- 8Excellent communication skills, able to articulate technical concepts to diverse audiences.
Salary Insight
Salary not disclosed in listing
Similar open positions
Explore active roles that match your skills and interests.

SmallArc, Inc
VerifiedPlatform Architect Senior Observability Platform Architect
Lead design and delivery of observability platforms for large scale systems. Own technical strategy and drive cross functional collaboration. Shape product roadmap based on insights from monitoring and logging data. Differentiate by delivering scalable solutions that improve system reliability and performance.
VictoriaMetrics
VerifiedSenior Solution Engineer, Observability & Linux
You will architect and demonstrate observability solutions for Fortune 500 companies using VictoriaMetrics, the open-source time-series database. You will bridge the gap between our Enterprise platform and customers' complex Linux infrastructures, handling the most technical engagements and partnering with our development team. This role blends deep technical consulting with hands-on POCs, and you will influence product roadmap with direct feedback.
Xora Innovation
VerifiedSenior Platform Engineer Cloud Xora Portfolio Company
Elemynt seeks a Senior Platform Engineer to architect and maintain the cloud infrastructure and core platform services that power our AI-driven materials platform. You will own networking identity access Kubernetes and infrastructure-as-code ensuring reproducibility and security across deployments. This hands-on role involves building scalable services observability pipelines and deployment processes while standing out in a fast-paced environment.
Virtasant
VerifiedSenior Platform Engineer Cloud Infrastructure Go
We seek a senior engineer to design and operate cloud-native platforms using Go. This role involves building Kubernetes clusters managing networking workload isolation and multi-region topologies implementing service mesh features like mTLS and traffic management optimizing containerized workloads developing Go services and middleware writing unit and integration tests leading incident response defining SLOs and building alerting systems managing infrastructure as code and CI/CD pipelines planning cloud migrations and ensuring observability through metrics and tracing collaborating with product and security teams.

BCforward
VerifiedTechnology & Information Architectures - Performance & Reliability Engineer
Own performance and reliability engineering initiatives scaling across hybrid cloud environments. Deliver high availability solutions leveraging Python AWS React Kubernetes and Spark. Drive observability and automation using industry best practices. Impact multiple global teams.
Cambia Health Solutions
VerifiedPlatform Infrastructure Engineer (AWS, Terraform)
You will own the design and implementation of secure, scalable cloud foundations and shared services across AWS and Azure in a high-uptime, high-security environment. You will lead modernization of platform operations through Infrastructure as Code, CI/CD, and observability standards, influencing leadership decisions and mentoring engineers. This role combines hands-on delivery with coaching, championing responsible automation including AI-accelerated workflows. You will drive enterprise best practices and reference architectures for hybrid integrations, reducing toil and improving reliability.