Staff Software Engineer - Adaptive Telemetry, Databases
Leads the architecture, delivery, reliability, and operation of large distributed telemetry database systems for Grafana Cloud. The role requires deep systems-design expertise, cloud-native platform experience, strong coding skills, and the ability to influence teams and mentor engineers.
About the job
Responsibilities
- Drive technical strategy and roadmap, defining architectural vision and prioritizing major product and platform improvements.
- Lead end-to-end delivery of large, cross-functional projects, including planning, design, execution, rollout, and long-term operation.
- Own architecture, reliability, performance, and cost for critical systems.
- Define SLOs and SLIs, lead high-severity incident response, conduct blameless post-mortems, and drive systemic fixes and automation.
- Improve observability, alerting, runbooks, capacity planning, and operational automation to reduce toil and MTTR.
- Align stakeholders across Product, Design, and Engineering, negotiate tradeoffs, and remove delivery blockers.
- Mentor senior and mid-level engineers, lead design reviews, and raise engineering standards.
- Communicate technical strategy to non-engineering stakeholders and represent the team in cross-team planning.
Requirements
- Proven delivery and operation of large distributed systems spanning multiple teams, with demonstrated technical leadership and impact.
- Deep systems-design understanding, including latency, consistency, availability, scalability, and cost tradeoffs.
- Experience with cloud-native architectures, microservices, containers, Kubernetes, infrastructure as code, and operational practices.
- Experience defining SLOs/SLIs, capacity planning, performance tuning, and end-to-end reliability ownership.
- Strong coding and technical design skills; experience with Go is useful, while Python, C, C++, Rust, or similar languages translate well.
- Comfort using AI-assisted and agentic development tools in engineering workflows.
- Familiarity with streaming or messaging systems and observability tooling.
- Ability to influence without authority and align cross-functional stakeholders in a remote-first environment.
- Strong written and verbal communication skills.
Compensation and Benefits
- Base compensation in Canada: CAD 186,368–CAD 223,642.
- Benefits include equity, bonus where applicable, and other company benefits.
Skills
Go, Python, C++, Rust, Kubernetes, Microservices, Infrastructure As Code, Distributed Systems, Kafka, Prometheus, Grafana, SLOs, Slis, Cloud-Native Architecture
Similar jobs
Backend Engineering jobsStaff Backend Engineer responsible for evolving Grafana into a scalable, multi-tenant observability application platform. The role requires production SaaS experience, distributed-systems expertise, strong backend coding skills, and familiarity with or willingness to learn Golang.
This Staff Backend Engineer will design and operate production backend services for an AI-native enterprise context system, including ingestion, indexing, retrieval APIs, and agent integrations. The role requires production software experience, cloud-native development, LLM and generative AI expertise, and strong ownership in an ambiguous early-stage environment.
Build large-scale backend infrastructure and data products for Databricks across areas such as log analytics, AI/BI, business semantics, and apps. The role requires staff-level software engineering experience, including 10+ years with Java, Scala, or C++ and expertise in distributed systems and cloud technologies.
Leads the technical vision and architecture for large-scale backend systems powering experimentation, personalization, analytics, and conversion optimization. The role requires 12+ years of software engineering experience, deep distributed-systems expertise, and cross-functional technical leadership.
Leads development of reliable, scalable database connectors and replication technology for enterprise data movement. The role requires strong Java or C/C++ experience, database internals expertise, distributed-systems design skills, and technical leadership through architecture, mentoring, and production investigations.