Staff Software Engineer, Observability
Leads the design, development, deployment, and operation of large-scale observability and distributed systems serving internal developers and customers. Requires staff-level expertise in Kubernetes, programming, query development, cloud infrastructure, and production reliability.
About the job
Responsibilities
- Lead the end-to-end software development lifecycle, including requirements gathering, design, implementation, deployment, operationalization, support, and maintenance.
- Lead feature design and reviews, incorporate stakeholder feedback, and drive consensus.
- Document design decisions and operational knowledge for successful deployment and management.
- Ensure unit, integration, performance, and production-readiness testing.
- Design and build horizontally scalable, multi-component distributed systems.
- Investigate issues methodically, identify root causes, and assess performance and reliability tradeoffs at scale.
- Participate in the on-call rotation.
- Develop expert knowledge of the assigned domain and the Temporal ecosystem.
- Understand internal developer and external customer needs to guide product development and feature design.
- Mentor teammates, participate in design reviews, and contribute to reliable-system design.
Requirements
- Demonstrated experience developing horizontally scalable, resilient, high-performance distributed systems in production.
- Experience designing, implementing, deploying, and supporting large-scale, geographically distributed observability systems, high-throughput data streaming or processing pipelines, or similar systems.
- Expert proficiency in one or more high-level programming languages, preferably Go.
- Expert-level Kubernetes skills.
- Expert-level query development skills, preferably SQL.
- Hands-on experience with a cloud provider, preferably AWS or Google Cloud.
- Thorough understanding of computer architecture, operating systems, and networking.
- Familiarity with monitoring, instrumentation, and infrastructure configuration best practices.
- Strong collaboration, communication, self-direction, and user-focused problem-solving skills.
Skills
Go, Kubernetes, SQL, AWS, GCP, ClickHouse, Prometheus, Grafana, Loki, Thanos, Temporal, Distributed Systems, Data Streaming, Networking, Operating Systems
Similar jobs
Backend Engineering jobsDesigns and operates Internet-scale scanning, DNS, attribution, and data pipelines, with deep ownership of distributed backend systems and production reliability. Requires 10+ years of software engineering experience, strong Go expertise, cloud and streaming infrastructure knowledge, and the ability to mentor engineers.
Staff backend engineer responsible for evolving Grafana into a scalable, multi-tenant observability application platform. The role requires production operations experience, distributed-systems expertise, strong communication, and familiarity with Go or willingness to learn it.
Leads architecture and development of mission-critical backend systems for spacecraft command, telemetry, mission planning, and operations. Requires 8+ years of software development experience, distributed-systems and cloud-native expertise, and technical leadership.
Technical leader for the Metadata team, designing distributed cloud subsystems and leading complex initiatives across discovery, catalog, lineage, and run history services. Requires 8+ years of software engineering experience, backend or systems expertise, cloud infrastructure experience, and strong architecture, reliability, and mentoring skills.
Technical leader for the Metadata team, owning architecture and delivery of distributed backend services for discovery, catalog, lineage, and run history. The role requires 8+ years of software engineering experience, strong cloud and systems expertise, and a record of leading multi-engineer initiatives.