Senior Software Engineer, Observability
Senior software engineer responsible for designing, deploying, and operating large-scale observability and distributed systems. The role requires expert Kubernetes and query-development skills, proficiency in a high-level language such as Go, and cloud experience.
About the job
Responsibilities
- Lead the end-to-end software development lifecycle, including requirements gathering, design, implementation, deployment, operationalization, support, and maintenance.
- Design and build multi-component, horizontally scalable distributed systems.
- Formulate feature designs, incorporate stakeholder feedback, and drive consensus.
- Document design decisions and operational knowledge.
- Provide unit, integration, performance, and production-readiness testing.
- Investigate incidents methodically and identify root causes.
- Evaluate performance and reliability tradeoffs at scale.
- Participate in the team on-call rotation.
- Understand internal developer and external customer needs and apply that knowledge to product and feature design.
- Participate in design reviews and share principles for building reliable systems at scale.
Requirements
- Demonstrated ability to develop resilient, high-performance distributed systems in production.
- Experience designing, implementing, deploying, and supporting large-scale, geographically distributed observability systems or high-throughput data streaming and processing pipelines.
- Expert proficiency in one or more high-level programming languages, preferably Go.
- Expert-level Kubernetes skills.
- Expert-level query development skills, preferably SQL.
- Hands-on experience with a cloud provider, preferably AWS or Google Cloud.
- Thorough understanding of computer architecture, operating systems, and networking.
- Familiarity with monitoring, instrumentation, and infrastructure configuration best practices.
- User-first mindset, pragmatic decision-making, self-direction, and strong collaboration and communication skills.
Compensation
- Annual salary: $140,800–$231,000.
Skills
Go, Kubernetes, SQL, AWS, GCP, ClickHouse, Prometheus, Grafana, Loki, Thanos, Temporal, Distributed Systems, Data Streaming, Networking, Operating Systems
Similar jobs
Backend Engineering jobsBuild and operate foundational infrastructure for Temporal Cloud, including multi-tenant control-plane systems, database scaling, HA/DR, deployment automation, and CI/CD. Requires senior backend engineering experience with cloud infrastructure, relational databases, high availability, and systems operating at scale.
Build and scale high-performance backend services and microservices for a platform serving millions of users. The role requires strong Node.js expertise, distributed-systems experience, and technical leadership, with Java, cloud, and AI-assisted development experience preferred.
Senior backend engineer building and scaling Node.js microservices, APIs, and distributed systems for a high-volume platform. Requires 6+ years of backend development experience, technical leadership, and expertise in performance, reliability, and modern development practices.
Owns the backend service layer for account activation and onboarding, building resilient Node.js and TypeScript workflows for identity verification, funding, and third-party integrations. Requires 6+ years of backend experience, event-driven architecture expertise, and production skills with cloud infrastructure and Kubernetes.
Leads development and operation of production Rust backend services for account creation, identity verification, and KYC/CIP decisioning. The role requires 6+ years of software development experience, strong async Rust expertise, secure data-handling experience, and the ability to drive cross-team technical design.