Senior Site Reliability Engineer
Improves the reliability, observability, and deployment safety of Block’s critical platforms while leading high-severity incident response and on-call operations. Requires strong production reliability experience, incident management skills, and 5+ years of software development experience.
About the job
Responsibilities
- Build and extend platforms to improve system reliability.
- Work on company-wide reliability goals.
- Standardize reliability tools across multiple platforms and organizations.
- Triage, coordinate, and lead stabilization of severity 0–1 incidents.
- Serve as primary on-call, maintaining structured escalation paths and exercising leadership escalation.
- Drive platform-wide reliability improvements, shared operational tooling, and deployment-safety patterns.
- Use AI-driven systems to improve signal detection, reduce alert noise, and accelerate root-cause analysis.
- Design and implement safe deployment patterns, including progressive delivery, automated rollback, and guardrails.
- Participate in primary platform on-call for 12 hours per day, one week every few weeks, supporting critical Tier 0 services.
- Lead incident command, coordinate mitigation, and drive escalation during high-severity events.
Requirements
- Drive complex systems to root cause and take the necessary steps to fix them.
- Demonstrated technical initiative and leadership on backend- or platform-focused projects.
- Familiarity with AI-driven tooling for observability, incident analysis, or automation.
- Experience running production on-call for high-availability systems.
- Strong incident management skills, including structured triage, mitigation under pressure, and blameless postmortems.
- Fluency with CI/CD pipelines, progressive rollout strategies, and rollback automation.
- Monitoring and observability expertise, including alerting for uptime, error rates, latency regression, and resource exhaustion.
- Ability to create and maintain evidence-based maturity assessments using trailing 90-day data windows.
- Experience managing vendors and dependencies, including validated escalation contacts reachable within 5 minutes.
- Curiosity, autonomy, accountability, and a desire to grow as an engineer.
- 5+ years of software development experience.
Technologies
- Kotlin
- Modern Java (11+)
- HTTP, JSON, gRPC, and Protocol Buffers
- MySQL, Vitess, and DynamoDB
- Event-driven architectures
- Datadog
- LaunchDarkly
- Terraform
- Kubernetes
- Istio/Envoy
- Amazon Web Services
Benefits
- Remote work
- Medical insurance
- Flexible time off
- Retirement savings plans
- Modern family planning benefits
Skills
Kotlin, Java, gRPC, Protocol Buffers, MySQL, Vitess, DynamoDB, Datadog, Launchdarkly, Terraform, Kubernetes, Istio, Envoy, Amazon Web Services, CI/CD
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.