# Senior Site Reliability Engineer

**Company:** [Square](https://hotfix.jobs/companies/square)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Kotlin, Java, gRPC, Protocol Buffers, MySQL, Vitess, DynamoDB, Datadog, Launchdarkly, Terraform, Kubernetes, Istio, Envoy, Amazon Web Services, CI/CD
**Posted:** 2026-08-03

> Improves the reliability, observability, and deployment safety of Block’s critical platforms while leading high-severity incident response and on-call operations. Requires strong production reliability experience, incident management skills, and 5+ years of software development experience.

## Job Description

## Responsibilities
- Build and extend platforms to improve system reliability.
- Work on company-wide reliability goals.
- Standardize reliability tools across multiple platforms and organizations.
- Triage, coordinate, and lead stabilization of severity 0–1 incidents.
- Serve as primary on-call, maintaining structured escalation paths and exercising leadership escalation.
- Drive platform-wide reliability improvements, shared operational tooling, and deployment-safety patterns.
- Use AI-driven systems to improve signal detection, reduce alert noise, and accelerate root-cause analysis.
- Design and implement safe deployment patterns, including progressive delivery, automated rollback, and guardrails.
- Participate in primary platform on-call for 12 hours per day, one week every few weeks, supporting critical Tier 0 services.
- Lead incident command, coordinate mitigation, and drive escalation during high-severity events.

## Requirements
- Drive complex systems to root cause and take the necessary steps to fix them.
- Demonstrated technical initiative and leadership on backend- or platform-focused projects.
- Familiarity with AI-driven tooling for observability, incident analysis, or automation.
- Experience running production on-call for high-availability systems.
- Strong incident management skills, including structured triage, mitigation under pressure, and blameless postmortems.
- Fluency with CI/CD pipelines, progressive rollout strategies, and rollback automation.
- Monitoring and observability expertise, including alerting for uptime, error rates, latency regression, and resource exhaustion.
- Ability to create and maintain evidence-based maturity assessments using trailing 90-day data windows.
- Experience managing vendors and dependencies, including validated escalation contacts reachable within 5 minutes.
- Curiosity, autonomy, accountability, and a desire to grow as an engineer.
- 5+ years of software development experience.

## Technologies
- Kotlin
- Modern Java (11+)
- HTTP, JSON, gRPC, and Protocol Buffers
- MySQL, Vitess, and DynamoDB
- Event-driven architectures
- Datadog
- LaunchDarkly
- Terraform
- Kubernetes
- Istio/Envoy
- Amazon Web Services

## Benefits
- Remote work
- Medical insurance
- Flexible time off
- Retirement savings plans
- Modern family planning benefits

## Similar jobs

- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior Production Engineer](https://hotfix.jobs/jobs/9ee5879e-954d-4681-ad0c-816d7151f874) - Lightspark - Remote - $200k – $238k/yr
- [Senior Cloud Software Engineer - Efficiency Engineering](https://hotfix.jobs/jobs/6d7b2812-de3f-4dfa-8586-9010e1e594e7) - Clickhouse - Remote
- [Senior Cloud Software Engineer - Efficiency Engineering](https://hotfix.jobs/jobs/56b82d25-159d-41f5-94fb-fb1f71dd6a5c) - Clickhouse - Remote
- [Release Engineer - Data Plane Internal Tooling and Productivity](https://hotfix.jobs/jobs/ee184847-0b76-42f4-8c1f-157f199626d3) - Clickhouse - Remote

**Apply:** https://hotfix.jobs/jobs/dd3adef6-726f-4a9c-a29a-3c5d5b382b97
**Canonical:** https://hotfix.jobs/jobs/dd3adef6-726f-4a9c-a29a-3c5d5b382b97