Skip to content
SquareSquare

Senior Site Reliability Engineer

Improves the reliability, observability, and deployment safety of Block’s critical platforms while leading high-severity incident response and on-call operations. Requires strong production reliability experience, incident management skills, and 5+ years of software development experience.

About the job

Responsibilities

  • Build and extend platforms to improve system reliability.
  • Work on company-wide reliability goals.
  • Standardize reliability tools across multiple platforms and organizations.
  • Triage, coordinate, and lead stabilization of severity 0–1 incidents.
  • Serve as primary on-call, maintaining structured escalation paths and exercising leadership escalation.
  • Drive platform-wide reliability improvements, shared operational tooling, and deployment-safety patterns.
  • Use AI-driven systems to improve signal detection, reduce alert noise, and accelerate root-cause analysis.
  • Design and implement safe deployment patterns, including progressive delivery, automated rollback, and guardrails.
  • Participate in primary platform on-call for 12 hours per day, one week every few weeks, supporting critical Tier 0 services.
  • Lead incident command, coordinate mitigation, and drive escalation during high-severity events.

Requirements

  • Drive complex systems to root cause and take the necessary steps to fix them.
  • Demonstrated technical initiative and leadership on backend- or platform-focused projects.
  • Familiarity with AI-driven tooling for observability, incident analysis, or automation.
  • Experience running production on-call for high-availability systems.
  • Strong incident management skills, including structured triage, mitigation under pressure, and blameless postmortems.
  • Fluency with CI/CD pipelines, progressive rollout strategies, and rollback automation.
  • Monitoring and observability expertise, including alerting for uptime, error rates, latency regression, and resource exhaustion.
  • Ability to create and maintain evidence-based maturity assessments using trailing 90-day data windows.
  • Experience managing vendors and dependencies, including validated escalation contacts reachable within 5 minutes.
  • Curiosity, autonomy, accountability, and a desire to grow as an engineer.
  • 5+ years of software development experience.

Technologies

  • Kotlin
  • Modern Java (11+)
  • HTTP, JSON, gRPC, and Protocol Buffers
  • MySQL, Vitess, and DynamoDB
  • Event-driven architectures
  • Datadog
  • LaunchDarkly
  • Terraform
  • Kubernetes
  • Istio/Envoy
  • Amazon Web Services

Benefits

  • Remote work
  • Medical insurance
  • Flexible time off
  • Retirement savings plans
  • Modern family planning benefits

Skills

Kotlin, Java, gRPC, Protocol Buffers, MySQL, Vitess, DynamoDB, Datadog, Launchdarkly, Terraform, Kubernetes, Istio, Envoy, Amazon Web Services, CI/CD

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Clickhouse

Clickhouse

Remote

Senior Cloud Software Engineer - Efficiency Engineering
No salary listedRemote5+ YOEDevOps / SRE

Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.

Clickhouse

Clickhouse

Remote

Senior Cloud Software Engineer - Efficiency Engineering
No salary listedRemote5+ YOEDevOps / SRE

Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.

Clickhouse

Clickhouse

Singapore
Release Engineer - Data Plane Internal Tooling and Productivity
No salary listedRemote5+ YOEDevOps / SRE

Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.