Skip to content
CoinbaseCoinbase

Senior Software Engineer, Core Reliability

Senior software engineer focused on improving production reliability, deployment safety, configuration and secrets management, and scalability across Coinbase’s service environment. The role requires 5+ years of experience with distributed systems, Ruby or Go, Terraform, cloud platforms, and observability tools.

About the job

Responsibilities

  • Design and deliver reliability projects and features that improve resiliency across Coinbase's service environment in partnership with other engineering teams.
  • Partner with critical T0/T1 services to understand architecture, improve scalability, and reduce operational toil.
  • Build and enhance systems that securely manage service configurations and secrets at scale.
  • Improve canary-based release systems and expand deployment capabilities to support thousands of services and hundreds of daily deployments with fewer incidents.
  • Drive reliability best practices and strengthen reliability culture across engineering teams.

Requirements

  • 5+ years of software engineering experience designing, building, and maintaining production services in service-oriented architectures.
  • Experience with Ruby, Go, Terraform, and cloud platforms such as AWS, GCP, or Azure.
  • Ability to design and operate reliable, high-throughput, low-latency distributed systems at scale.
  • Track record of writing well-tested, production-quality code.
  • Experience with observability and monitoring tools such as Kibana and Datadog.
  • Ability to debug complex production issues, tune system performance, and reduce incident frequency.
  • Experience communicating architecture decisions to cross-functional engineering stakeholders.
  • Ability to participate in on-call rotations and respond to issues outside normal business hours.
  • Responsible use of generative AI with human oversight to deliver business-ready outputs and improve workflow efficiency, cost, and quality.

Compensation and Benefits

  • Annual base salary: 191,100–191,100 CAD.
  • Total compensation may also include equity, bonus eligibility, and medical, dental, and vision benefits.

Skills

Ruby, Go, Terraform, AWS, GCP, Microsoft Azure, Distributed Systems, Service-Oriented Architecture, Kibana, Datadog, Observability, Canary Deployments, Secrets Management, Production Monitoring, Generative AI

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Kindred

Kindred

United States
Senior Infrastructure Engineer
$170k+/yrRemote5+ YOEDevOps / SRE

Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Gumloop

Gumloop

San Francisco, CA
Senior Infrastructure Engineer
$150k+/yrOn-siteDevOps / SRE

Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.

Okta

Okta

Bellevue, WA
Senior Site Reliability Engineer
$147k+/yrHybrid5+ YOEDevOps / SRE

The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.