Senior Software Engineer, Core Reliability
Senior software engineer focused on improving production reliability, deployment safety, configuration and secrets management, and scalability across Coinbase’s service environment. The role requires 5+ years of experience with distributed systems, Ruby or Go, Terraform, cloud platforms, and observability tools.
About the job
Responsibilities
- Design and deliver reliability projects and features that improve resiliency across Coinbase's service environment in partnership with other engineering teams.
- Partner with critical T0/T1 services to understand architecture, improve scalability, and reduce operational toil.
- Build and enhance systems that securely manage service configurations and secrets at scale.
- Improve canary-based release systems and expand deployment capabilities to support thousands of services and hundreds of daily deployments with fewer incidents.
- Drive reliability best practices and strengthen reliability culture across engineering teams.
Requirements
- 5+ years of software engineering experience designing, building, and maintaining production services in service-oriented architectures.
- Experience with Ruby, Go, Terraform, and cloud platforms such as AWS, GCP, or Azure.
- Ability to design and operate reliable, high-throughput, low-latency distributed systems at scale.
- Track record of writing well-tested, production-quality code.
- Experience with observability and monitoring tools such as Kibana and Datadog.
- Ability to debug complex production issues, tune system performance, and reduce incident frequency.
- Experience communicating architecture decisions to cross-functional engineering stakeholders.
- Ability to participate in on-call rotations and respond to issues outside normal business hours.
- Responsible use of generative AI with human oversight to deliver business-ready outputs and improve workflow efficiency, cost, and quality.
Compensation and Benefits
- Annual base salary: 191,100–191,100 CAD.
- Total compensation may also include equity, bonus eligibility, and medical, dental, and vision benefits.
Skills
Ruby, Go, Terraform, AWS, GCP, Microsoft Azure, Distributed Systems, Service-Oriented Architecture, Kibana, Datadog, Observability, Canary Deployments, Secrets Management, Production Monitoring, Generative AI
Similar jobs
DevOps / SRE jobsThe Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.
The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.