Skip to content
CoinbaseCoinbase

Senior Software Engineer, Compute Platform

Build and operate Kubernetes-based compute orchestration infrastructure used across Coinbase, while developing developer tooling, automation, and AI-enabled workflows. The role requires 5+ years of software engineering experience, including substantial experience operating Kubernetes or comparable systems in production.

About the job

Responsibilities

  • Own the design, build, and operation of Kubernetes cluster management tooling and automation that keeps the compute platform reliable and self-healing at scale.
  • Build developer-facing tooling and workflows that improve how engineers interact with Kubernetes, including AI-driven processes and support.
  • Deliver compute capabilities such as one-off jobs, cron scheduling, deployment strategies, EFS support, and automated right-sizing.
  • Automate operational toil, reduce on-call burden, and improve platform observability and incident response.
  • Partner with Security, Reliability, and Observability teams to meet standards for security, uptime, and performance.

Requirements

  • 5+ years of software engineering experience, including 3+ years building and operating Kubernetes or similar compute orchestration systems such as Mesos, Nomad, or ECS.
  • Production experience with AWS and/or Google Cloud infrastructure services, including EC2, EKS, IAM, VPC, and networking, at scale.
  • Ability to design, implement, and operate distributed infrastructure systems; diagnose complex failures; and drive root-cause resolution.
  • Hands-on experience with the CNCF ecosystem, including Helm, Prometheus, Argo CD, and Envoy.
  • Experience applying AI tooling to infrastructure workflows to improve automation, developer productivity, or operational efficiency.
  • Responsible use of generative AI with human oversight to deliver business-ready outputs and measurable improvements in workflow efficiency, cost, and quality.

Compensation and Benefits

  • Annual base salary: 191,100–191,100 CAD.
  • Total compensation may also include equity, bonus eligibility, and medical, dental, and vision benefits.

Skills

Kubernetes, AWS, GCP, Amazon Ec2, Amazon Eks, IAM, Vpc, Networking, Helm, Prometheus, Argo Cd, Envoy, Istio, Distributed Systems, Generative AI

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Kindred

Kindred

United States
Senior Infrastructure Engineer
$170k+/yrRemote5+ YOEDevOps / SRE

Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Gumloop

Gumloop

San Francisco, CA
Senior Infrastructure Engineer
$150k+/yrOn-siteDevOps / SRE

Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.

Okta

Okta

Bellevue, WA
Senior Site Reliability Engineer
$147k+/yrHybrid5+ YOEDevOps / SRE

The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.