Staff Software Engineer, Infrastructure
Leads the design and production adoption of self-service infrastructure platforms, multi-region cloud foundations, and reliable delivery workflows. The role requires 8+ years of hands-on software or platform engineering experience, strong Go or comparable programming skills, and deep expertise in a core infrastructure domain.
About the job
Responsibilities
- Turn ambiguous infrastructure problems into proposals, RFCs, and architecture reviews that align teams.
- Design self-service platform capabilities and APIs, primarily in Go, for onboarding, provisioning, deployment, observability defaults, and day-2 operations.
- Establish delivery standards using Terraform, GitOps with Argo CD, progressive rollout, and testing.
- Build trusted continuous-deployment workflows.
- Evolve multi-tenant EKS foundations for reliability, security, scale, and cost, including Envoy Gateway ingress, traffic routing, and multi-region, cross-account connectivity.
- Improve SLOs, alerting, and incident follow-up using Grafana Cloud.
- Drive platform adoption and measure provisioning speed, team self-service, and operational reliability.
- Shape safe, auditable, human-reviewed AI-assisted operational workflows.
- Improve alert enrichment, incident context gathering, runbook-assisted diagnosis, remediation recommendations, and onboarding assistants.
- Participate in on-call after onboarding and shadowing, and improve on-call health through better alerts, runbooks, reduced toil, and blameless postmortems.
Requirements
- 8+ years of professional, hands-on, full-time software engineering experience in backend, infrastructure, or platform engineering.
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Strong software engineering ability in Go or a similar language, including design, testing, debugging, review, and maintainability.
- Experience designing, shipping, and operating cloud services or infrastructure platforms in production.
- Deep expertise in at least one of Kubernetes, networking, cloud platforms, reliability engineering, or developer platforms.
- Solid Linux, networking, and production-operations fundamentals.
- Experience setting technical direction and leading cross-team work.
- Clear written and verbal communication in a remote environment, including RFCs, design documents, and incident writeups.
Nice to Have
- EKS and ingress, CNI, or service-mesh experience.
- Observability experience with OpenTelemetry, Prometheus, or Grafana.
- CI/CD and progressive delivery experience, including GitHub Actions, Argo CD, or canary deployments.
- Experience leading migrations or adoption programs across teams.
Compensation & Benefits
- Canada compensation: CA$238,250–CA$382,250 plus equity.
- United States compensation: $170,350–$275,550 plus equity.
- Freedom and flexibility.
- Quarterly Whaleness Days and an end-of-year Whaleness break.
- Home office setup support.
- 16 weeks of paid parental leave after six months of employment.
- Technology stipend equivalent to $100 USD net per month.
- PTO plan.
- Training stipend for conferences, courses, and classes.
- Equity participation.
- Medical benefits, retirement, and holidays vary by country.
- Remote-first culture with offices in Seattle and Paris.
Skills
Go, Kubernetes, Terraform, Argo Cd, Amazon Eks, Envoy Gateway, Grafana Cloud, OpenTelemetry, Prometheus, GitHub Actions, Linux, GitOps, Cross-Account Networking, Progressive Delivery, SLOs
Similar jobs
DevOps / SRE jobsOwns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.
Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.
Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.