Senior engineer building self-service internal platform capabilities at Docker, focusing on multi-region networking, continuous deployment, EKS foundations, and AI-assisted operations to enable faster, safer provisioning for engineering teams.
161k – 261k/yr
Remote6+ YOEDevOps / SRE
About the role
Responsibilities
Turn ambiguous infrastructure problems into clear designs and working systems, contributing to RFCs and architecture reviews.
Build self-service capabilities and platform APIs (primarily in Go) for onboarding, provisioning, deployment, observability defaults, and day-2 operations.
Apply and help shape delivery standards with Terraform, GitOps on Argo CD, progressive rollout, and strong testing.
Strengthen the multi-tenant EKS foundations for reliability, security, scale, and cost, including Envoy Gateway ingress, traffic routing, and multi-region, cross-account connectivity.
Improve SLOs, alerting, and incident follow-up on Grafana Cloud.
Help shape AI-assisted and agentic workflows for alert enrichment, incident context-gathering, runbook-assisted diagnosis and remediation, and onboarding assistants.
Participate in on-call rotation, improve alerts, runbooks, and postmortems.
Qualifications
6+ years of hands-on software engineering in backend, infrastructure, or platform engineering.
Strong software engineering in Go or similar language, including design, testing, debugging, review, and maintainability.
Track record of building, shipping, and operating cloud services or infrastructure in production.
Deep expertise in at least one of: Kubernetes, networking, cloud platforms, reliability engineering, or developer platforms.
Solid Linux and production-ops fundamentals.
Experience shaping technical direction and working across teams.
Clear written and verbal communication in a remote environment.
Bachelor’s in CS/Engineering or equivalent.
Nice-to-Haves
EKS and ingress/CNI/service-mesh experience.
Observability with OpenTelemetry, Prometheus, Grafana.
CI/CD and progressive delivery (GitHub Actions, Argo CD, canaries).
Driving migrations or adoption programs across teams.
Compensation & Benefits
United States: $160,900 – $260,700 + equity
Freedom & flexibility; fit your work around your life.
Designated quarterly Whaleness Days plus end of year Whaleness break.
Home office setup.
16 weeks of paid Parental leave (after 6 months of employment).
Technology stipend equivalent to $100 USD net/month.
PTO plan.
Training stipend for conferences, courses and classes.
Equity.
Medical benefits, retirement and holidays vary by country.
Senior SRE improves platform reliability using AI-driven automation, leads incident response and oncall for critical services, and implements safe deployment practices. Requires 5+ years experience in backend/platform engineering, observability, and high-availability systems.
161k – 284k/yrOn-site5+ YOEDevOps / SRE
Senior Software Engineer, Production Engineering
HarveySan Francisco, CA +1
Build and operate Harvey's core production infrastructure powering AI workloads, including Kubernetes, compute fleets, networking, and orchestration platforms. Drive reliability, scalability, security, and cost efficiency for rapidly growing LLM infrastructure while partnering across engineering teams.
161k – 242k/yrHybrid5+ YOEDevOps / SRE
Senior Site Reliability Engineer
TalkiatryUnited States
Join as the first SRE to define reliability practices, SLOs, observability, and toil reduction across six product teams at a leading mental health platform. 7+ years software/infra engineering with hands-on SRE experience required; product teams retain on-call ownership.
160k – 185k/yrRemote7+ YOEDevOps / SRE
Senior Infrastructure Engineer
AurelianSeattle, WA
Senior Infrastructure Engineer building analytics, observability, and developer tooling for Aurelian's real-time AI agents used in 911 emergency response centers. Requires 4+ years in infrastructure/platform/backend roles with experience in reliability and scale.
160k – 220k/yrOn-site4+ YOEDevOps / SRE
Senior Microsoft Cloud Infrastructure Engineer
CrusoeSan Francisco, CA
Senior Cloud Infrastructure Engineer owning design, implementation, and management of Microsoft 365, Entra ID, Azure, and Azure Arc hybrid infrastructure. Requires 8+ years infrastructure experience with deep Azure/M365 expertise, IaC, Windows Server admin, and hands-on data center hardware work.